New Model Releases A
Showing 61–90 of 326
-
TransMem: Transforming Hidden States into Memory for Large Language Models
-
GoldenRetriever: Non-Interactive Homomorphic Encrypted Retrieval for Privacy-Preserving RAG
-
Mixture-of-Translators: Translating KV Caches Across Heterogeneous Large Language Models
-
Thinking Machines、軽量モデル「Inkling-Small」正式公開 サイズ4分の1で「Inkling」に匹敵する性能Thinking Machines releases Inkling-Small, matching Inkling at 1/4 the sizeThinking Machines Lab released the final version of Inkling-Small, an open-weight AI model. At a quarter the size of its predecessor, the company says data improvements and reinforcement learning let it match the larger Inkling on tasks such as code generation.
-
FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation
-
Google、ロボット向けAI「Gemini Robotics 2」発表 ヒューマノイドの全身制御や指先作業を実現Google unveils Gemini Robotics 2 for whole-body and fine fingertip controlGoogle and Google DeepMind announced Gemini Robotics 2, a family of robotics AI models supporting humanoid whole-body control, fine fingertip manipulation, and multi-robot collaboration. The lineup includes the ER 2 reasoning model that acts as a high-level brain, plus lighter variants.
-
Claudeが評価環境から実在企業に不正アクセス――Anthropic、3件のインシデントを公表Anthropic: Claude mistakenly accessed three real companies' infra during evalAnthropic disclosed that, during a cybersecurity evaluation, its Claude model reached the open internet through a misconfigured path and mistakenly accessed the production infrastructure of three real organizations. It published the three incidents, where an exercise environment unexpectedly touched live systems.
-
Token-Level Diagnosis of Sycophancy in LLMs with Attribution-Guided Steering
-
Cohere signs EU Code of Practice on Transparency of AI-Generated ContentCohere signs EU Code of Practice on AI content transparencyCohere said it signed the EU Code of Practice on Transparency of AI-Generated Content, joining other companies committing to clearer labeling and provenance for AI outputs. The move signals alignment with Europe's emerging AI governance framework.
-
Advancing the price-performance frontier with GPT‑5.6OpenAI slashes GPT-5.6 prices: Luna down 80%, Terra down 20%OpenAI announced steep price cuts for GPT-5.6, with Luna dropping 80% and Terra 20%. The company credits GPT-5.6 Sol for enabling the reduction by optimizing load balancing and even the model's forward pass, the computation that turns inputs into next-token predictions.
-
OpenAI、「GPT-5.6 Luna」を80%値下げ モデル自身による効率化でコスト削減OpenAI cuts 'GPT-5.6 Luna' price by 80% via model-driven efficiencyOpenAI cut the price of 'Luna' in its GPT-5.6 family by 80%, saying efficiency gains achieved by the model itself lowered costs. The move makes a high-performance model considerably cheaper, reflecting OpenAI's recent emphasis on price-performance.
-
llm 0.32rc2llm 0.32rc2 switches its default model to GPT-5.6 LunaSimon Willison released llm 0.32rc2, fixing a dependency issue and changing the default model for users who have not set one from GPT-4o mini to the newer, more capable GPT-5.6 Luna. Luna is slightly more expensive but a notable upgrade.
-
TextCloak: Thwarting Unauthorized LLM Exploitation via RL-Driven Unlearnable Text
-
Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions
-
Rolling With Resistance: Preference-Optimized LLM Counselors Can Trade Goal Persistence for Relational Attunement in Motivational Interviewing
-
Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation
-
Self-Supervised Skill Optimization
-
AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis
-
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
-
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
-
JetBrains、AIが少ないトークンでコンテキストを取得しやすく、よりよいコード生成を可能にする「JetBrains Context」発表JetBrains unveils 'JetBrains Context' to feed AI agents code context efficientlyJetBrains announced JetBrains Context, a service that builds an intelligence layer over code repositories. By supplying AI agents with the right code context using fewer tokens, it aims to enable better code generation from agentic coding tools.
-
VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation
-
$β$-OPSD: Deriving with Policy Optimization, Training with Self-Distillation
-
DualG-MRAG: Decoupling Macro-Reasoning and Micro-Matching for Multimodal Retrieval-Augmented Generation
-
Rethinking Inference-Time Scaling in Local Computer-Use Agents: Failure Modes and Compute Tradeoffs
-
Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering
-
ORCA-bench: How Ready Are Language Model Agents for Oncall?
-
ScaFE: Data-Efficient Scar Classification with LLM-Generated Clinical Feature Programs
-
Graph Neural Network Force Fields for Spin Dynamics in Metallic Magnets
-
MANTA: Multi-Agent Network Topology Adaptation for Self-Evolving Multi-Agent Systems