New Model Releases A
Showing 271–300 of 315
-
Stemma: Induced Decision Regions Reveal LLM Provenance
-
Runtime Uncertainty Monitoring for LLM-Based Multi-Agent Systems Using Bayesian Networks
-
A2TTA: Anchored-and-Agile Test-Time Adaptation for Evolving Traffic Sensor Networks
-
OmniQEC: discovering practical quantum error-correcting codes by an AI scientist
-
DRIFT: Direct-Recursive Intervention-Conditioned Forecasting of ICU Physiological Trajectories
-
China to launch new cable laying vessel next month, Global Marine Group commissions new vesselChina and Global Marine to add new subsea cable-laying vesselsChina is set to launch a new cable-laying vessel next month, while Global Marine Group has commissioned its own new ship. The two vessels are expected to debut in 2026 and 2029 respectively, expanding global subsea cable installation capacity.
-
Lowering the implementation barrier of neutral-atom quantum computing with agentic workflows
-
Johnson Controls releases absorption chiller reference designJohnson Controls unveils absorption chiller reference design for DCsJohnson Controls has released a reference design for an absorption chiller, outlining how data centers can reuse waste heat for cooling. The approach aims to improve energy efficiency and cut cooling costs amid rising thermal management demands.
-
Speculate While You Reason: Teaching Agents to Predict Their Next Tool Call via Joint Agent-Speculator RL
-
WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing
-
From Deterministic to Generative Deep Learning for Urban Air Quality Reconstruction from Sparse Observations
-
Rashomon Alignment
-
Contextual Deconvolution for Variance-Stable Demand Sensing: Kernel-Modulated Operators in Promotional Retail
-
Localized Adaptation Reveals Distinct Learning Signatures in Transformers
-
AIriskEval-edu Demo: Auditing of Pedagogical Risks in Educational Explanations
-
Beyond Self-Knowledge: Propagating Uncertainty Across Reasoning and Retrieval in LLMs
-
Forensic Reproducibility Audit of a Radiology Vision-Language Model Benchmark: From Intended Protocol to Released Artifact
-
Multi-Scale Structural Features for Continual, Comprehensible Visual Recognition in a Developmental Learning Framework
-
AMPBench-MT: A Homology-Controlled Benchmark for Antimicrobial Peptide Potency, Spectrum, and Safety Prediction
-
Phase Structure in Rotary Attention: A Spectral Framework for Semantic Continuity and Execution-Boundary Governance
-
「痺れるほどにミスを繰り返す」Gemini 3.6 Flashは変わった? 公開から1週間、当初のおバカ回答を今検証するITmedia re-tests Google's Gemini 3.6 Flash a week after launchITmedia revisits Google's Gemini 3.6 Flash about a week after its release. Just after launch, the model drew attention on X for shaky accuracy, including botched numerical comparisons and an answer about a fictional fish. The piece re-runs those early wrong answers to check whether its behavior has improved.
-
Memory for Large Language Models
-
CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition
-
Where Steering Signals Come From: Activation Source Selection in Activation Steering
-
VisualPatchWorld: Code World Models as Latent Structured Representations for Planning
-
A Cross-lingual Comparison of Human and Classification Model Entrainment Behavior in Code-switched Speech Settings
-
MyoCardBench: A Real-World Data Benchmark for Evaluating Large Language Models in Clinically Authentic Cardiovascular Care Scenarios
-
moonshotai/Kimi-K3Moonshot releases weights for Kimi K3, a 2.8-trillion-parameter modelMoonshot AI released the weights for Kimi K3, a 2.8-trillion-parameter LLM that runs to 1.56TB on Hugging Face. With K2 in July 2025 the company added a modified MIT license requiring commercial products with over 100M monthly active users or $20M in monthly revenue to prominently display 'Kimi K2' in their UI. Reported via Simon Willison's link-blog; details past the license clause are truncated in the source excerpt and unconfirmed.
-
An opinionated guide to which AI to use to do stuffSimon Willison on Ethan Mollick's updated guide to which AI to useSimon Willison links to Ethan Mollick's evolving guide on which AI to use for tasks. A year ago it centered on chat (ChatGPT, Claude, Gemini); today it emphasizes agentic systems doing the equivalent of hours of human work. Gemini has dropped off Ethan's list, while modes such as ChatGPT Work/Codex and Claude Cowork/Code are explained. Note: the excerpt is truncated at the end, so later details could not be verified.