NVIDIA laid out guidance for energy storage and power design for AI factories (large data centers). The signal pairs NVIDIA and Anthropic official sources with thin involvement from simon_willison and arXiv—vendor technical guidance, announcement-led but stepping into the physical constraint of power. The focus is how to support surging AI power consumption through power design including batteries; the through-line is the power and infrastructure side—rather than models or software—coming to the fore as the bottleneck. The compute conversation has descended to 'how to supply the electricity.' But the content is mainly design guidance—real deployment impact and standardization are to confirm.
NVIDIA outlines AI factory storage
NVIDIA outlines AI factory storage
Designing Production-Ready Battery Energy Storage Systems for AI Factories
NVIDIA details production-ready battery energy storage for AI factories
Results from the first Anthropic Public Record
Anthropic shares first Public Record survey of 52,000 Americans on AI
NVIDIA details deploying MiniMax M3 for long-context agentic workflows
Claude Fable is relentlessly proactive
Simon Willison: Claude Fable 5 is 'relentlessly proactive'
Academic (arxiv etc.) 119 ▾
AdaSR: Adaptive Streaming Reasoning with Hierarchical Relative Policy Optimization
AdaSR enables adaptive streaming reasoning for reasoning models
CORA aligns reasoning and answers in multimodal RLVR
Why generating 'trivia' is provably necessary for valuable mathematics
Beyond task performance: Decoding bioacoustic embeddings with speech features
Decoding what bioacoustic embeddings encode via speech features
Graph-structured combinatorial semi-bandits with nonlinear rewards
Which Directions Matter? Sparse Design for Affine Robust Optimization
Sparse design identifies which directions matter in robust optimization
Graph Diffusion Residuals for Control-Function Instrumental Variables
Graph diffusion residuals for control-function instrumental variables
Characterizing Cultural Localization in AI-Generated Stories
Characterizing cultural localization in AI-generated stories
Regulating the Machine Contributor: Governance and Policy Alignment in Open Source
Governance and policy alignment for AI contributors in open source
AudioDER: A Deduplication-Enhanced Reasoning Dataset for Post-Training Large Audio-Language Models
AudioDER: a deduplication-enhanced reasoning dataset for audio LLMs
SIMMER: Benchmarking Latent Failures in LLM Executable Planning with a World Model
SIMMER: benchmarking latent failures in LLM executable planning
ORCA: A Platform for Open-Source Dexterity Research
ORCA: an open-source platform for dexterity research
Rethinking Global Average Pooling: Your Classifier Is Secretly a Multi-Instance Learner
Rethinking GAP: your classifier is secretly a multi-instance learner
Provably Safe, Yet Scalable Reinforcement Learning
Provably safe yet scalable reinforcement learning
From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI
From chatbot to digital colleague: the shift to persistent autonomous AI
Manipulating audio-model explanations while predictions stay unchanged
Learning to Hear Hesitation: Continual Learning for Disfluency-Aware ASR
Learning to hear hesitation: continual learning for disfluency-aware ASR
A Low-Rank Subspace Analysis of LLM Interventions
A low-rank subspace analysis of LLM behavioral interventions
Discovery under Hypothesis Redundancy: A Geometric Theory of Discovery Bottlenecks
A geometric theory of discovery bottlenecks under hypothesis redundancy
Can Deep Neural Networks Improve Compression of Very Large Scientific Data?
Can deep neural networks improve compression of very large scientific data?
Achieving Precise Text-To-Cypher Via Grounded Knowledge Graph Data Generation
Precise Text-to-Cypher via grounded knowledge graph data generation
ScoreGate: adaptive chunk selection for RAG via dual-score fusion
Decoupled Mixture-of-Experts for Parametric Knowledge Injection
Decoupled mixture-of-experts for parametric knowledge injection
Spatio-Temporal Audio Language Modeling for Dynamic Sound Sources
Spatio-temporal audio language modeling for dynamic sound sources
Knowledge Graph Enhanced Memory-Augmented Retrieval for Long Context Modeling
Knowledge-graph-enhanced memory-augmented retrieval for long context
The Holistic Storage of Verb+Up Phrases in Text-based and Audio-based Language Models
Holistic storage of verb-up phrases in text and audio language models
EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments
EvoArena tests LLM agents in evolving environments with patch-based memory
Learning to Reason by Analogy via Retrieval-Augmented Reinforcement Fine-Tuning
RA-RFT teaches LLMs to reason by analogy via retrieval-augmented RL fine-tuning
Influcoder: Distilling Decoders' Gradient Influence Rankings into an Encoder for Data Attribution
Influcoder distills gradient influence rankings for fast data attribution
HyperTool: Beyond Step-Wise Tool Calls for Tool-Augmented Agents
HyperTool folds tool workflows into code, lifting MCP-Universe to 35.29%
Operadic consistency: a label-free signal for compositional reasoning failures in LLMs
Operadic consistency flags LLM compositional reasoning errors label-free
SkMTEB: Slovak Massive Text Embedding Benchmark and Model Adaptation
SkMTEB: first comprehensive Slovak text embedding benchmark and compact models
From Tokens to Faces: Investigating Discrete Speech Representations for 3D Facial Animation
Phonetic encoding key for speech-driven 3D facial animation, study finds
AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility
AgentBeats proposes agentified, protocol-standardized agent assessment (AAA)
Distribution-Agnostic Robust Trajectory Optimization via Chance-Constrained Reinforcement Learning
Chance-constrained RL robustifies trajectories for Earth-Mars transfer
Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models
Chain-of-thought reasoning crosses a 'commitment boundary,' study shows
Reward Modeling for Multi-Agent Orchestration
OrchRM: self-supervised reward modeling for multi-agent orchestration
EvTexture++: Event-Driven Texture Enhancement for Video Super-Resolution
EvTexture++ uses event signals for texture enhancement in video super-resolution
Uncertainty-Aware Hybrid Retrieval for Long-Document RAG
UMG-RAG treats chunk granularity as query-specific reliability for RAG
AgentRivet: an automated system for producing Rivet routines from journal publications
AgentRivet auto-generates missing Rivet routines from physics papers
Graphical Causal Reasoning for Root Cause Analysis in Cloud Networks
Causal graph traversal recalls 85.7% of cloud incident root causes
Measurement-Calibrated Multi-Camera Fusion for Vision-Based Indoor Localization
Measurement-calibrated multi-camera fusion improves indoor visual localization
Leveraging Audio-LLMs to Filter Speech-to-Speech Training Data
Audio-LLM filters speech translation data, gaining up to +1.4 ASR-BLEU
Heterogeneous LiDAR fusion and re-ranking boost place recognition in fields
Ontology Memory-Augmented ASR Correction for Long Text-Speech Interleaved Conversations
Ontology memory grounds ASR correction in long text-speech conversations
Clustering Node Attributed Networks with Graph Neural Networks and Self Learning
Self-learning rounds couple GNN embeddings and graph clustering
Mod-Guide integrates minority perspectives into LLM content moderation
SmartFont: Dynamic Condition Allocation for Few-Shot Font Generation
SmartFont allocates global and local conditions for few-shot font generation
An LLM System for Autonomous Variational Quantum Circuit Design
LLM agent framework autonomously designs variational quantum circuits
From Passive Generation to Investigation: A Proactive Scientific Peer Review Agent
ProReviewer agent proactively investigates papers for deeper peer review
Training-free sampling tweak helps diffusion models reach rare samples
PowerPhase benchmarks probabilistic forecasting on 37K grid channels
SkillCAT: Contrastive Assessment and Topology-Aware Skill Self-Evolution for LLM Agents
SkillCAT verifies and routes self-evolved skills for LLM agents
Physics-Guided Spatiotemporal Learning for Coastal Wave Peak Period Estimation from Video
Physics-guided video learning estimates coastal wave peak periods
Clipping Makes Distributed and Federated Asynchronous SGD Robust to Stragglers
Theory shows clipping makes asynchronous SGD robust to stragglers
TimeLens: on-device artifact recognition with RAG for Egyptian museum
ComAct: Reframing Professional Software Manipulation via COM-as-Action Paradigm
ComAct drives industrial CAD software via COM-as-Action paradigm
PolyAlign: Conditional Human-Distribution Alignment
PolyAlign matches LLMs to context-specific human response distributions
SICI: A Semantic-Pragmatic Complexity Index Reveals Regime Shifts in LLM Stance Detection
SICI complexity index reveals regime shifts in LLM stance detection
MemRefine: LLM-Guided Compression for Long-Term Agent Memory
MemRefine compresses agent memory to budget with LLM-judged merging
NTS-CoT curbs hallucinations in LLM news timeline summarization
Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models
Reroute replaces visual-token removal with recoverable routing in VLMs
Context-Driven Incremental Compression for Multi-Turn Dialogue Generation
C-DIC compresses multi-turn dialogue context incrementally for stability
DIRECT: When and Where Should You Allocate Test-Time Compute in Embodied Planners?
DIRECT allocates test-time compute per prompt for embodied planners
Which Models Are Our Models Built On? Auditing Invisible Dependencies in Modern LLMs
ModSleuth audits invisible recursive dependencies in modern LLMs
Verifiable Environments Are LEGO Bricks: Recursive Composition for Reasoning Generalization
RACES recursively composes verifiable RL environments to scale reasoning
Adjoint Method versus Physics-Informed Neural Networks in PDE-Constrained Inverse Problems
Adjoint optimization vs PINNs: a fair test on PDE inverse problems
Fourier Features Let Agents Learn High Precision Policies with Imitation Learning
Fourier features give point-cloud policies high-precision control
PROJECTMEM: A Local-First, Event-Sourced Memory and Judgment Layer for AI Coding Agents
projectmem adds a local-first memory layer for AI coding agents
Measuring Epistemic Resilience of LLMs Under Misleading Medical Context
MedMisBench shows LLM medical judgment collapses under misleading context
Standard Interpretable Model designs interpretability via Lagrangian mechanics
INT8 quantization of Ideogram 4.0 holds FP8 quality on consumer GPUs
DiffCold: A Diffusion-based Generative Model for Cold-Start Item Recommendation
DiffCold uses diffusion to resolve the cold-start recommendation seesaw
Re-evaluating Confidence Remasking in Masked Diffusion Language Models
Re-evaluation: WINO remasking adds little in masked diffusion LLMs
MLT-Dedup speeds large-scale online video deduplication
OpenMedReason: Scientific Reasoning Supervision for Medical Vision-Language Models
OpenMedReason releases a ~450K medical multimodal reasoning corpus
A Controlled Study of Decoding-Time Truthfulness Methods on Instruction-Tuned LLMs
CHAIR detects hallucinations from per-layer token logits
nD-RoPE: A Generalized RoPE for n-Dimensional Position Embedding
nD-RoPE generalizes RoPE to n-dimensional position embedding
Augmenting Molecular Language Models with Local $n$-gram Memory
MolGram adds local n-gram memory to boost molecular language models
InDex adapts VLA models to dexterous hands via semantic inheritance
DAM-VLA: Decoupled Asynchronous Multimodal Vision Language Action model
DAM-VLA's asynchronous modality buffers double manipulation success
FORT-Searcher: Synthesizing Shortcut-Resistant Search Tasks for Training Deep Search Agents
FORT synthesizes shortcut-resistant tasks to train deep search agents
On the Limits of LLM-as-Judge for Scientific Novelty Assessment
RQ-Bench exposes limits of LLM-as-judge for scientific novelty
Simplicity Suffices for Parameter Noise Injection in Stochastic Gradient Descent
Simple isotropic noise injection suffices for SGD, study finds
Decoding Multimodal Cues: Unveiling the Implicit Meaning Behind Hateful Videos
IARE enables explainable hateful-video detection with rationales
Multi-turn RAG with sparse retrieval and listwise reranking for SemEval-2026
Toward Generalist Autonomous Research via Hypothesis-Tree Refinement
Arbor runs autonomous research via hypothesis-tree refinement
An Ontology-Guided Multi-Anchor Graph Retrieval Framework for Traffic Legal Liability Determination
OMAGR uses ontology-guided multi-anchor retrieval for traffic-law liability
Grammar-Constrained Decoding Can Jailbreak LLMs into Generating Malicious Code
CodeSpear jailbreaks LLMs into malicious code via grammar-constrained decoding
Lius improves low-resource Kupang Malay translation via continual tuning
EEVEE: Towards Test-time Prompt Learning in the Real World for Self-Improving Agents
EEVEE: multi-dataset test-time prompt learning for self-improving agents
Data Journalist Agent: Transforming Data into Verifiable Multimodal Stories
Data2Story: a multi-agent framework for verifiable, multimodal data stories
Flaws in the LLM Automation Narrative
Human experts still outperform a frontier LLM on a coding data-analysis task
ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecurity
ABC-Bench evaluates LLM agents' biosecurity-relevant capabilities
First-Order Trajectory Matching: Fast Ensemble Predictions of Chaotic, Turbulent, Stochastic Systems
FTM learns probability-current velocity from trajectories for cheap forecasts
DMT: demographic-conditioned Transformer for cuffless BP from PPG
TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning
TRACE: tree-structured rollout budget allocation for agentic RL
Towards Autonomous Accelerator Design: FPGA Accelerator Generation with SECDA
SECDA-DSE: LLM-guided design space exploration for FPGA accelerators
PhantomBench: Benchmarking the Non-existential Threat of Language Models
PhantomBench: testing LLM hallucination on non-existent concepts
RoboNaldo: motion-guided curriculum RL for powerful humanoid soccer shots
VISTA: A Versatile Interactive User Simulation Toolkit for Agent Evaluation
VISTA: a versatile user-simulation toolkit for agent evaluation
A History-Aware Visually Grounded Critic for Computer Use Agents
HiViG: a history-aware, visually grounded critic for computer-use agents
T1-Bench: Benchmarking Multi-Scenario Agents in Real-World Domains
T1-Bench: high-fidelity evaluation of multi-domain agents across 25 domains
Generative Archetype-Grounded Item Representations for Sequential Recommendation
GenAIR: archetype-grounded item representations for sequential recommendation
LLM-augmented XAI generates natural-language explanations for networks
Democratising Camera Trap AI: An Open-Source Model for Detecting UK Mammals
Open-source UK mammal detector reaches 0.984 [email protected]
DocTrace: structure-aware on-demand hypergraph memory for long-document QA
Task Robustness via Re-Labelling Vision-Action Robot Data
TREAD: re-labelling robot vision-action data with VLMs for task robustness
Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution
Role-Agent: a single LLM as both agent and environment for co-evolution
Range Penalization: Theoretical Insights with Applications in Federated Learning
Range penalization for quantization-friendly federated learning
Ethical and Technical Limits of Deepfake Speech Datasets
Auditing 39 deepfake speech datasets reveals fairness gaps and overlap
Pose-ICL: 3D-Aware In-Context Learning for Pose-Controllable Subject Customization
Pose-ICL: 3D-aware in-context learning for pose-controllable generation
Janus: A Benchmark for Goal-Conditioned Information Distortion in LLMs
JANUS: measuring goal-conditioned information distortion in LLMs
Attention-Discounted Adaptive Sampler for Masked Diffusion Language Models
ADAS: attention-discounted reranking for masked diffusion decoding
Recovering the Zipfian Distribution in Unsupervised Term Discovery
Graph clustering beats K-means in unsupervised term discovery
N-GRPO: Embedding-Level Neighbor Mixing for Enhanced Policy Optimization
N-GRPO mixes neighbor embeddings to boost policy optimization
Infini Memory: Maintainable Topic Documents for Long-Term LLM Agent Memory
Infini Memory: topic documents for long-term LLM agent memory
How Does Reasoning Flow? Tracing Attention-Induced Information Flow for Targeted RL in LLMs
FlowTracer traces attention flow for token-level RL credit
ParaBridge: Bridging Paralinguistic Perception and Dialogue Behavior in Speech Language Models
ParaBridge links paralinguistic cues to dialogue behavior