NVIDIA took the top spot on an agentic-coding benchmark. The composition leans academic—one official NVIDIA source and four arXiv papers—so it reads as research and evaluation rather than a product launch, anchored on a quantitative benchmark ranking. The arena measures 'implementation ability as an agent'—not just writing code but iterating over plan, execute, and fix; the through-line is a shift of evaluation from single-model generation quality toward autonomously running long workflows. But a benchmark lead is a result under specific conditions—real-world effectiveness, reproducibility, and the durability of the gap over other models can't be concluded from ranking alone.
NVIDIA tops agentic coding benchmark
NVIDIA tops agentic coding benchmark
NVIDIA Achieves Leading Agentic Coding Performance on First Agentic AI Benchmark
NVIDIA tops first agentic AI benchmark for agentic coding performance
Academic (arxiv etc.) 58 ▾
CORA aligns reasoning and answers in multimodal RLVR
Route-specialized dual adapters for memory-assisted knowledge editing
Abstracting Cross-Domain Action Sequences into Interpretable Workflows
Abstracting cross-domain action sequences into interpretable workflows
Zero-shot generalization of transformer neural operators to larger domains
Zero-shot generalization of transformer neural operators to larger domains
From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI
From chatbot to digital colleague: the shift to persistent autonomous AI
A Fixed-Point Neural Operator for Size- and Functional-Transferable Hamiltonian Prediction
A fixed-point neural operator for transferable Hamiltonian prediction
EM-NeSy: Expectation Maximization for Neurosymbolic Learning
EM-NeSy applies expectation maximization to neurosymbolic learning
A theoretical model for task routing in mixture-of-expert transformers
A theoretical model of task routing in mixture-of-expert transformers
Elastic Queries Reinforcement Learning: Self-Aware Policy Execution for VLA Models
Elastic Queries RL: self-aware policy execution for VLA models
ScoreGate: adaptive chunk selection for RAG via dual-score fusion
Decoupled Mixture-of-Experts for Parametric Knowledge Injection
Decoupled mixture-of-experts for parametric knowledge injection
Implicit Reasoning for Large Language Model-based Generative Recommendation
Implicit reasoning for LLM-based generative recommendation
CoRe: a reward-finetuned LLM query rewriter for web video search
Knowledge Graph Enhanced Memory-Augmented Retrieval for Long Context Modeling
Knowledge-graph-enhanced memory-augmented retrieval for long context
Operadic consistency: a label-free signal for compositional reasoning failures in LLMs
Operadic consistency flags LLM compositional reasoning errors label-free
Valid Inference with Synthetic Data via Task Exchangeability
Task exchangeability enables valid inference from synthetic data with guarantees
Beyond Uniform Tokens: Adaptive Compression for Time Series Language Models
Adaptive token compression streamlines time series language models
Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models
Chain-of-thought reasoning crosses a 'commitment boundary,' study shows
SCSB prunes bagging ensembles up to 96% while improving calibration
Timeflies jointly models future observation existence and values in forecasting
A2D2: Fine-Tuning Any-Length Discrete Diffusion for Adaptive Decoding
A2D2 unifies reward-guided fine-tuning for any-length discrete diffusion
Genomic priors solve the cold-start problem in personalized health AI
NetCause: Counterfactual Learning for Root Cause Analysis in Large-Scale Networks
NetCause ranks network incident root causes via counterfactuals
Graphical Causal Reasoning for Root Cause Analysis in Cloud Networks
Causal graph traversal recalls 85.7% of cloud incident root causes
Heterogeneous LiDAR fusion and re-ranking boost place recognition in fields
GF-DiT: Scheduling Parallelism for Diffusion Transformer Serving
GF-DiT makes GPU parallelism schedulable for diffusion transformers
Optical Implementation of Equilibrium Propagation Using Spatial Photonic Ising Machines
Equilibrium propagation realized on spatial photonic Ising machines
Accelerating Speculative Diffusions via Block Verification
Block verification speeds up speculative sampling for diffusion models
PolyFlow embeds polytope constraints into flow matching, projection-free
MiniMax Sparse Attention enables efficient ultra-long-context LLMs
SmartFont: Dynamic Condition Allocation for Few-Shot Font Generation
SmartFont allocates global and local conditions for few-shot font generation
Hölder++: Improving the Quality-Coherence Trade-off in Multimodal VAEs
Hölder++ improves quality-coherence trade-off in multimodal VAEs
VideoMDM: Towards 3D Human Motion Generation From 2D Supervision
VideoMDM learns 3D human motion priors from 2D video supervision
SkillCAT: Contrastive Assessment and Topology-Aware Skill Self-Evolution for LLM Agents
SkillCAT verifies and routes self-evolved skills for LLM agents
SICI: A Semantic-Pragmatic Complexity Index Reveals Regime Shifts in LLM Stance Detection
SICI complexity index reveals regime shifts in LLM stance detection
MiniPIC: Flexible Position-Independent Caching in <100LOC
MiniPIC adds position-independent KV caching to vLLM in <100 LOC
Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models
Reroute replaces visual-token removal with recoverable routing in VLMs
Context-Driven Incremental Compression for Multi-Turn Dialogue Generation
C-DIC compresses multi-turn dialogue context incrementally for stability
Doc-to-Atom: Learning to Compile and Compose Memory Atoms
Doc2Atom decomposes documents into composable micro-LoRA memory atoms
System Report for CCL25-Eval Task 5: New Dataset and LoRA-Fine-Tuned Qwen2.5
PoetryQwen specializes classical Chinese poetry appreciation via CCPoetry-49K
TAHOE: Text-to-SQL with Automated Hint Optimization from Experience
Tahoe learns hints from experience to optimize production Text-to-SQL
Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling
Bebop boosts MTP acceptance via rejection sampling to speed RL training
Latent World Recovery for Multimodal Learning with Missing Modalities
LWR recovers a latent world for multimodal learning with missing modalities
CHORUS: Decentralized Multi-Embodiment Collaboration with One VLA Policy
CHORUS controls multi-robot teams with one decentralized VLA policy
ALIGNBEAM : Inference-Time Alignment Transfer via Cross-Vocabulary Logit Mixing
ALIGNBEAM transfers safety logits across model vocabularies
Measuring Semantic Progress in Multi-turn Dialogue via Information Gain
An information-gain metric measures semantic progress in multi-turn dialogue
Harness In-Context Operator Learning with Chain of Operators
CHOP chains operators to generalize a frozen ICON to OOD tasks
Mathematical perspective on genetic algorithms with optimization guided operators
A mathematical model of genetic algorithms with optimization-guided operators
VIA-SD: Verification via Intra-Model Routing for Speculative Decoding
VIA-SD speeds speculative decoding via intra-model verifier routing
Re-evaluating Confidence Remasking in Masked Diffusion Language Models
Re-evaluation: WINO remasking adds little in masked diffusion LLMs
Can News Predict the Market? Limits of Zero-Shot Financial NLP and the Role of Explainable AI
Zero-shot financial NLP fails to beat baselines at predicting markets
Adaptive Multi-Resolution Procedural Knowledge Compression for Large Language Models
SKIM compresses procedural LLM skills via adaptive multi-resolution tokens
A Resource for Enthymeme Detection in Controversial Political Discourse
New dataset enables enthymeme detection in political discourse
Language robustness in VLA models is a step-wise control problem
Beyond representational alignment with brain-guided language models for robust reasoning
Brain-guided language models strengthen robust deductive reasoning
Fine-tuning Multi-modal LLMs with ART: Art-based Reinforcement Training
ART fine-tunes frozen MLLMs by optimizing only the visual input
MultiToP patches visual tokens to cut video-LMM hallucinations
Fast Speech Foundation Model Distillation Using Interleaved Stacking
Interleaved stacking accelerates speech foundation model distillation