NVIDIA × Inference & Efficiency

NVIDIA tops agentic coding benchmark

NVIDIA tops agentic coding benchmark

✎ Story body

NVIDIA took the top spot on an agentic-coding benchmark. The composition leans academic—one official NVIDIA source and four arXiv papers—so it reads as research and evaluation rather than a product launch, anchored on a quantitative benchmark ranking. The arena measures 'implementation ability as an agent'—not just writing code but iterating over plan, execute, and fix; the through-line is a shift of evaluation from single-model generation quality toward autonomously running long workflows. But a benchmark lead is a result under specific conditions—real-world effectiveness, reproducibility, and the durability of the gap over other models can't be concluded from ranking alone.

▲ Official & Press
Official

NVIDIA Achieves Leading Agentic Coding Performance on First Agentic AI Benchmark

NVIDIA Developer Blog ・ 2026-06-12 ・ 📌

NVIDIA tops first agentic AI benchmark for agentic coding performance

Academic (arxiv etc.) 58 ▾
Academic

CORA: Analyzing and bridging thinking-answer gap in Multimodal RLVR via Consistency-Oriented Reasoning Alignment

arXiv cs.CL (Computation and Language) ・ 2026-06-12

CORA aligns reasoning and answers in multimodal RLVR

Academic

When to Write and When to Suppress: Route-Specialized Dual Adapters for Memory-Assisted Knowledge Editing

arXiv cs.LG (Machine Learning) ・ 2026-06-12

Route-specialized dual adapters for memory-assisted knowledge editing

Academic

Abstracting Cross-Domain Action Sequences into Interpretable Workflows

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-12

Abstracting cross-domain action sequences into interpretable workflows

Academic

Zero-shot generalization of transformer neural operators to larger domains

arXiv cs.LG (Machine Learning) ・ 2026-06-12

Zero-shot generalization of transformer neural operators to larger domains

Academic

From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-12

From chatbot to digital colleague: the shift to persistent autonomous AI

Academic

A Fixed-Point Neural Operator for Size- and Functional-Transferable Hamiltonian Prediction

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-12

A fixed-point neural operator for transferable Hamiltonian prediction

Academic

EM-NeSy: Expectation Maximization for Neurosymbolic Learning

arXiv cs.LG (Machine Learning) ・ 2026-06-12

EM-NeSy applies expectation maximization to neurosymbolic learning

Academic

A theoretical model for task routing in mixture-of-expert transformers

arXiv cs.LG (Machine Learning) ・ 2026-06-12

A theoretical model of task routing in mixture-of-expert transformers

Academic

Elastic Queries Reinforcement Learning: Self-Aware Policy Execution for VLA Models

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-12

Elastic Queries RL: self-aware policy execution for VLA models

Academic

ScoreGate: Adaptive Chunk Selection for Retrieval-Augmented Generation via Dual-Score Statistical Fusion

arXiv cs.CL (Computation and Language) ・ 2026-06-12

ScoreGate: adaptive chunk selection for RAG via dual-score fusion

Academic

Decoupled Mixture-of-Experts for Parametric Knowledge Injection

arXiv cs.CL (Computation and Language) ・ 2026-06-12

Decoupled mixture-of-experts for parametric knowledge injection

Academic

Implicit Reasoning for Large Language Model-based Generative Recommendation

arXiv cs.CL (Computation and Language) ・ 2026-06-12

Implicit reasoning for LLM-based generative recommendation

Academic

CoRe: A Continuously Reward-Finetuned LLM Query Rewriter for Multi-Stage Context-Aware Relevance in Web-Scale Video Search

arXiv cs.CL (Computation and Language) ・ 2026-06-12

CoRe: a reward-finetuned LLM query rewriter for web video search

Academic

Knowledge Graph Enhanced Memory-Augmented Retrieval for Long Context Modeling

arXiv cs.CL (Computation and Language) ・ 2026-06-12

Knowledge-graph-enhanced memory-augmented retrieval for long context

Academic

Operadic consistency: a label-free signal for compositional reasoning failures in LLMs

arXiv cs.CL (Computation and Language) ・ 2026-06-11

Operadic consistency flags LLM compositional reasoning errors label-free

Academic

Valid Inference with Synthetic Data via Task Exchangeability

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-11

Task exchangeability enables valid inference from synthetic data with guarantees

Academic

Beyond Uniform Tokens: Adaptive Compression for Time Series Language Models

arXiv cs.CL (Computation and Language) ・ 2026-06-11

Adaptive token compression streamlines time series language models

Academic

Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-11

Chain-of-thought reasoning crosses a 'commitment boundary,' study shows

Academic

Simplex-Constrained Sparse Bagging: Transitioning from Uniform Priors to Sparse Posteriors in Ensemble Learning

arXiv cs.LG (Machine Learning) ・ 2026-06-11

SCSB prunes bagging ensembles up to 96% while improving calibration

Academic

Existence Precedes Value: Joint Modeling of Observational Existence and Evolving States in Time Series Forecasting

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-11

Timeflies jointly models future observation existence and values in forecasting

Academic

A2D2: Fine-Tuning Any-Length Discrete Diffusion for Adaptive Decoding

arXiv cs.LG (Machine Learning) ・ 2026-06-11

A2D2 unifies reward-guided fine-tuning for any-length discrete diffusion

Academic

Is It You or Your Environment? A Bayesian Inference Framework for Genomically-Anchored Personalized Physiological Interpretation

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-11

Genomic priors solve the cold-start problem in personalized health AI

Academic

NetCause: Counterfactual Learning for Root Cause Analysis in Large-Scale Networks

arXiv cs.LG (Machine Learning) ・ 2026-06-11

NetCause ranks network incident root causes via counterfactuals

Academic

Graphical Causal Reasoning for Root Cause Analysis in Cloud Networks

arXiv cs.LG (Machine Learning) ・ 2026-06-11

Causal graph traversal recalls 85.7% of cloud incident root causes

Academic

Heterogeneous LiDAR Early Fusion and Learned Re-Ranking Strategy for Robust Long-Term Place Recognition in Unstructured Environments

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-11

Heterogeneous LiDAR fusion and re-ranking boost place recognition in fields

Academic

GF-DiT: Scheduling Parallelism for Diffusion Transformer Serving

arXiv cs.LG (Machine Learning) ・ 2026-06-11

GF-DiT makes GPU parallelism schedulable for diffusion transformers

Academic

Optical Implementation of Equilibrium Propagation Using Spatial Photonic Ising Machines

arXiv cs.LG (Machine Learning) ・ 2026-06-11

Equilibrium propagation realized on spatial photonic Ising machines

Academic

Accelerating Speculative Diffusions via Block Verification

arXiv cs.LG (Machine Learning) ・ 2026-06-11

Block verification speeds up speculative sampling for diffusion models

Academic

PolyFlow: Safe and Efficient Polytope-Constrained Flow Matching with Constraint Embedding and Projection-free Update

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-11

PolyFlow embeds polytope constraints into flow matching, projection-free

Academic

MiniMax Sparse Attention

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-11

MiniMax Sparse Attention enables efficient ultra-long-context LLMs

Academic

SmartFont: Dynamic Condition Allocation for Few-Shot Font Generation

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-11

SmartFont allocates global and local conditions for few-shot font generation

Academic

Hölder++: Improving the Quality-Coherence Trade-off in Multimodal VAEs

arXiv cs.LG (Machine Learning) ・ 2026-06-11

Hölder++ improves quality-coherence trade-off in multimodal VAEs

Academic

VideoMDM: Towards 3D Human Motion Generation From 2D Supervision

arXiv cs.LG (Machine Learning) ・ 2026-06-11

VideoMDM learns 3D human motion priors from 2D video supervision

Academic

SkillCAT: Contrastive Assessment and Topology-Aware Skill Self-Evolution for LLM Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-11

SkillCAT verifies and routes self-evolved skills for LLM agents

Academic

SICI: A Semantic-Pragmatic Complexity Index Reveals Regime Shifts in LLM Stance Detection

arXiv cs.CL (Computation and Language) ・ 2026-06-11

SICI complexity index reveals regime shifts in LLM stance detection

Academic

MiniPIC: Flexible Position-Independent Caching in <100LOC

arXiv cs.CL (Computation and Language) ・ 2026-06-11

MiniPIC adds position-independent KV caching to vLLM in <100 LOC

Academic

Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-10

Reroute replaces visual-token removal with recoverable routing in VLMs

Academic

Context-Driven Incremental Compression for Multi-Turn Dialogue Generation

arXiv cs.CL (Computation and Language) ・ 2026-06-10

C-DIC compresses multi-turn dialogue context incrementally for stability

Academic

Doc-to-Atom: Learning to Compile and Compose Memory Atoms

arXiv cs.CL (Computation and Language) ・ 2026-06-10

Doc2Atom decomposes documents into composable micro-LoRA memory atoms

Academic

System Report for CCL25-Eval Task 5: New Dataset and LoRA-Fine-Tuned Qwen2.5

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-10

PoetryQwen specializes classical Chinese poetry appreciation via CCPoetry-49K

Academic

TAHOE: Text-to-SQL with Automated Hint Optimization from Experience

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-10

Tahoe learns hints from experience to optimize production Text-to-SQL

Academic

Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling

arXiv cs.CL (Computation and Language) ・ 2026-06-10

Bebop boosts MTP acceptance via rejection sampling to speed RL training

Academic

Latent World Recovery for Multimodal Learning with Missing Modalities

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-10

LWR recovers a latent world for multimodal learning with missing modalities

Academic

CHORUS: Decentralized Multi-Embodiment Collaboration with One VLA Policy

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-10

CHORUS controls multi-robot teams with one decentralized VLA policy

Academic

ALIGNBEAM : Inference-Time Alignment Transfer via Cross-Vocabulary Logit Mixing

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-10

ALIGNBEAM transfers safety logits across model vocabularies

Academic

Measuring Semantic Progress in Multi-turn Dialogue via Information Gain

arXiv cs.CL (Computation and Language) ・ 2026-06-10

An information-gain metric measures semantic progress in multi-turn dialogue

Academic

Harness In-Context Operator Learning with Chain of Operators

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-10

CHOP chains operators to generalize a frozen ICON to OOD tasks

Academic

Mathematical perspective on genetic algorithms with optimization guided operators

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-10

A mathematical model of genetic algorithms with optimization-guided operators

Academic

VIA-SD: Verification via Intra-Model Routing for Speculative Decoding

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-10

VIA-SD speeds speculative decoding via intra-model verifier routing

Academic

Re-evaluating Confidence Remasking in Masked Diffusion Language Models

arXiv cs.LG (Machine Learning) ・ 2026-06-10

Re-evaluation: WINO remasking adds little in masked diffusion LLMs

Academic

Can News Predict the Market? Limits of Zero-Shot Financial NLP and the Role of Explainable AI

arXiv cs.CL (Computation and Language) ・ 2026-06-10

Zero-shot financial NLP fails to beat baselines at predicting markets

Academic

Adaptive Multi-Resolution Procedural Knowledge Compression for Large Language Models

arXiv cs.CL (Computation and Language) ・ 2026-06-10

SKIM compresses procedural LLM skills via adaptive multi-resolution tokens

Academic

A Resource for Enthymeme Detection in Controversial Political Discourse

arXiv cs.CL (Computation and Language) ・ 2026-06-10

New dataset enables enthymeme detection in political discourse

Academic

When Does Language Matter? Multilingual Instructions Reveal Step-wise Language Sensitivity in Vision-Language-Action Models

arXiv cs.CL (Computation and Language) ・ 2026-06-10

Language robustness in VLA models is a step-wise control problem

Academic

Beyond representational alignment with brain-guided language models for robust reasoning

arXiv cs.CL (Computation and Language) ・ 2026-06-10

Brain-guided language models strengthen robust deductive reasoning

Academic

Fine-tuning Multi-modal LLMs with ART: Art-based Reinforcement Training

arXiv cs.CL (Computation and Language) ・ 2026-06-10

ART fine-tunes frozen MLLMs by optimizing only the visual input

Academic

MultiToP: Learning to Patch Visual Tokens to Mitigate Hallucinations in Video Large Multimodal Models

arXiv cs.CL (Computation and Language) ・ 2026-06-10

MultiToP patches visual tokens to cut video-LMM hallucinations

Academic

Fast Speech Foundation Model Distillation Using Interleaved Stacking

arXiv cs.CL (Computation and Language) ・ 2026-06-10

Interleaved stacking accelerates speech foundation model distillation

← Story Archive