NVIDIA × 推論・効率化

NVIDIA、エージェント型コーディングで首位

NVIDIA、エージェント型コーディングで首位

✎ ストーリー本文

NVIDIA がエージェント型コーディングのベンチマークで首位に立った。発生元は NVIDIA 公式1件に arXiv 4件が続く学術寄りの構成で、製品発表というより研究と評価が主導する動き――ベンチマーク上の順位という定量的な指標が起点になっている。コードを書くだけでなく、計画・実行・修正を繰り返す「エージェントとしての実装力」を測る土俵で、本流はモデル単体の生成品質から、長い工程を自律的に回す能力へ評価軸が移りつつある点にある。ただしベンチマーク首位は特定条件下での結果であり、実際の開発現場での有効性や再現性、他モデルとの差の持続性は、順位だけからは断定できない。

▲ 公式・報道
公式

NVIDIA Achieves Leading Agentic Coding Performance on First Agentic AI Benchmark

NVIDIA Developer Blog ・ 2026-06-12 ・ 📌

NVIDIA、初のエージェント型AIベンチマークでコーディング性能首位を達成

学術(arxiv ほか) 58本 ▾
学術

CORA: Analyzing and bridging thinking-answer gap in Multimodal RLVR via Consistency-Oriented Reasoning Alignment

arXiv cs.CL (Computation and Language) ・ 2026-06-12

CORA、マルチモーダルRLVRの「思考と回答のずれ」を是正

学術

When to Write and When to Suppress: Route-Specialized Dual Adapters for Memory-Assisted Knowledge Editing

arXiv cs.LG (Machine Learning) ・ 2026-06-12

知識編集の書込/抑制を切替える二重アダプタを提案

学術

Abstracting Cross-Domain Action Sequences into Interpretable Workflows

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-12

操作ログを解釈可能なワークフローへ抽象化する手法を提案

学術

Zero-shot generalization of transformer neural operators to larger domains

arXiv cs.LG (Machine Learning) ・ 2026-06-12

Transformerニューラル演算子の大領域へのゼロショット汎化

学術

From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-12

LLMが「チャットボット」から永続自律AIへ移行する転換を概念化

学術

A Fixed-Point Neural Operator for Size- and Functional-Transferable Hamiltonian Prediction

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-12

サイズ・汎関数転移可能なハミルトニアン予測の固定点ニューラル演算子

学術

EM-NeSy: Expectation Maximization for Neurosymbolic Learning

arXiv cs.LG (Machine Learning) ・ 2026-06-12

EM-NeSy、EM法でニューロシンボリック学習を一般化

学術

A theoretical model for task routing in mixture-of-expert transformers

arXiv cs.LG (Machine Learning) ・ 2026-06-12

MoE Transformerのタスクルーティングを理論モデル化

学術

Elastic Queries Reinforcement Learning: Self-Aware Policy Execution for VLA Models

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-12

VLAモデルの推論頻度を自律調整する「Elastic Queries RL」

学術

ScoreGate: Adaptive Chunk Selection for Retrieval-Augmented Generation via Dual-Score Statistical Fusion

arXiv cs.CL (Computation and Language) ・ 2026-06-12

ScoreGate、RAGの取得チャンク数を適応的に選択

学術

Decoupled Mixture-of-Experts for Parametric Knowledge Injection

arXiv cs.CL (Computation and Language) ・ 2026-06-12

分離型MoEでパラメトリックな知識注入を改善

学術

Implicit Reasoning for Large Language Model-based Generative Recommendation

arXiv cs.CL (Computation and Language) ・ 2026-06-12

暗黙的推論でLLMの生成的推薦を強化

学術

CoRe: A Continuously Reward-Finetuned LLM Query Rewriter for Multi-Stage Context-Aware Relevance in Web-Scale Video Search

arXiv cs.CL (Computation and Language) ・ 2026-06-12

CoRe、報酬微調整のクエリ書換でWeb動画検索を改善

学術

Knowledge Graph Enhanced Memory-Augmented Retrieval for Long Context Modeling

arXiv cs.CL (Computation and Language) ・ 2026-06-12

知識グラフ強化の記憶検索で長文脈を一貫処理

学術

Operadic consistency: a label-free signal for compositional reasoning failures in LLMs

arXiv cs.CL (Computation and Language) ・ 2026-06-11

LLMの合成推論の失敗をラベル不要で検出する「オペラド整合性」を提案

学術

Valid Inference with Synthetic Data via Task Exchangeability

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-11

合成データで妥当な推論を保証するtask exchangeabilityを提案

学術

Beyond Uniform Tokens: Adaptive Compression for Time Series Language Models

arXiv cs.CL (Computation and Language) ・ 2026-06-11

時系列LLMの非対称トークン圧縮で効率化、予測から異常検知まで有効

学術

Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-11

思考連鎖に『コミットメント境界』、以降の手順は無影響と判明

学術

Simplex-Constrained Sparse Bagging: Transitioning from Uniform Priors to Sparse Posteriors in Ensemble Learning

arXiv cs.LG (Machine Learning) ・ 2026-06-11

単体制約スパースバギングSCSB、アンサンブルを最大96%圧縮し較正改善

学術

Existence Precedes Value: Joint Modeling of Observational Existence and Evolving States in Time Series Forecasting

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-11

Timeflies、未来の観測有無と値を同時推定する予測枠組み

学術

A2D2: Fine-Tuning Any-Length Discrete Diffusion for Adaptive Decoding

arXiv cs.LG (Machine Learning) ・ 2026-06-11

A2D2、可変長離散拡散モデルの報酬誘導ファインチューニングを統一

学術

Is It You or Your Environment? A Bayesian Inference Framework for Genomically-Anchored Personalized Physiological Interpretation

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-11

ゲノムを事前分布に使い個別生理解釈の冷却開始を解決

学術

NetCause: Counterfactual Learning for Root Cause Analysis in Large-Scale Networks

arXiv cs.LG (Machine Learning) ・ 2026-06-11

NetCause、反実仮想学習でクラウド網障害の根本原因を順位付け

学術

Graphical Causal Reasoning for Root Cause Analysis in Cloud Networks

arXiv cs.LG (Machine Learning) ・ 2026-06-11

クラウド網障害の根本原因分析、因果グラフ探索で85.7%の再現率

学術

Heterogeneous LiDAR Early Fusion and Learned Re-Ranking Strategy for Robust Long-Term Place Recognition in Unstructured Environments

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-11

異種LiDAR早期融合と再ランクで非構造環境の場所認識を頑健化

学術

GF-DiT: Scheduling Parallelism for Diffusion Transformer Serving

arXiv cs.LG (Machine Learning) ・ 2026-06-11

GF-DiT、拡散Transformer配信のGPU並列度を動的スケジューリング

学術

Optical Implementation of Equilibrium Propagation Using Spatial Photonic Ising Machines

arXiv cs.LG (Machine Learning) ・ 2026-06-11

空間フォトニックイジングマシンで平衡伝播を光学実装

学術

Accelerating Speculative Diffusions via Block Verification

arXiv cs.LG (Machine Learning) ・ 2026-06-11

拡散モデルに投機的復号のブロック検証を導入、受理率を理論的に改善

学術

PolyFlow: Safe and Efficient Polytope-Constrained Flow Matching with Constraint Embedding and Projection-free Update

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-11

PolyFlow、多面体制約を埋込む射影不要のflow matching

学術

MiniMax Sparse Attention

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-11

MiniMax、超長文脈向けブロック疎注意MSAを発表

学術

SmartFont: Dynamic Condition Allocation for Few-Shot Font Generation

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-11

SmartFont、大域と局所条件を多段配分する少数事例フォント生成

学術

Hölder++: Improving the Quality-Coherence Trade-off in Multimodal VAEs

arXiv cs.LG (Machine Learning) ・ 2026-06-11

Hölder++、マルチモーダルVAEの品質と整合性の両立を改善

学術

VideoMDM: Towards 3D Human Motion Generation From 2D Supervision

arXiv cs.LG (Machine Learning) ・ 2026-06-11

VideoMDM、3D正解データなしで動画の2D姿勢から3D動作生成を学習

学術

SkillCAT: Contrastive Assessment and Topology-Aware Skill Self-Evolution for LLM Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-11

SkillCAT、LLMエージェントのスキル自己進化を検証付き3段階に分離

学術

SICI: A Semantic-Pragmatic Complexity Index Reveals Regime Shifts in LLM Stance Detection

arXiv cs.CL (Computation and Language) ・ 2026-06-11

意味・語用論的複雑性指標SICI、LLM立場検出のレジーム変化を解明

学術

MiniPIC: Flexible Position-Independent Caching in <100LOC

arXiv cs.CL (Computation and Language) ・ 2026-06-11

MiniPIC、100行未満でvLLMに位置非依存KVキャッシュを実装

学術

Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-10

Reroute、視覚トークンを除去でなく再ルーティングしVLM推論を効率化

学術

Context-Driven Incremental Compression for Multi-Turn Dialogue Generation

arXiv cs.CL (Computation and Language) ・ 2026-06-10

C-DIC、多ターン対話の文脈を逐次圧縮し長対話を安定化

学術

Doc-to-Atom: Learning to Compile and Compose Memory Atoms

arXiv cs.CL (Computation and Language) ・ 2026-06-10

Doc2Atom、文書を知識アトムに分解し合成的パラメトリック記憶を構築

学術

System Report for CCL25-Eval Task 5: New Dataset and LoRA-Fine-Tuned Qwen2.5

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-10

PoetryQwen、CCPoetry-49Kで古典中国詩の鑑賞を専門化

学術

TAHOE: Text-to-SQL with Automated Hint Optimization from Experience

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-10

Tahoe、経験からヒントを学びText-to-SQLを本番最適化

学術

Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling

arXiv cs.CL (Computation and Language) ・ 2026-06-10

Bebop、棄却サンプリングでMTP受理率を改善しRL学習を加速

学術

Latent World Recovery for Multimodal Learning with Missing Modalities

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-10

LWR、欠損モダリティ下で潜在世界を復元しマルチモーダル学習

学術

CHORUS: Decentralized Multi-Embodiment Collaboration with One VLA Policy

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-10

CHORUS、単一VLA方策で分散的な多体協調を実現

学術

ALIGNBEAM : Inference-Time Alignment Transfer via Cross-Vocabulary Logit Mixing

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-10

ALIGNBEAM、語彙を越えて安全アンカーのlogitを移植

学術

Measuring Semantic Progress in Multi-turn Dialogue via Information Gain

arXiv cs.CL (Computation and Language) ・ 2026-06-10

多ターン対話の意味的進展を情報利得で測る指標を提案

学術

Harness In-Context Operator Learning with Chain of Operators

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-10

CHOP、演算子連鎖で凍結ICONを分布外タスクへ汎化

学術

Mathematical perspective on genetic algorithms with optimization guided operators

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-10

最適化誘導演算子を持つ遺伝的アルゴリズムの数理的視座

学術

VIA-SD: Verification via Intra-Model Routing for Speculative Decoding

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-10

VIA-SD、モデル内ルーティングで投機的デコードを高速化

学術

Re-evaluating Confidence Remasking in Masked Diffusion Language Models

arXiv cs.LG (Machine Learning) ・ 2026-06-10

拡散言語モデルの remasking 手法 WINO、再評価で効果は限定的

学術

Can News Predict the Market? Limits of Zero-Shot Financial NLP and the Role of Explainable AI

arXiv cs.CL (Computation and Language) ・ 2026-06-10

ゼロショット金融NLPの限界、ニュースで株価予測は困難と示す

学術

Adaptive Multi-Resolution Procedural Knowledge Compression for Large Language Models

arXiv cs.CL (Computation and Language) ・ 2026-06-10

SKIM、手続き的スキルを適応的多解像度で圧縮しコスト削減

学術

A Resource for Enthymeme Detection in Controversial Political Discourse

arXiv cs.CL (Computation and Language) ・ 2026-06-10

論争的政治言説のエンチメーム検出データセットを公開

学術

When Does Language Matter? Multilingual Instructions Reveal Step-wise Language Sensitivity in Vision-Language-Action Models

arXiv cs.CL (Computation and Language) ・ 2026-06-10

VLAモデルの言語頑健性は段階ごとの制御問題と判明

学術

Beyond representational alignment with brain-guided language models for robust reasoning

arXiv cs.CL (Computation and Language) ・ 2026-06-10

脳誘導の言語モデルで頑健な演繹推論を強化

学術

Fine-tuning Multi-modal LLMs with ART: Art-based Reinforcement Training

arXiv cs.CL (Computation and Language) ・ 2026-06-10

ART、視覚入力のみ最適化で凍結MLLMをソフトトークン微調整

学術

MultiToP: Learning to Patch Visual Tokens to Mitigate Hallucinations in Video Large Multimodal Models

arXiv cs.CL (Computation and Language) ・ 2026-06-10

MultiToP、視覚トークンを修復し動画LMMの幻覚を緩和

学術

Fast Speech Foundation Model Distillation Using Interleaved Stacking

arXiv cs.CL (Computation and Language) ・ 2026-06-10

交互スタッキングで音声基盤モデルの蒸留学習を高速化

← ストーリー アーカイブ