NVIDIA × 安全性・評価

NVIDIA、AR/XR向けエージェント基盤を公開

NVIDIA、AR/XR向けエージェント基盤を公開

✎ ストーリー本文

NVIDIA が AR グラスや XR デバイス向けのエージェント基盤を公開し、Hugging Face のロボット連携や DeepMind のエージェント安全化と足並みがそろった。発生元は5件すべてが公式・プラットフォーム側で、学術やコミュニティの反応はまだ薄い――研究が降りてきた週というより、基盤ベンダー各社が「XRと身体性を持つエージェント」に向けて土台を並べ始めた週だ。話題の中心はモデルの賢さより、デバイス上でエージェントをどう動かし守るかという実装層にある。ただし現時点は発表とSDK提示が主で、実機採用の広がりは次段階の確認点になる。

▲ 公式・報道
公式

Building AI Agents for AR Glasses and XR Devices with NVIDIA XR AI

NVIDIA Developer Blog ・ 2026-06-16 ・ 📌

NVIDIA、ARグラス/XR向けAIエージェント構築基盤「XR AI」を発表

コミュニティ

Show HN: Are You in the Weights?

Hacker News (Front Page) ・ 2026-06-18

Show HN『Are You in the Weights?』、LLMがあなたを認識するか測定

報道

かんぽ生命、AIで営業支援 “郵便局での一言”拾って保険提案へ 寸劇で分かる活用例

ITmedia AI+ ・ 2026-06-17

かんぽ生命、AIエージェントで営業支援を本格化

公式

From the Hugging Face Hub to robot hardware with Strands Agents and LeRobot

Hugging Face Blog ・ 2026-06-17

Strands AgentsとLeRobot、Hugging Faceのモデルをロボット実機へ

報道

「ポケカ対戦AIエージェント」開発コンテスト開始 「不完全情報ゲーム」をどう制するか

ITmedia AI+ ・ 2026-06-17

ポケモンカード対戦AIエージェントの開発コンテスト開始、不完全情報ゲームに挑む

公式

Agentic Resource Discovery: Let agents search

Hugging Face Blog ・ 2026-06-17

Hugging Face、エージェントが自ら資源を探索する手法を提案

報道

GitLab、AIエージェント向けの次世代Git互換ソースコード管理サービス「Project Switch」発表。最大で50倍高速かつ半分のトークンで利用可能に

Publickey ・ 2026-06-16

GitLab、AIエージェント向けGit互換管理サービス「Project Switch」発表

コミュニティ

Quoting Georgi Gerganov

Simon Willison's Weblog ・ 2026-06-16

Simon Willison、llama.cpp 開発者 Georgi Gerganov の発言を引用紹介

公式

How to Optimize Transformer-Based Models for Low-Precision Training

NVIDIA Developer Blog ・ 2026-06-16

NVIDIA、Transformer モデルの低精度学習を最適化する解説記事を公開

公式

Securing the future of AI agents

Google DeepMind Blog ・ 2026-06-16

Google DeepMind、AI エージェントを守る AI Control Roadmap を提示

公式

NVIDIA Blackwell Tops MLPerf Training 6.0 with Industry-Leading Scale and Performance

NVIDIA Developer Blog ・ 2026-06-16

NVIDIA、Blackwell が MLPerf Training 6.0 で首位と発表

報道

Stack Overflow、AIエージェント同士が掲示板で技術情報を共有する「Stack Overflow for Agents」ベータ公開

Publickey ・ 2026-06-15

Stack Overflow、AIエージェント向け情報共有サービスをベータ公開

コミュニティ

Building llm-driven “ai” still requires domain knowledge

Lobste.rs (AI tagged) ・ 2026-06-15

LLM駆動ツール開発でもドメイン知識の言語化が不可欠と論じる

報道

Sakana AI、初の商用プロダクト「Marlin」リリース その実力は?【出力レポート全文掲載】

ITmedia AI+ ・ 2026-06-15

Sakana AI、初の商用プロダクト「Sakana Marlin」を提供開始

コミュニティ

Why AI hasn’t replaced software engineers, and won’t

Simon Willison's Weblog ・ 2026-06-14

「AIはソフトウェア技術者を代替しない」と論じるエッセイ

学術(arxiv ほか) 130本 ▾
学術

Multi-LCB: Extending LiveCodeBench to Multiple Programming Languages

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-18

LiveCodeBを多言語に拡張したコード評価ベンチMulti-LCB

学術

Probe-and-Refine Tuning of Repository Guidance for Coding Agents

arXiv cs.LG (Machine Learning) ・ 2026-06-18

コーディングエージェント向けにリポジトリ指示文を調整する手法

学術

Entropy Estimation in Multi-Qutrit Systems via Variational and Classical Neural Networks

arXiv cs.LG (Machine Learning) ・ 2026-06-18

変分量子アルゴリズムとCNNでマルチqutritのエントロピー推定

学術

Scalable Training of Spatially Grounded 2D Vision-Language Models for Radiology

arXiv cs.CL (Computation and Language) ・ 2026-06-18

放射線科向け空間接地VLMの大規模学習とデータセットRefRad2D

学術

Judging to Improve: A De-biased VLM-as-3D-Judge Protocol for Single-Image 3D Generation

arXiv cs.LG (Machine Learning) ・ 2026-06-18

脱バイアスのVLM3D判定器で単一画像からの3D生成を専門化

学術

SoftSkill: Behavioral Compression for Contextual Adaptation

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-18

文脈適応のための行動圧縮手法SoftSkill

学術

Navigating Unreliable Parametric and Contextual Knowledge: Explicit Knowledge Conflict Resolution for LLM Inference

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-18

LLM推論で信頼できない知識の衝突を明示的に解消する手法

学術

SPOT-E: Test-Time Entropy Shaping with Visual Spotlights for Frozen VLMs

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-18

凍結VLM向けの視覚スポットライトによるテスト時エントロピー整形SPOT-E

学術

ScholarQuest: A Taxonomy-Guided Benchmark for Agentic Academic Paper Search in Open Literature Environments

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-18

開かれた文献環境でのエージェント論文探索ベンチScholarQuest

学術

MedRLM: Recursive Multimodal Health Intelligence for Long-Context Clinical Reasoning, Sensor-Guided Screening, Evidence-Grounded Decision Support, and Community-to-Tertiary Referral Optimization

arXiv cs.CL (Computation and Language) ・ 2026-06-18

長文脈の臨床推論に向けた再帰的マルチモーダル医療AI「MedRLM」

学術

When Does Streaming Tool Use Help? Characterizing Tool-Intent Stabilization in Streaming Retrieval-Augmented Generation

arXiv cs.CL (Computation and Language) ・ 2026-06-18

ストリーミングRAGでツール先行実行が効く条件を特徴づける

学術

IHUBERT: Vector-Based Semantic Deduplication and Domain-Balanced Pretraining for Persian Resources

arXiv cs.CL (Computation and Language) ・ 2026-06-18

ペルシャ語向け意味的重複除去とドメイン均衡事前学習のIHUBERT

学術

Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines

arXiv cs.CL (Computation and Language) ・ 2026-06-18

AI検索エンジン全体でブランド可視性を測る生成エンジン最適化

学術

Think Again or Think Longer? Selective Verification for Budget-Aware Reasoning

arXiv cs.CL (Computation and Language) ・ 2026-06-18

予算を意識した推論に向けた選択的な検証手法を提案

学術

CombEval: A Framework for Evaluating Combinatorial Counting in Large Language Models

arXiv cs.CL (Computation and Language) ・ 2026-06-18

LLMの組合せ計数能力を評価する枠組みCombEvalを提案

学術

AgentFinVQA: A Deployable Multi-Agent Pipeline for Auditable Financial Chart QA

arXiv cs.CL (Computation and Language) ・ 2026-06-18

監査可能な金融チャートQA向け多エージェントAgentFinVQA

学術

NRITYAM: Language Models Meet Art and Heritage of Dance

arXiv cs.CL (Computation and Language) ・ 2026-06-18

世界の舞踊文化でLMの文化理解を測るベンチマークNRITYAM

学術

Data Intelligence Agents: Interpreting, Modeling, and Querying Enterprise Data via Autonomous Coding Agents

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

自律コーディングエージェントで企業データを解釈・問い合わせ

学術

Explaining Attention with Program Synthesis

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

プログラム合成で注意機構を説明し解釈可能性を追求

学術

Enhancing Decision-Making with Large Language Models through Multi-Agent Fictitious Play

arXiv cs.CL (Computation and Language) ・ 2026-06-17

マルチエージェントの架空プレイで LLM の意思決定を強化

学術

Optimal scenario design for climate emulation

arXiv cs.LG (Machine Learning) ・ 2026-06-17

気候エミュレーションの精度を高める最適シナリオ設計を提案

学術

Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models

arXiv cs.LG (Machine Learning) ・ 2026-06-17

VLA モデルは常識を保持しているか、知識保持度を測る研究

学術

Trade-offs in Medical LLM Adaptation: An Empirical Study in French QA

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

医療 LLM のドメイン適応、仏語 QA で利点と代償を実証研究

学術

OneCanvas: 3D Scene Understanding via Panoramic Reprojection

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

OneCanvas、パノラマ再投影で VLM の 3D シーン理解を実現

学術

Transformer Geometry Observatory TGO-I: Spectral Geometry Observatory

arXiv cs.LG (Machine Learning) ・ 2026-06-17

TGO-I、スペクトル幾何で Vision Transformer の内部構造を解析

学術

TxBench-PP: Analyzing AI Agent Performance on Small-Molecule Preclinical Pharmacology

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

TxBench-PP、小分子前臨床薬理での AI エージェント性能を評価

学術

RECOM: A Validity Discrimination Tradeoff in Automatic Metrics for Open Ended Reddit Question Answering

arXiv cs.CL (Computation and Language) ・ 2026-06-17

RECOM、自動評価指標の妥当性と識別性のトレードオフを分析

学術

Learning to Annotate Delayed and False AEB Events: A Practical System for Extreme Class Imbalance and Asymmetric Label Noise

arXiv cs.LG (Machine Learning) ・ 2026-06-17

遅延・誤作動 AEB 事象を学習注釈、極端な不均衡に対応する実用系

学術

Hardware- and Vision-in-the-Loop Validation of Deep Monocular Pose Estimation for Autonomous Maritime UAV Flight

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

単眼姿勢推定を HIL/VIL 検証、艦上 UAV の自律飛行へ

学術

User as Engram: Internalizing Per-User Memory as Local Parametric Edits

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

ユーザー記憶を局所的なパラメータ編集として内在化する手法

学術

IndicContextEval: A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages

arXiv cs.CL (Computation and Language) ・ 2026-06-17

IndicContextEval、音声 LLM の文脈活用を 8 印度語で評価

学術

AdsMind: A Physics-Grounded Multi-Agent System for Self-Correcting Discovery of Adsorption Configurations on Heterogeneous Catalyst Surfaces

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

AdsMind、物理基盤のマルチエージェントで触媒の吸着配置を探索

学術

Complementary Attention Head Pruning for Efficient Transformers

arXiv cs.LG (Machine Learning) ・ 2026-06-17

相補的な注意ヘッド剪定で Transformer を効率化

学術

A Technical Taxonomy of LLM Agent Communication Protocols

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

LLM エージェントの通信プロトコルを技術的に分類・整理

学術

Seeing Before Reasoning: Decoupling Perception and Reasoning for Shortcut-Resilient Multimodal On-Policy Self-Distillation

arXiv cs.LG (Machine Learning) ・ 2026-06-17

知覚と推論を分離し、近道に強い多モーダル自己蒸留を実現

学術

Towards an Agent-First Web: Redesigning the Web for AI Agents

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

エージェント優先の Web へ、AI 向けに Web を再設計する提案

学術

RODS: Reward-Driven Online Data Synthesis for Multi-Turn Tool-Use Agents

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

RODS、報酬駆動のオンラインデータ合成で多ターンツール利用を強化

学術

Where Did the Variability Go? From Vibe Coding to Product Lines by Regeneration

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

Vibe コーディングの多様性はどこへ、再生成で製品ライン化

学術

A Hybrid LSTM--Vision Transformer Architecture for Predicting HRRR Forecast Errors

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

LSTM と ViT の融合で高解像度数値予報 HRRR の誤差を予測

学術

Spotlight: Synergizing Seed Exploration and Spot GPUs for DiT RL Post-Training

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

Spotlight、シード探索とスポット GPU で DiT の RL 事後学習を低コスト化

学術

TRAP: Benchmark for Task-completion and Resistance to Active Privacy-extraction

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

TRAP、課題遂行とプライバシー抽出耐性を測るエージェント評価

学術

Beyond Tokenization: Direct Timestep Embedding and Contrastive Alignment for Time-Series Question Answering

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

時系列を直接埋め込み時系列質問応答を高めるTSQA手法

学術

CAPRA: Scaling Feedback on Software Architecture Deliverables with a Multi-Agent LLM System

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

ソフト設計成果物の添削を自動化するマルチエージェントLLM「CAPRA」

学術

GraphPO: Graph-based Policy Optimization for Reasoning Models

arXiv cs.CL (Computation and Language) ・ 2026-06-17

推論モデル向けグラフベース方策最適化「GraphPO」

学術

RTSGameBench: An RTS Benchmark for Strategic Reasoning by Vision-Language Models

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

視覚言語モデルの戦略的推論を測るRTSベンチマーク

学術

Decoupling Search from Reasoning: A Vendor-Agnostic Grounding Architecture for LLM Agents

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

検索と推論を分離するベンダー非依存のLLMエージェント基盤

学術

SciRisk-Bench: A Risk-Dimension-Aware Benchmark for AI4Science Safety

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

AI for Scienceの安全性をリスク次元別に測る「SciRisk-Bench」

学術

REVES: REvision and VErification--Augmented Training for Test-Time Scaling

arXiv cs.CL (Computation and Language) ・ 2026-06-17

逐次修正によるテスト時スケーリングを強化する「REVES」

学術

Beyond Reward Engineering: A Data Recipe for Long-Context Reinforcement Learning

arXiv cs.CL (Computation and Language) ・ 2026-06-17

長文脈強化学習のためのデータレシピ

学術

GateMem: Benchmarking Memory Governance in Multi-Principal Shared-Memory Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-17

多主体共有メモリのガバナンスを測る「GateMem」

学術

LegalWorld: A Life-Cycle Interactive Environment for Legal Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-17

法務エージェント向けライフサイクル型環境「LegalWorld」

学術

LLMs Struggle to Measure What Distinguishes Students of Different Proficiency Levels: A Study of Item Discrimination in Reading Comprehension Assessment

arXiv cs.CL (Computation and Language) ・ 2026-06-17

LLMは読解問題の識別力指標の測定に苦戦

学術

Attention as Frustrated Synchronization

arXiv cs.CL (Computation and Language) ・ 2026-06-17

注意機構を「フラストレートした同期」として捉える理論

学術

ForecastBench-Sim: A Simulated-World Forecasting Benchmark

arXiv cs.CL (Computation and Language) ・ 2026-06-17

模擬世界で予測力を測る「ForecastBench-Sim」

学術

Variable-Width Transformers

arXiv cs.CL (Computation and Language) ・ 2026-06-16

幅可変Transformer、層ごとに幅を変え22%省FLOPs

学術

ReproRepo: Scaling Reproducibility Audits with GitHub Repository Issues

arXiv cs.CL (Computation and Language) ・ 2026-06-16

ReproRepo、GitHub課題で再現性監査をスケール

学術

EvolveNav: Proactive Preflection and Self-Evolving Memory for Zero-Shot Object Goal Navigation

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-16

軌跡記憶を自己進化させるゼロショット物体探索ナビゲーションを提案

学術

Adaptive Volumetric Mechanical Property Fields Invariant to Resolution

arXiv cs.LG (Machine Learning) ・ 2026-06-16

AdaVoMP、3D物体の力学特性場を解像度不変に予測

学術

Learning Red Agent Policy from Observations for Neurosymbolic Autonomous Cyber Agents

arXiv cs.LG (Machine Learning) ・ 2026-06-16

観測から赤エージェント方策を学ぶ自律サイバー防御

学術

Looped World Models

arXiv cs.CL (Computation and Language) ・ 2026-06-16

Looped World Models、反復的潜在精緻化で深さと効率を両立

学術

Fixed-Point Reasoners: Stable and Adaptive Deep Looped Transformers

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-16

ループ型Transformerの信号伝播問題を改善する手法FPRMを提案

学術

RubricsTree: Scalable and Evolving Open-Ended Evaluation of Personal Health Agents across Health Memory and Medical Skills

arXiv cs.CL (Computation and Language) ・ 2026-06-16

RubricsTree、個人健康エージェントの開放型評価を拡張

学術

Learning from the Self-future: On-policy Self-distillation for dLLMs

arXiv cs.CL (Computation and Language) ・ 2026-06-16

拡散LLM向けのオンポリシー自己蒸留OPSDを探究

学術

DRFLOW: A Deep Research Benchmark for Personalized Workflow Prediction

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-16

個別化ワークフロー予測を測るDeep Researchベンチマークを提案

学術

All Smoke, No Alarm: Oracle Signals in Agent-Authored Test Code

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-16

AIエージェント生成テストコードの検証力の弱さを分析した研究

学術

WEQA: Wearable hEalth Question Answering with Query-Adaptive Agentic Reasoning

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-16

装着型健康データのQAを行うエージェント推論手法WEQAを提案

学術

Memory as a Wasting Asset: Pricing Flash Endurance for Embodied Agents, and the Limits of Doing So

arXiv cs.LG (Machine Learning) ・ 2026-06-16

フラッシュ耐久を消耗資産として価格付けする実体エージェント論

学術

Your AI Travel Agent Would Book You a Bullfight: An Agentic Benchmark for Implicit Animal Welfare in Frontier AI Models

arXiv cs.CL (Computation and Language) ・ 2026-06-16

動物福祉の暗黙的配慮を測るエージェント型ベンチマーク

学術

Knowledge Reutilization in Meta-Reinforcement Learning

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-16

メタ強化学習で知識を再利用する転移フレームワークを提案

学術

Ternary Mamba: Grouped Quantization-Aware Training of W1.58A16 State Space Models

arXiv cs.LG (Machine Learning) ・ 2026-06-16

Ternary Mamba、1.58ビット重みのQATで状態空間を量子化

学術

HistoRAG: Embedding Historical Methodology in Retrieval-Augmented Generation Through Critical Technical Practice

arXiv cs.CL (Computation and Language) ・ 2026-06-16

HistoRAG、歴史方法論をRAGに組み込む批判的実践

学術

Agentic AI-based Framework for Mitigating Premature Diagnostic Handoff and Silent Hallucination in Healthcare Applications

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-16

医療AIの早期診断委譲と静かな幻覚を抑える多エージェント枠組み

学術

PseudoBench: Measuring How Agentic Auto-Research Fuels Pseudoscience

arXiv cs.CL (Computation and Language) ・ 2026-06-16

PseudoBench、自律研究エージェントが擬似科学を助長する度合いを測定

学術

Compositional Skill Routing for LLM Agents: Decompose, Retrieve, and Compose

arXiv cs.CL (Computation and Language) ・ 2026-06-16

LLMエージェント向け合成的スキルルーティング

学術

ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-16

ProvenanceGuard、MCPエージェント向け出所考慮の事実検証

学術

LoopCoder-v2: Only Loop Once for Efficient Test-Time Computation Scaling

arXiv cs.LG (Machine Learning) ・ 2026-06-16

LoopCoder-v2、一度のループで効率的テスト時計算スケール

学術

Recursive Scaling in Masked Diffusion Models

arXiv cs.LG (Machine Learning) ・ 2026-06-16

マスク拡散モデルにおける再帰的スケーリングを検討

学術

LLM Consumer Behavior Theory: Foundations of a Novel Research Field

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-16

エージェント市場の消費行動を扱う新研究領域LLM消費行動論を提唱

学術

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models

arXiv cs.CL (Computation and Language) ・ 2026-06-16

動的ロールアウト編集でRL推論モデルの過剰思考を抑制

学術

Monotonic Kolmogorov-Arnold Networks: A Theoretical and Empirical Study of Monotonicity as an Inductive Bias

arXiv cs.LG (Machine Learning) ・ 2026-06-16

単調KAN、帰納バイアスとしての単調性を理論・実験で検討

学術

GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine?

arXiv cs.CL (Computation and Language) ・ 2026-06-16

GameCraft-Bench、実ゲームエンジンで遊べるゲームを作れるか

学術

Environment-Grounded Automated Prompt Optimization for LLM Game Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-16

環境に接地した自動プロンプト最適化でLLMゲームエージェント

学術

From Drift to Coherence: Stabilizing Beliefs in LLMs

arXiv cs.LG (Machine Learning) ・ 2026-06-16

ドリフトから整合へ、LLMの信念を安定化

学術

A Framework for Evaluating Agentic Skills at Scale

arXiv cs.CL (Computation and Language) ・ 2026-06-16

エージェントのスキルを大規模に評価する枠組み

学術

Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering

arXiv cs.CL (Computation and Language) ・ 2026-06-16

立場論文、コーディングベンチはエージェント的開発と乖離

学術

Vision-language models for chest radiography do not always need the image

arXiv cs.CL (Computation and Language) ・ 2026-06-16

胸部X線の視覚言語モデルは画像を常に要しない

学術

EComAgentBench: Benchmarking Shopping Agents on Long-Horizon Tasks with Distributed Hidden Intent

arXiv cs.CL (Computation and Language) ・ 2026-06-16

EComAgentBench、隠れた意図を含む長期課題で買い物エージェント評価

学術

LLMs Infer Cultural Context but Fail to Apply It When Responding

arXiv cs.CL (Computation and Language) ・ 2026-06-16

LLMは文化的文脈を推測できても応答で適用できない

学術

SuCo: Sufficiency-guided Continuous Adaptive Reasoning

arXiv cs.CL (Computation and Language) ・ 2026-06-16

SuCo、十分性に導かれた連続適応的推論

学術

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning

arXiv cs.CL (Computation and Language) ・ 2026-06-16

EnvRL、環境ダイナミクスから学ぶエージェント強化学習

学術

MambaCount: Efficient Text-guided Open-vocabulary Object Counting with Spatial Sparse State Space Duality Block

arXiv cs.CL (Computation and Language) ・ 2026-06-16

MambaCount、状態空間双対ブロックで開語彙物体計数

学術

Beyond Domains: Reusing Web Skills via Transferable Interaction Patterns

arXiv cs.CL (Computation and Language) ・ 2026-06-16

領域を超えて転移可能な相互作用パターンでWebスキルを再利用

学術

OPD-Evolver: Cultivating Holistic Agent Evolver via On-Policy Distillation

arXiv cs.CL (Computation and Language) ・ 2026-06-16

OPD-Evolver、オンポリシー蒸留で自己進化エージェントを育成

学術

Context-Aware RL for Agentic and Multimodal LLMs

arXiv cs.CL (Computation and Language) ・ 2026-06-15

文脈選択を報酬化するRL手法ContextRLを提案

学術

Benchmarking LLM Agents on Meta-Analysis Articles from Nature Portfolio

arXiv cs.CL (Computation and Language) ・ 2026-06-15

Nature系メタ分析論文でLLMエージェントを評価するベンチマーク

学術

DEEPRUBRIC: Evidence-Tree Rubric Supervision for Efficient Reinforcement Learning of Deep Research Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-15

DeepRubric、評価基準を逆生成し深層リサーチエージェントのRLを効率化

学術

HAMON: Passive Optical Sequence Mixing for Long-Horizon Forecasting

arXiv cs.LG (Machine Learning) ・ 2026-06-15

HAMON、受動的な光学回路で長期時系列予測 ─ デジタル混合層が不要

学術

ExpRL: Exploratory RL for LLM Mid-Training

arXiv cs.LG (Machine Learning) ・ 2026-06-15

ExpRL、人手QAデータを「報酬足場」に使うLLM中間学習向けRLを提案

学術

TokenPilot: Cache-Efficient Context Management for LLM Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-15

TokenPilot、キャッシュを保つ文脈管理でLLMエージェントの推論コスト6割減

学術

Agent trajectories as programs: fingerprinting and programming coding-agent behavior

arXiv cs.LG (Machine Learning) ・ 2026-06-15

コーディングエージェントを手続き的に同定する「指紋」手法を提案

学術

Greed Is Learned: Visible Incentives as Reward-Hacking Triggers

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-15

報酬指標の可視化が RL 方策を「報酬チャネル依存」にし安全整合を崩すと報告

学術

LESS Is More: Mutual-Stability Sampling for Diffusion Language Models

arXiv cs.CL (Computation and Language) ・ 2026-06-15

拡散言語モデル向け学習不要の適応サンプラ『LESS』を提案

学術

Binary Tracking for Spatial QA and Navigation with Open Vision-Language Models

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-15

オープン VLM で動く空間質問応答・ナビ手法 Binary Tracking を提案

学術

Semantic Flip: Synthetic OOD Generation for Robust Refusal in Embodied Question Answering and Spatial Localization

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-15

身体化エージェントの拒否応答を強化する合成 OOD 生成手法 Semantic Flip を提案

学術

Compositional Reasoning Depth Predicts Clinical AI Failure: Empirical Evidence Consistent with Transformer Compositionality Limits in Electronic Health Record Question Answering

arXiv cs.CL (Computation and Language) ・ 2026-06-15

推論ホップ数が臨床 AI の誤りを予測、Transformer の合成性限界を示唆

学術

HawkesNest: A Multi-Axis Synthetic Benchmark for Spatiotemporal Pattern Complexity

arXiv cs.LG (Machine Learning) ・ 2026-06-15

HawkesNest: 時空間点過程モデル評価の合成ベンチマークを提案

学術

Robust Dual-Signal Fusion: Hybrid Neuro-Symbolic Gating with Compressed Chain-of-Thought Refinement for Irony Detection in Social Media Texts

arXiv cs.CL (Computation and Language) ・ 2026-06-15

圧縮 CoT の神経記号ハイブリッドでゼロショット皮肉検出を強化

学術

Data-Driven Decoding of Russell's Circumplex Model of Affect

arXiv cs.CL (Computation and Language) ・ 2026-06-15

Transformer 埋め込みが感情の Russell 円環モデル幾何を再現するか検証

学術

Beyond Models: Reflections on Engineering AI-enabled Systems in a Project-Based Course

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-15

AI搭載システムの工学教育を扱うプロジェクト型講義の実践を報告

学術

Does Traversal Order Matter? A Systematic Study of Tree Traversal Methods in Transformer Grammars

arXiv cs.CL (Computation and Language) ・ 2026-06-15

Transformer Grammars の木探索順序を比較分析する論文

学術

Tying the Loop -- Tied Expert Layers in Mixture-of-Experts Language Models

arXiv cs.CL (Computation and Language) ・ 2026-06-15

MoE で専門家パラメータを層間共有する手法を提案する論文

学術

How Much Can We Trust LLM Search Agents? Measuring Endorsement Vulnerability to Web Content Manipulation

arXiv cs.CL (Computation and Language) ・ 2026-06-15

LLM 検索エージェントの推薦汚染耐性を測る枠組みを提案

学術

GIST-CMTF: Goal-State Inference for Causal Minimal Tool Filtering in LLM Agents

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-15

ツール選択の誤目標実行を抑えるエージェント手法GIST-CMTFを提案

学術

Scaling LLM Reasoning from Minimal Labels: A Semi-Supervised Framework with a Lightweight Verifier

arXiv cs.CL (Computation and Language) ・ 2026-06-15

少量ラベルで LLM 推論を拡張する半教師あり枠組みを提案

学術

LabOSBench: Benchmarking Computer Use Agents for Scientific Instrument Control

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-15

科学機器を操作するエージェント評価へ、模擬ベンチLabOSBenchを提案

学術

OpenClaw-Skill: Collective Skill Tree Search for Agentic Large Language Models

arXiv cs.CL (Computation and Language) ・ 2026-06-15

エージェント向けに再利用可能スキルを木探索で構築するCSTSを提案

学術

Skill-to-LoRA: From Using Skills to Learning Behaviors for Token-Efficient LLM Agents

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-15

SKILL.md文書をLoRAに置換しトークン効率を高める手法S2Lを提案

学術

MyPCBench: A Benchmark for Personally Intelligent Computer-Use Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-15

個人秘書としてのPC操作エージェントを測る基準MyPCBenchを提案

学術

Misinformation Propagation in Benign Multi-Agent Systems

arXiv cs.CL (Computation and Language) ・ 2026-06-15

多エージェント系で誤情報が伝播し性能を低下させる現象を分析

学術

Multi-Turn Reflective Masking Elicits Reasoning in Mask Diffusion Models

arXiv cs.CL (Computation and Language) ・ 2026-06-15

マスク拡散モデルに反復的な局所修正の推論力を引き出す手法を提案

学術

Multimodal Evaluator Preference Collapse: Cross-Modal Contagion in Self-Evolving Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-15

自己進化エージェントの評価選好崩壊と跨モーダル伝播を扱う論文

学術

FraudSMSWalker: Benchmarking Agentic Large Language Models for SMS-to-Webpage Fraud Detection

arXiv cs.CL (Computation and Language) ・ 2026-06-15

SMS経由の詐欺判定を測るベンチマークFraudSMSWalkerを提案

学術

Islamic Large Language Models: From Knowledge Acquisition to Trustworthy and Hallucination-Resistant AI

arXiv cs.CL (Computation and Language) ・ 2026-06-15

イスラム知識を扱う信頼性の高いLLMの研究動向を概観する論文

学術

VeriGraph: Towards Verifiable Data-Analytic Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-15

データ分析エージェントの推論を検証可能にする VeriGraph を提案

学術

SING: Synthetic Intention Graph for Scalable Active Tool Discovery in LLM Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-15

LLM エージェントのツール探索を拡張する手法 SING を提案

学術

Uncertainty Is Not a Safety Net for Clinical VQA, but Can It Anticipate Model Failure?

arXiv cs.CL (Computation and Language) ・ 2026-06-15

臨床 VQA で不確実性推定は安全網にならないと検証

学術

Can LLM Agents Infer World Models? Evidence from Agentic Automata Learning

arXiv cs.CL (Computation and Language) ・ 2026-06-15

LLM エージェントは世界モデルを推論できるか、オートマトン学習で検証

学術

The BD-LSC Dataset: Facilitating the Benchmarking of Models for Lexical Semantic Change Detection in Slang and Standard Usage

arXiv cs.CL (Computation and Language) ・ 2026-06-15

語義変化検出の新ベンチマーク BD-LSC データセットを公開

学術

Can LLM Coding Agents Reason About Time Series?

arXiv cs.CL (Computation and Language) ・ 2026-06-15

LLM コーディングエージェントは時系列を推論できるか検証

学術

daVinci-kernel: Co-Evolving Skill Selection, Summarization, and Utilization via RL for GPU Kernel Optimization

arXiv cs.CL (Computation and Language) ・ 2026-06-15

GPUカーネル最適化向けスキル共進化RL「daVinci-kernel」を提案

← ストーリー アーカイブ