NVIDIA × Safety & Evaluation

NVIDIA unveils XR AI agent stack

NVIDIA unveils XR AI agent stack

✎ Story body

NVIDIA unveiled an agent stack for AR glasses and XR devices, moving in step with Hugging Face's robot-hardware integration and DeepMind's work on securing agents. All five sources are official platform vendors, with academia and community still quiet—less a week of research trickling down than one where infrastructure vendors began laying groundwork for XR and embodied agents. The focus is the implementation layer—how to run and secure agents on-device—rather than raw model capability. For now it's announcements and SDKs; real hardware adoption is the next thing to confirm.

▲ Official & Press
Official

Building AI Agents for AR Glasses and XR Devices with NVIDIA XR AI

NVIDIA Developer Blog ・ 2026-06-16 ・ 📌

NVIDIA unveils XR AI to build AI agents for AR glasses and XR devices

Community

Show HN: Are You in the Weights?

Hacker News (Front Page) ・ 2026-06-18

Show HN: 'Are You in the Weights?' checks if LLMs recognize you

Press

かんぽ生命、AIで営業支援 “郵便局での一言”拾って保険提案へ 寸劇で分かる活用例

ITmedia AI+ ・ 2026-06-17

Japan Post Insurance adds AI agents to its sales workflow

Official

From the Hugging Face Hub to robot hardware with Strands Agents and LeRobot

Hugging Face Blog ・ 2026-06-17

From Hugging Face Hub to robot hardware with Strands Agents and LeRobot

Press

「ポケカ対戦AIエージェント」開発コンテスト開始 「不完全情報ゲーム」をどう制するか

ITmedia AI+ ・ 2026-06-17

Contest launches to build AI agents for Pokemon TCG, an imperfect-info game

Official

Agentic Resource Discovery: Let agents search

Hugging Face Blog ・ 2026-06-17

Hugging Face proposes agentic resource discovery via search

Press

GitLab、AIエージェント向けの次世代Git互換ソースコード管理サービス「Project Switch」発表。最大で50倍高速かつ半分のトークンで利用可能に

Publickey ・ 2026-06-16

GitLab unveils 'Project Switch,' a Git-compatible SCM service for AI agents

Community

Quoting Georgi Gerganov

Simon Willison's Weblog ・ 2026-06-16

Simon Willison quotes Georgi Gerganov (llama.cpp / ggml author)

Official

How to Optimize Transformer-Based Models for Low-Precision Training

NVIDIA Developer Blog ・ 2026-06-16

NVIDIA guide on optimizing transformer models for low-precision training

Official

Securing the future of AI agents

Google DeepMind Blog ・ 2026-06-16

DeepMind outlines an AI Control Roadmap to secure AI agents

Official

NVIDIA Blackwell Tops MLPerf Training 6.0 with Industry-Leading Scale and Performance

NVIDIA Developer Blog ・ 2026-06-16

NVIDIA says Blackwell tops MLPerf Training 6.0 benchmark

Press

Stack Overflow、AIエージェント同士が掲示板で技術情報を共有する「Stack Overflow for Agents」ベータ公開

Publickey ・ 2026-06-15

Stack Overflow launches 'Stack Overflow for Agents' beta

Community

Building llm-driven “ai” still requires domain knowledge

Lobste.rs (AI tagged) ・ 2026-06-15

Building LLM-driven tools still hinges on capturing domain knowledge

Press

Sakana AI、初の商用プロダクト「Marlin」リリース その実力は?【出力レポート全文掲載】

ITmedia AI+ ・ 2026-06-15

Sakana AI launches its first commercial product, Sakana Marlin

Community

Why AI hasn’t replaced software engineers, and won’t

Simon Willison's Weblog ・ 2026-06-14

Essay argues AI hasn't replaced software engineers, and won't

Academic (arxiv etc.) 130 ▾
Academic

Multi-LCB: Extending LiveCodeBench to Multiple Programming Languages

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-18

Multi-LCB: extending LiveCodeBench to multiple programming languages

Academic

Probe-and-Refine Tuning of Repository Guidance for Coding Agents

arXiv cs.LG (Machine Learning) ・ 2026-06-18

Probe-and-Refine: tuning repository guidance for coding agents

Academic

Entropy Estimation in Multi-Qutrit Systems via Variational and Classical Neural Networks

arXiv cs.LG (Machine Learning) ・ 2026-06-18

Estimating entropy in multi-qutrit systems with VQAs and CNNs

Academic

Scalable Training of Spatially Grounded 2D Vision-Language Models for Radiology

arXiv cs.CL (Computation and Language) ・ 2026-06-18

RefRad2D: training spatially grounded radiology VLMs at scale

Academic

Judging to Improve: A De-biased VLM-as-3D-Judge Protocol for Single-Image 3D Generation

arXiv cs.LG (Machine Learning) ・ 2026-06-18

Using a de-biased VLM 3D judge to improve single-image 3D generation

Academic

SoftSkill: Behavioral Compression for Contextual Adaptation

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-18

SoftSkill: behavioral compression for contextual adaptation

Academic

Navigating Unreliable Parametric and Contextual Knowledge: Explicit Knowledge Conflict Resolution for LLM Inference

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-18

Explicit knowledge conflict resolution for LLM inference

Academic

SPOT-E: Test-Time Entropy Shaping with Visual Spotlights for Frozen VLMs

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-18

SPOT-E: test-time entropy shaping with visual spotlights for frozen VLMs

Academic

ScholarQuest: A Taxonomy-Guided Benchmark for Agentic Academic Paper Search in Open Literature Environments

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-18

ScholarQuest: a taxonomy-guided benchmark for agentic paper search

Academic

MedRLM: Recursive Multimodal Health Intelligence for Long-Context Clinical Reasoning, Sensor-Guided Screening, Evidence-Grounded Decision Support, and Community-to-Tertiary Referral Optimization

arXiv cs.CL (Computation and Language) ・ 2026-06-18

MedRLM: recursive multimodal AI for long-context clinical reasoning

Academic

When Does Streaming Tool Use Help? Characterizing Tool-Intent Stabilization in Streaming Retrieval-Augmented Generation

arXiv cs.CL (Computation and Language) ・ 2026-06-18

When does streaming tool use help in streaming RAG?

Academic

IHUBERT: Vector-Based Semantic Deduplication and Domain-Balanced Pretraining for Persian Resources

arXiv cs.CL (Computation and Language) ・ 2026-06-18

IHUBERT: a Persian language model with semantic dedup pretraining

Academic

Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines

arXiv cs.CL (Computation and Language) ・ 2026-06-18

Measuring brand visibility across AI search engines at scale

Academic

Think Again or Think Longer? Selective Verification for Budget-Aware Reasoning

arXiv cs.CL (Computation and Language) ・ 2026-06-18

Selective verification for budget-aware test-time reasoning

Academic

CombEval: A Framework for Evaluating Combinatorial Counting in Large Language Models

arXiv cs.CL (Computation and Language) ・ 2026-06-18

CombEval: evaluating combinatorial counting in LLMs

Academic

AgentFinVQA: A Deployable Multi-Agent Pipeline for Auditable Financial Chart QA

arXiv cs.CL (Computation and Language) ・ 2026-06-18

AgentFinVQA: an auditable multi-agent pipeline for financial chart QA

Academic

NRITYAM: Language Models Meet Art and Heritage of Dance

arXiv cs.CL (Computation and Language) ・ 2026-06-18

NRITYAM: a benchmark for cultural comprehension of dance traditions

Academic

Data Intelligence Agents: Interpreting, Modeling, and Querying Enterprise Data via Autonomous Coding Agents

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

Data Intelligence Agents query enterprise data autonomously

Academic

Explaining Attention with Program Synthesis

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

Explaining attention via program synthesis for interpretability

Academic

Enhancing Decision-Making with Large Language Models through Multi-Agent Fictitious Play

arXiv cs.CL (Computation and Language) ・ 2026-06-17

Multi-agent fictitious play boosts LLM decision-making

Academic

Optimal scenario design for climate emulation

arXiv cs.LG (Machine Learning) ・ 2026-06-17

Optimal scenario design improves climate emulation surrogates

Academic

Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models

arXiv cs.LG (Machine Learning) ・ 2026-06-17

Measuring commonsense and knowledge retention in VLA models

Academic

Trade-offs in Medical LLM Adaptation: An Empirical Study in French QA

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

Trade-offs in medical LLM adaptation, studied on French QA

Academic

OneCanvas: 3D Scene Understanding via Panoramic Reprojection

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

OneCanvas enables VLM 3D scene understanding via panoramic reprojection

Academic

Transformer Geometry Observatory TGO-I: Spectral Geometry Observatory

arXiv cs.LG (Machine Learning) ・ 2026-06-17

TGO-I: a spectral geometry observatory for Vision Transformers

Academic

TxBench-PP: Analyzing AI Agent Performance on Small-Molecule Preclinical Pharmacology

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

TxBench-PP evaluates AI agents on preclinical pharmacology

Academic

RECOM: A Validity Discrimination Tradeoff in Automatic Metrics for Open Ended Reddit Question Answering

arXiv cs.CL (Computation and Language) ・ 2026-06-17

RECOM analyzes validity vs discrimination in automatic metrics

Academic

Learning to Annotate Delayed and False AEB Events: A Practical System for Extreme Class Imbalance and Asymmetric Label Noise

arXiv cs.LG (Machine Learning) ・ 2026-06-17

Annotating rare delayed and false AEB events under class imbalance

Academic

Hardware- and Vision-in-the-Loop Validation of Deep Monocular Pose Estimation for Autonomous Maritime UAV Flight

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

Hardware/vision-in-the-loop validation of monocular UAV pose estimation

Academic

User as Engram: Internalizing Per-User Memory as Local Parametric Edits

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

User as Engram: per-user memory as local parametric edits

Academic

IndicContextEval: A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages

arXiv cs.CL (Computation and Language) ・ 2026-06-17

IndicContextEval: audio-LLM context use across 8 Indic languages

Academic

AdsMind: A Physics-Grounded Multi-Agent System for Self-Correcting Discovery of Adsorption Configurations on Heterogeneous Catalyst Surfaces

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

AdsMind: physics-grounded multi-agent search for adsorption configs

Academic

Complementary Attention Head Pruning for Efficient Transformers

arXiv cs.LG (Machine Learning) ・ 2026-06-17

Complementary attention-head pruning for efficient Transformers

Academic

A Technical Taxonomy of LLM Agent Communication Protocols

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

A technical taxonomy of LLM agent communication protocols

Academic

Seeing Before Reasoning: Decoupling Perception and Reasoning for Shortcut-Resilient Multimodal On-Policy Self-Distillation

arXiv cs.LG (Machine Learning) ・ 2026-06-17

Decoupling perception and reasoning for shortcut-resilient self-distillation

Academic

Towards an Agent-First Web: Redesigning the Web for AI Agents

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

Towards an agent-first web: redesigning the web for AI agents

Academic

RODS: Reward-Driven Online Data Synthesis for Multi-Turn Tool-Use Agents

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

RODS: reward-driven online data synthesis for tool-use agents

Academic

Where Did the Variability Go? From Vibe Coding to Product Lines by Regeneration

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

From vibe coding to product lines via regeneration

Academic

A Hybrid LSTM--Vision Transformer Architecture for Predicting HRRR Forecast Errors

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

Hybrid LSTM–Vision Transformer predicts HRRR forecast errors

Academic

Spotlight: Synergizing Seed Exploration and Spot GPUs for DiT RL Post-Training

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

Spotlight cuts DiT RL post-training cost with spot GPUs

Academic

TRAP: Benchmark for Task-completion and Resistance to Active Privacy-extraction

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

TRAP benchmarks agents on task completion and privacy resistance

Academic

Beyond Tokenization: Direct Timestep Embedding and Contrastive Alignment for Time-Series Question Answering

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

Direct timestep embedding and contrastive alignment for time-series QA

Academic

CAPRA: Scaling Feedback on Software Architecture Deliverables with a Multi-Agent LLM System

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

CAPRA: a multi-agent LLM system for software architecture feedback

Academic

GraphPO: Graph-based Policy Optimization for Reasoning Models

arXiv cs.CL (Computation and Language) ・ 2026-06-17

GraphPO: graph-based policy optimization for reasoning models

Academic

RTSGameBench: An RTS Benchmark for Strategic Reasoning by Vision-Language Models

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

RTSGameBench: an RTS benchmark for strategic reasoning by VLMs

Academic

Decoupling Search from Reasoning: A Vendor-Agnostic Grounding Architecture for LLM Agents

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

Decoupling search from reasoning: a vendor-agnostic grounding architecture

Academic

SciRisk-Bench: A Risk-Dimension-Aware Benchmark for AI4Science Safety

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

SciRisk-Bench: a risk-dimension-aware benchmark for AI4Science safety

Academic

REVES: REvision and VErification--Augmented Training for Test-Time Scaling

arXiv cs.CL (Computation and Language) ・ 2026-06-17

REVES: revision- and verification-augmented training for test-time scaling

Academic

Beyond Reward Engineering: A Data Recipe for Long-Context Reinforcement Learning

arXiv cs.CL (Computation and Language) ・ 2026-06-17

Beyond reward engineering: a data recipe for long-context RL

Academic

GateMem: Benchmarking Memory Governance in Multi-Principal Shared-Memory Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-17

GateMem: benchmarking memory governance in shared-memory agents

Academic

LegalWorld: A Life-Cycle Interactive Environment for Legal Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-17

LegalWorld: a life-cycle interactive environment for legal agents

Academic

LLMs Struggle to Measure What Distinguishes Students of Different Proficiency Levels: A Study of Item Discrimination in Reading Comprehension Assessment

arXiv cs.CL (Computation and Language) ・ 2026-06-17

LLMs struggle to measure item discrimination in reading assessment

Academic

Attention as Frustrated Synchronization

arXiv cs.CL (Computation and Language) ・ 2026-06-17

Attention as frustrated synchronization

Academic

ForecastBench-Sim: A Simulated-World Forecasting Benchmark

arXiv cs.CL (Computation and Language) ・ 2026-06-17

ForecastBench-Sim: a simulated-world forecasting benchmark

Academic

Variable-Width Transformers

arXiv cs.CL (Computation and Language) ・ 2026-06-16

Variable-width transformer cuts FLOPs ~22% via x-shaped layer widths

Academic

ReproRepo: Scaling Reproducibility Audits with GitHub Repository Issues

arXiv cs.CL (Computation and Language) ・ 2026-06-16

ReproRepo scales reproducibility audits using GitHub repo issues

Academic

EvolveNav: Proactive Preflection and Self-Evolving Memory for Zero-Shot Object Goal Navigation

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-16

EvolveNav: a self-evolving framework for zero-shot object-goal navigation

Academic

Adaptive Volumetric Mechanical Property Fields Invariant to Resolution

arXiv cs.LG (Machine Learning) ・ 2026-06-16

AdaVoMP predicts resolution-invariant mechanical property fields for 3D

Academic

Learning Red Agent Policy from Observations for Neurosymbolic Autonomous Cyber Agents

arXiv cs.LG (Machine Learning) ・ 2026-06-16

Learning red-agent policy from observations for cyber-defense RL

Academic

Looped World Models

arXiv cs.CL (Computation and Language) ・ 2026-06-16

Looped World Models refine latents iteratively for efficient sim

Academic

Fixed-Point Reasoners: Stable and Adaptive Deep Looped Transformers

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-16

Fixed-Point Reasoners: stabilizing deep looped Transformers (FPRM)

Academic

RubricsTree: Scalable and Evolving Open-Ended Evaluation of Personal Health Agents across Health Memory and Medical Skills

arXiv cs.CL (Computation and Language) ・ 2026-06-16

RubricsTree: scalable open-ended evaluation of personal health agents

Academic

Learning from the Self-future: On-policy Self-distillation for dLLMs

arXiv cs.CL (Computation and Language) ・ 2026-06-16

On-policy self-distillation explored for diffusion LLMs

Academic

DRFLOW: A Deep Research Benchmark for Personalized Workflow Prediction

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-16

DRFLOW: a deep research benchmark for personalized workflow prediction

Academic

All Smoke, No Alarm: Oracle Signals in Agent-Authored Test Code

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-16

Study finds agent-authored test code often lacks real verification logic

Academic

WEQA: Wearable hEalth Question Answering with Query-Adaptive Agentic Reasoning

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-16

WEQA: query-adaptive agentic reasoning for wearable health QA

Academic

Memory as a Wasting Asset: Pricing Flash Endurance for Embodied Agents, and the Limits of Doing So

arXiv cs.LG (Machine Learning) ・ 2026-06-16

Pricing flash endurance as a wasting asset for embodied agents

Academic

Your AI Travel Agent Would Book You a Bullfight: An Agentic Benchmark for Implicit Animal Welfare in Frontier AI Models

arXiv cs.CL (Computation and Language) ・ 2026-06-16

An agentic benchmark for implicit animal welfare in frontier AI

Academic

Knowledge Reutilization in Meta-Reinforcement Learning

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-16

A meta-knowledge reutilization framework for meta-RL across agents

Academic

Ternary Mamba: Grouped Quantization-Aware Training of W1.58A16 State Space Models

arXiv cs.LG (Machine Learning) ・ 2026-06-16

Ternary Mamba: grouped QAT for W1.58A16 state space models

Academic

HistoRAG: Embedding Historical Methodology in Retrieval-Augmented Generation Through Critical Technical Practice

arXiv cs.CL (Computation and Language) ・ 2026-06-16

HistoRAG embeds historical methodology into RAG via critical practice

Academic

Agentic AI-based Framework for Mitigating Premature Diagnostic Handoff and Silent Hallucination in Healthcare Applications

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-16

A multi-agent framework against premature handoff and silent hallucination

Academic

PseudoBench: Measuring How Agentic Auto-Research Fuels Pseudoscience

arXiv cs.CL (Computation and Language) ・ 2026-06-16

PseudoBench measures how agentic auto-research fuels pseudoscience

Academic

Compositional Skill Routing for LLM Agents: Decompose, Retrieve, and Compose

arXiv cs.CL (Computation and Language) ・ 2026-06-16

Compositional skill routing for LLM agents: decompose, retrieve, compose

Academic

ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-16

ProvenanceGuard: source-aware factuality verification for MCP agents

Academic

LoopCoder-v2: Only Loop Once for Efficient Test-Time Computation Scaling

arXiv cs.LG (Machine Learning) ・ 2026-06-16

LoopCoder-v2: loop once for efficient test-time compute scaling

Academic

Recursive Scaling in Masked Diffusion Models

arXiv cs.LG (Machine Learning) ・ 2026-06-16

Recursive scaling in masked diffusion models

Academic

LLM Consumer Behavior Theory: Foundations of a Novel Research Field

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-16

LLM Consumer Behavior Theory: a new field for agentic markets

Academic

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models

arXiv cs.CL (Computation and Language) ・ 2026-06-16

Dynamic rollout editing reduces overthinking in RL reasoning models

Academic

Monotonic Kolmogorov-Arnold Networks: A Theoretical and Empirical Study of Monotonicity as an Inductive Bias

arXiv cs.LG (Machine Learning) ・ 2026-06-16

Monotonic KANs: monotonicity as an inductive bias, studied theoretically

Academic

GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine?

arXiv cs.CL (Computation and Language) ・ 2026-06-16

GameCraft-Bench: can agents build playable games end-to-end?

Academic

Environment-Grounded Automated Prompt Optimization for LLM Game Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-16

Environment-grounded automated prompt optimization for LLM game agents

Academic

From Drift to Coherence: Stabilizing Beliefs in LLMs

arXiv cs.LG (Machine Learning) ・ 2026-06-16

From drift to coherence: stabilizing beliefs in LLMs

Academic

A Framework for Evaluating Agentic Skills at Scale

arXiv cs.CL (Computation and Language) ・ 2026-06-16

A framework for evaluating agentic skills at scale

Academic

Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering

arXiv cs.CL (Computation and Language) ・ 2026-06-16

Position: coding benchmarks are misaligned with agentic software engineering

Academic

Vision-language models for chest radiography do not always need the image

arXiv cs.CL (Computation and Language) ・ 2026-06-16

Vision-language models for chest radiography do not always need the image

Academic

EComAgentBench: Benchmarking Shopping Agents on Long-Horizon Tasks with Distributed Hidden Intent

arXiv cs.CL (Computation and Language) ・ 2026-06-16

EComAgentBench: shopping agents on long-horizon tasks with hidden intent

Academic

LLMs Infer Cultural Context but Fail to Apply It When Responding

arXiv cs.CL (Computation and Language) ・ 2026-06-16

LLMs infer cultural context but fail to apply it when responding

Academic

SuCo: Sufficiency-guided Continuous Adaptive Reasoning

arXiv cs.CL (Computation and Language) ・ 2026-06-16

SuCo: sufficiency-guided continuous adaptive reasoning

Academic

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning

arXiv cs.CL (Computation and Language) ・ 2026-06-16

EnvRL learns from environment dynamics in agentic RL

Academic

MambaCount: Efficient Text-guided Open-vocabulary Object Counting with Spatial Sparse State Space Duality Block

arXiv cs.CL (Computation and Language) ・ 2026-06-16

MambaCount: efficient open-vocabulary counting via state-space duality

Academic

Beyond Domains: Reusing Web Skills via Transferable Interaction Patterns

arXiv cs.CL (Computation and Language) ・ 2026-06-16

Reusing web skills via transferable interaction patterns

Academic

OPD-Evolver: Cultivating Holistic Agent Evolver via On-Policy Distillation

arXiv cs.CL (Computation and Language) ・ 2026-06-16

OPD-Evolver cultivates self-evolving agents via on-policy distillation

Academic

Context-Aware RL for Agentic and Multimodal LLMs

arXiv cs.CL (Computation and Language) ・ 2026-06-15

ContextRL rewards picking the right context to ground answers

Academic

Benchmarking LLM Agents on Meta-Analysis Articles from Nature Portfolio

arXiv cs.CL (Computation and Language) ・ 2026-06-15

A benchmark for LLM agents on Nature Portfolio meta-analyses

Academic

DEEPRUBRIC: Evidence-Tree Rubric Supervision for Efficient Reinforcement Learning of Deep Research Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-15

DeepRubric: evidence-tree rubrics to boost deep-research agent RL

Academic

HAMON: Passive Optical Sequence Mixing for Long-Horizon Forecasting

arXiv cs.LG (Machine Learning) ・ 2026-06-15

HAMON: a passive optical core for long-horizon forecasting

Academic

ExpRL: Exploratory RL for LLM Mid-Training

arXiv cs.LG (Machine Learning) ・ 2026-06-15

ExpRL uses human QA as reward scaffolds for LLM mid-training RL

Academic

TokenPilot: Cache-Efficient Context Management for LLM Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-15

TokenPilot cuts LLM-agent context costs ~61% while preserving prompt cache

Academic

Agent trajectories as programs: fingerprinting and programming coding-agent behavior

arXiv cs.LG (Machine Learning) ・ 2026-06-15

Coding agents have behavioral fingerprints identifiable from trajectories

Academic

Greed Is Learned: Visible Incentives as Reward-Hacking Triggers

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-15

Greed Is Learned: RL agents get addicted to visible reward channels

Academic

LESS Is More: Mutual-Stability Sampling for Diffusion Language Models

arXiv cs.CL (Computation and Language) ・ 2026-06-15

LESS: a training-free adaptive sampler for diffusion language models

Academic

Binary Tracking for Spatial QA and Navigation with Open Vision-Language Models

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-15

Binary Tracking: open vision-language models for spatial QA and navigation

Academic

Semantic Flip: Synthetic OOD Generation for Robust Refusal in Embodied Question Answering and Spatial Localization

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-15

Semantic Flip: synthetic OOD generation for robust refusal in embodied agents

Academic

Compositional Reasoning Depth Predicts Clinical AI Failure: Empirical Evidence Consistent with Transformer Compositionality Limits in Electronic Health Record Question Answering

arXiv cs.CL (Computation and Language) ・ 2026-06-15

Reasoning hop-count predicts clinical AI failure in EHR QA

Academic

HawkesNest: A Multi-Axis Synthetic Benchmark for Spatiotemporal Pattern Complexity

arXiv cs.LG (Machine Learning) ・ 2026-06-15

HawkesNest: a synthetic benchmark for spatiotemporal point process models

Academic

Robust Dual-Signal Fusion: Hybrid Neuro-Symbolic Gating with Compressed Chain-of-Thought Refinement for Irony Detection in Social Media Texts

arXiv cs.CL (Computation and Language) ・ 2026-06-15

RDS Fusion: neuro-symbolic gating with compressed CoT for irony detection

Academic

Data-Driven Decoding of Russell's Circumplex Model of Affect

arXiv cs.CL (Computation and Language) ・ 2026-06-15

Do Transformer embeddings recover Russell's circumplex affect geometry?

Academic

Beyond Models: Reflections on Engineering AI-enabled Systems in a Project-Based Course

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-15

Reflections on teaching the engineering of AI-enabled systems in a course

Academic

Does Traversal Order Matter? A Systematic Study of Tree Traversal Methods in Transformer Grammars

arXiv cs.CL (Computation and Language) ・ 2026-06-15

Paper: compares tree traversal orders in Transformer Grammars

Academic

Tying the Loop -- Tied Expert Layers in Mixture-of-Experts Language Models

arXiv cs.CL (Computation and Language) ・ 2026-06-15

Paper: Expert Tying shares MoE expert params across layers

Academic

How Much Can We Trust LLM Search Agents? Measuring Endorsement Vulnerability to Web Content Manipulation

arXiv cs.CL (Computation and Language) ・ 2026-06-15

Paper: framework measures LLM search-agent endorsement risk

Academic

GIST-CMTF: Goal-State Inference for Causal Minimal Tool Filtering in LLM Agents

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-15

GIST-CMTF adds goal-state inference to causal minimal tool filtering

Academic

Scaling LLM Reasoning from Minimal Labels: A Semi-Supervised Framework with a Lightweight Verifier

arXiv cs.CL (Computation and Language) ・ 2026-06-15

Paper: semi-supervised LLM reasoning from minimal labels

Academic

LabOSBench: Benchmarking Computer Use Agents for Scientific Instrument Control

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-15

LabOSBench: a simulated testbed for computer-use agents controlling instruments

Academic

OpenClaw-Skill: Collective Skill Tree Search for Agentic Large Language Models

arXiv cs.CL (Computation and Language) ・ 2026-06-15

OpenClaw-Skill: collective skill tree search for LLM agents

Academic

Skill-to-LoRA: From Using Skills to Learning Behaviors for Token-Efficient LLM Agents

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-15

S2L replaces runtime SKILL.md text with skill-specific LoRA adapters

Academic

MyPCBench: A Benchmark for Personally Intelligent Computer-Use Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-15

MyPCBench: benchmarking personal computer-use agents

Academic

Misinformation Propagation in Benign Multi-Agent Systems

arXiv cs.CL (Computation and Language) ・ 2026-06-15

Study on misinformation propagation in benign multi-agent systems

Academic

Multi-Turn Reflective Masking Elicits Reasoning in Mask Diffusion Models

arXiv cs.CL (Computation and Language) ・ 2026-06-15

Reflective Masking elicits iterative reasoning in mask diffusion models

Academic

Multimodal Evaluator Preference Collapse: Cross-Modal Contagion in Self-Evolving Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-15

Paper on evaluator preference collapse in self-evolving agents

Academic

FraudSMSWalker: Benchmarking Agentic Large Language Models for SMS-to-Webpage Fraud Detection

arXiv cs.CL (Computation and Language) ・ 2026-06-15

FraudSMSWalker benchmark targets URL-masked SMS-to-webpage fraud

Academic

Islamic Large Language Models: From Knowledge Acquisition to Trustworthy and Hallucination-Resistant AI

arXiv cs.CL (Computation and Language) ・ 2026-06-15

Survey reviews Islamic LLMs and trustworthy, hallucination-resistant AI

Academic

VeriGraph: Towards Verifiable Data-Analytic Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-15

VeriGraph: a traceable neuro-symbolic framework for verifiable data agents

Academic

SING: Synthetic Intention Graph for Scalable Active Tool Discovery in LLM Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-15

SING: synthetic intention graph for scalable active tool discovery

Academic

Uncertainty Is Not a Safety Net for Clinical VQA, but Can It Anticipate Model Failure?

arXiv cs.CL (Computation and Language) ・ 2026-06-15

Uncertainty estimation fails as a safety net for clinical VQA

Academic

Can LLM Agents Infer World Models? Evidence from Agentic Automata Learning

arXiv cs.CL (Computation and Language) ・ 2026-06-15

Can LLM agents infer world models? Evidence from automata learning

Academic

The BD-LSC Dataset: Facilitating the Benchmarking of Models for Lexical Semantic Change Detection in Slang and Standard Usage

arXiv cs.CL (Computation and Language) ・ 2026-06-15

BD-LSC: a new benchmark dataset for lexical semantic change detection

Academic

Can LLM Coding Agents Reason About Time Series?

arXiv cs.CL (Computation and Language) ・ 2026-06-15

Can LLM coding agents reason about time series? A benchmark study

Academic

daVinci-kernel: Co-Evolving Skill Selection, Summarization, and Utilization via RL for GPU Kernel Optimization

arXiv cs.CL (Computation and Language) ・ 2026-06-15

daVinci-kernel: an RL framework co-evolving skills for GPU kernel tuning

← Story Archive