NVIDIA unveiled an agent stack for AR glasses and XR devices, moving in step with Hugging Face's robot-hardware integration and DeepMind's work on securing agents. All five sources are official platform vendors, with academia and community still quiet—less a week of research trickling down than one where infrastructure vendors began laying groundwork for XR and embodied agents. The focus is the implementation layer—how to run and secure agents on-device—rather than raw model capability. For now it's announcements and SDKs; real hardware adoption is the next thing to confirm.
NVIDIA unveils XR AI agent stack
NVIDIA unveils XR AI agent stack
Building AI Agents for AR Glasses and XR Devices with NVIDIA XR AI
NVIDIA unveils XR AI to build AI agents for AR glasses and XR devices
Show HN: Are You in the Weights?
Show HN: 'Are You in the Weights?' checks if LLMs recognize you
かんぽ生命、AIで営業支援 “郵便局での一言”拾って保険提案へ 寸劇で分かる活用例
Japan Post Insurance adds AI agents to its sales workflow
From the Hugging Face Hub to robot hardware with Strands Agents and LeRobot
From Hugging Face Hub to robot hardware with Strands Agents and LeRobot
「ポケカ対戦AIエージェント」開発コンテスト開始 「不完全情報ゲーム」をどう制するか
Contest launches to build AI agents for Pokemon TCG, an imperfect-info game
Agentic Resource Discovery: Let agents search
Hugging Face proposes agentic resource discovery via search
GitLab、AIエージェント向けの次世代Git互換ソースコード管理サービス「Project Switch」発表。最大で50倍高速かつ半分のトークンで利用可能に
GitLab unveils 'Project Switch,' a Git-compatible SCM service for AI agents
Simon Willison quotes Georgi Gerganov (llama.cpp / ggml author)
How to Optimize Transformer-Based Models for Low-Precision Training
NVIDIA guide on optimizing transformer models for low-precision training
Securing the future of AI agents
DeepMind outlines an AI Control Roadmap to secure AI agents
NVIDIA Blackwell Tops MLPerf Training 6.0 with Industry-Leading Scale and Performance
NVIDIA says Blackwell tops MLPerf Training 6.0 benchmark
Stack Overflow、AIエージェント同士が掲示板で技術情報を共有する「Stack Overflow for Agents」ベータ公開
Stack Overflow launches 'Stack Overflow for Agents' beta
Building llm-driven “ai” still requires domain knowledge
Building LLM-driven tools still hinges on capturing domain knowledge
Sakana AI、初の商用プロダクト「Marlin」リリース その実力は?【出力レポート全文掲載】
Sakana AI launches its first commercial product, Sakana Marlin
Why AI hasn’t replaced software engineers, and won’t
Essay argues AI hasn't replaced software engineers, and won't
Academic (arxiv etc.) 130 ▾
Multi-LCB: Extending LiveCodeBench to Multiple Programming Languages
Multi-LCB: extending LiveCodeBench to multiple programming languages
Probe-and-Refine Tuning of Repository Guidance for Coding Agents
Probe-and-Refine: tuning repository guidance for coding agents
Entropy Estimation in Multi-Qutrit Systems via Variational and Classical Neural Networks
Estimating entropy in multi-qutrit systems with VQAs and CNNs
Scalable Training of Spatially Grounded 2D Vision-Language Models for Radiology
RefRad2D: training spatially grounded radiology VLMs at scale
Judging to Improve: A De-biased VLM-as-3D-Judge Protocol for Single-Image 3D Generation
Using a de-biased VLM 3D judge to improve single-image 3D generation
SoftSkill: Behavioral Compression for Contextual Adaptation
SoftSkill: behavioral compression for contextual adaptation
Explicit knowledge conflict resolution for LLM inference
SPOT-E: Test-Time Entropy Shaping with Visual Spotlights for Frozen VLMs
SPOT-E: test-time entropy shaping with visual spotlights for frozen VLMs
ScholarQuest: a taxonomy-guided benchmark for agentic paper search
MedRLM: recursive multimodal AI for long-context clinical reasoning
When does streaming tool use help in streaming RAG?
IHUBERT: Vector-Based Semantic Deduplication and Domain-Balanced Pretraining for Persian Resources
IHUBERT: a Persian language model with semantic dedup pretraining
Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines
Measuring brand visibility across AI search engines at scale
Think Again or Think Longer? Selective Verification for Budget-Aware Reasoning
Selective verification for budget-aware test-time reasoning
CombEval: A Framework for Evaluating Combinatorial Counting in Large Language Models
CombEval: evaluating combinatorial counting in LLMs
AgentFinVQA: A Deployable Multi-Agent Pipeline for Auditable Financial Chart QA
AgentFinVQA: an auditable multi-agent pipeline for financial chart QA
NRITYAM: Language Models Meet Art and Heritage of Dance
NRITYAM: a benchmark for cultural comprehension of dance traditions
Data Intelligence Agents query enterprise data autonomously
Explaining Attention with Program Synthesis
Explaining attention via program synthesis for interpretability
Enhancing Decision-Making with Large Language Models through Multi-Agent Fictitious Play
Multi-agent fictitious play boosts LLM decision-making
Optimal scenario design for climate emulation
Optimal scenario design improves climate emulation surrogates
Measuring commonsense and knowledge retention in VLA models
Trade-offs in Medical LLM Adaptation: An Empirical Study in French QA
Trade-offs in medical LLM adaptation, studied on French QA
OneCanvas: 3D Scene Understanding via Panoramic Reprojection
OneCanvas enables VLM 3D scene understanding via panoramic reprojection
Transformer Geometry Observatory TGO-I: Spectral Geometry Observatory
TGO-I: a spectral geometry observatory for Vision Transformers
TxBench-PP: Analyzing AI Agent Performance on Small-Molecule Preclinical Pharmacology
TxBench-PP evaluates AI agents on preclinical pharmacology
RECOM analyzes validity vs discrimination in automatic metrics
Annotating rare delayed and false AEB events under class imbalance
Hardware/vision-in-the-loop validation of monocular UAV pose estimation
User as Engram: Internalizing Per-User Memory as Local Parametric Edits
User as Engram: per-user memory as local parametric edits
IndicContextEval: audio-LLM context use across 8 Indic languages
AdsMind: physics-grounded multi-agent search for adsorption configs
Complementary Attention Head Pruning for Efficient Transformers
Complementary attention-head pruning for efficient Transformers
A Technical Taxonomy of LLM Agent Communication Protocols
A technical taxonomy of LLM agent communication protocols
Decoupling perception and reasoning for shortcut-resilient self-distillation
Towards an Agent-First Web: Redesigning the Web for AI Agents
Towards an agent-first web: redesigning the web for AI agents
RODS: Reward-Driven Online Data Synthesis for Multi-Turn Tool-Use Agents
RODS: reward-driven online data synthesis for tool-use agents
Where Did the Variability Go? From Vibe Coding to Product Lines by Regeneration
From vibe coding to product lines via regeneration
A Hybrid LSTM--Vision Transformer Architecture for Predicting HRRR Forecast Errors
Hybrid LSTM–Vision Transformer predicts HRRR forecast errors
Spotlight: Synergizing Seed Exploration and Spot GPUs for DiT RL Post-Training
Spotlight cuts DiT RL post-training cost with spot GPUs
TRAP: Benchmark for Task-completion and Resistance to Active Privacy-extraction
TRAP benchmarks agents on task completion and privacy resistance
Direct timestep embedding and contrastive alignment for time-series QA
CAPRA: Scaling Feedback on Software Architecture Deliverables with a Multi-Agent LLM System
CAPRA: a multi-agent LLM system for software architecture feedback
GraphPO: Graph-based Policy Optimization for Reasoning Models
GraphPO: graph-based policy optimization for reasoning models
RTSGameBench: An RTS Benchmark for Strategic Reasoning by Vision-Language Models
RTSGameBench: an RTS benchmark for strategic reasoning by VLMs
Decoupling Search from Reasoning: A Vendor-Agnostic Grounding Architecture for LLM Agents
Decoupling search from reasoning: a vendor-agnostic grounding architecture
SciRisk-Bench: A Risk-Dimension-Aware Benchmark for AI4Science Safety
SciRisk-Bench: a risk-dimension-aware benchmark for AI4Science safety
REVES: REvision and VErification--Augmented Training for Test-Time Scaling
REVES: revision- and verification-augmented training for test-time scaling
Beyond Reward Engineering: A Data Recipe for Long-Context Reinforcement Learning
Beyond reward engineering: a data recipe for long-context RL
GateMem: Benchmarking Memory Governance in Multi-Principal Shared-Memory Agents
GateMem: benchmarking memory governance in shared-memory agents
LegalWorld: A Life-Cycle Interactive Environment for Legal Agents
LegalWorld: a life-cycle interactive environment for legal agents
LLMs struggle to measure item discrimination in reading assessment
Attention as Frustrated Synchronization
Attention as frustrated synchronization
ForecastBench-Sim: A Simulated-World Forecasting Benchmark
ForecastBench-Sim: a simulated-world forecasting benchmark
Variable-width transformer cuts FLOPs ~22% via x-shaped layer widths
ReproRepo: Scaling Reproducibility Audits with GitHub Repository Issues
ReproRepo scales reproducibility audits using GitHub repo issues
EvolveNav: Proactive Preflection and Self-Evolving Memory for Zero-Shot Object Goal Navigation
EvolveNav: a self-evolving framework for zero-shot object-goal navigation
Adaptive Volumetric Mechanical Property Fields Invariant to Resolution
AdaVoMP predicts resolution-invariant mechanical property fields for 3D
Learning Red Agent Policy from Observations for Neurosymbolic Autonomous Cyber Agents
Learning red-agent policy from observations for cyber-defense RL
Looped World Models refine latents iteratively for efficient sim
Fixed-Point Reasoners: Stable and Adaptive Deep Looped Transformers
Fixed-Point Reasoners: stabilizing deep looped Transformers (FPRM)
RubricsTree: scalable open-ended evaluation of personal health agents
Learning from the Self-future: On-policy Self-distillation for dLLMs
On-policy self-distillation explored for diffusion LLMs
DRFLOW: A Deep Research Benchmark for Personalized Workflow Prediction
DRFLOW: a deep research benchmark for personalized workflow prediction
All Smoke, No Alarm: Oracle Signals in Agent-Authored Test Code
Study finds agent-authored test code often lacks real verification logic
WEQA: Wearable hEalth Question Answering with Query-Adaptive Agentic Reasoning
WEQA: query-adaptive agentic reasoning for wearable health QA
Memory as a Wasting Asset: Pricing Flash Endurance for Embodied Agents, and the Limits of Doing So
Pricing flash endurance as a wasting asset for embodied agents
An agentic benchmark for implicit animal welfare in frontier AI
Knowledge Reutilization in Meta-Reinforcement Learning
A meta-knowledge reutilization framework for meta-RL across agents
Ternary Mamba: Grouped Quantization-Aware Training of W1.58A16 State Space Models
Ternary Mamba: grouped QAT for W1.58A16 state space models
HistoRAG embeds historical methodology into RAG via critical practice
A multi-agent framework against premature handoff and silent hallucination
PseudoBench: Measuring How Agentic Auto-Research Fuels Pseudoscience
PseudoBench measures how agentic auto-research fuels pseudoscience
Compositional Skill Routing for LLM Agents: Decompose, Retrieve, and Compose
Compositional skill routing for LLM agents: decompose, retrieve, compose
ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents
ProvenanceGuard: source-aware factuality verification for MCP agents
LoopCoder-v2: Only Loop Once for Efficient Test-Time Computation Scaling
LoopCoder-v2: loop once for efficient test-time compute scaling
Recursive Scaling in Masked Diffusion Models
Recursive scaling in masked diffusion models
LLM Consumer Behavior Theory: Foundations of a Novel Research Field
LLM Consumer Behavior Theory: a new field for agentic markets
Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models
Dynamic rollout editing reduces overthinking in RL reasoning models
Monotonic KANs: monotonicity as an inductive bias, studied theoretically
GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine?
GameCraft-Bench: can agents build playable games end-to-end?
Environment-Grounded Automated Prompt Optimization for LLM Game Agents
Environment-grounded automated prompt optimization for LLM game agents
From Drift to Coherence: Stabilizing Beliefs in LLMs
From drift to coherence: stabilizing beliefs in LLMs
A Framework for Evaluating Agentic Skills at Scale
A framework for evaluating agentic skills at scale
Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering
Position: coding benchmarks are misaligned with agentic software engineering
Vision-language models for chest radiography do not always need the image
Vision-language models for chest radiography do not always need the image
EComAgentBench: Benchmarking Shopping Agents on Long-Horizon Tasks with Distributed Hidden Intent
EComAgentBench: shopping agents on long-horizon tasks with hidden intent
LLMs Infer Cultural Context but Fail to Apply It When Responding
LLMs infer cultural context but fail to apply it when responding
SuCo: Sufficiency-guided Continuous Adaptive Reasoning
SuCo: sufficiency-guided continuous adaptive reasoning
EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning
EnvRL learns from environment dynamics in agentic RL
MambaCount: efficient open-vocabulary counting via state-space duality
Beyond Domains: Reusing Web Skills via Transferable Interaction Patterns
Reusing web skills via transferable interaction patterns
OPD-Evolver: Cultivating Holistic Agent Evolver via On-Policy Distillation
OPD-Evolver cultivates self-evolving agents via on-policy distillation
Context-Aware RL for Agentic and Multimodal LLMs
ContextRL rewards picking the right context to ground answers
Benchmarking LLM Agents on Meta-Analysis Articles from Nature Portfolio
A benchmark for LLM agents on Nature Portfolio meta-analyses
DeepRubric: evidence-tree rubrics to boost deep-research agent RL
HAMON: Passive Optical Sequence Mixing for Long-Horizon Forecasting
HAMON: a passive optical core for long-horizon forecasting
ExpRL: Exploratory RL for LLM Mid-Training
ExpRL uses human QA as reward scaffolds for LLM mid-training RL
TokenPilot: Cache-Efficient Context Management for LLM Agents
TokenPilot cuts LLM-agent context costs ~61% while preserving prompt cache
Agent trajectories as programs: fingerprinting and programming coding-agent behavior
Coding agents have behavioral fingerprints identifiable from trajectories
Greed Is Learned: Visible Incentives as Reward-Hacking Triggers
Greed Is Learned: RL agents get addicted to visible reward channels
LESS Is More: Mutual-Stability Sampling for Diffusion Language Models
LESS: a training-free adaptive sampler for diffusion language models
Binary Tracking for Spatial QA and Navigation with Open Vision-Language Models
Binary Tracking: open vision-language models for spatial QA and navigation
Semantic Flip: synthetic OOD generation for robust refusal in embodied agents
Reasoning hop-count predicts clinical AI failure in EHR QA
HawkesNest: A Multi-Axis Synthetic Benchmark for Spatiotemporal Pattern Complexity
HawkesNest: a synthetic benchmark for spatiotemporal point process models
RDS Fusion: neuro-symbolic gating with compressed CoT for irony detection
Data-Driven Decoding of Russell's Circumplex Model of Affect
Do Transformer embeddings recover Russell's circumplex affect geometry?
Beyond Models: Reflections on Engineering AI-enabled Systems in a Project-Based Course
Reflections on teaching the engineering of AI-enabled systems in a course
Does Traversal Order Matter? A Systematic Study of Tree Traversal Methods in Transformer Grammars
Paper: compares tree traversal orders in Transformer Grammars
Tying the Loop -- Tied Expert Layers in Mixture-of-Experts Language Models
Paper: Expert Tying shares MoE expert params across layers
Paper: framework measures LLM search-agent endorsement risk
GIST-CMTF: Goal-State Inference for Causal Minimal Tool Filtering in LLM Agents
GIST-CMTF adds goal-state inference to causal minimal tool filtering
Scaling LLM Reasoning from Minimal Labels: A Semi-Supervised Framework with a Lightweight Verifier
Paper: semi-supervised LLM reasoning from minimal labels
LabOSBench: Benchmarking Computer Use Agents for Scientific Instrument Control
LabOSBench: a simulated testbed for computer-use agents controlling instruments
OpenClaw-Skill: Collective Skill Tree Search for Agentic Large Language Models
OpenClaw-Skill: collective skill tree search for LLM agents
Skill-to-LoRA: From Using Skills to Learning Behaviors for Token-Efficient LLM Agents
S2L replaces runtime SKILL.md text with skill-specific LoRA adapters
MyPCBench: A Benchmark for Personally Intelligent Computer-Use Agents
MyPCBench: benchmarking personal computer-use agents
Misinformation Propagation in Benign Multi-Agent Systems
Study on misinformation propagation in benign multi-agent systems
Multi-Turn Reflective Masking Elicits Reasoning in Mask Diffusion Models
Reflective Masking elicits iterative reasoning in mask diffusion models
Multimodal Evaluator Preference Collapse: Cross-Modal Contagion in Self-Evolving Agents
Paper on evaluator preference collapse in self-evolving agents
FraudSMSWalker: Benchmarking Agentic Large Language Models for SMS-to-Webpage Fraud Detection
FraudSMSWalker benchmark targets URL-masked SMS-to-webpage fraud
Survey reviews Islamic LLMs and trustworthy, hallucination-resistant AI
VeriGraph: Towards Verifiable Data-Analytic Agents
VeriGraph: a traceable neuro-symbolic framework for verifiable data agents
SING: Synthetic Intention Graph for Scalable Active Tool Discovery in LLM Agents
SING: synthetic intention graph for scalable active tool discovery
Uncertainty Is Not a Safety Net for Clinical VQA, but Can It Anticipate Model Failure?
Uncertainty estimation fails as a safety net for clinical VQA
Can LLM Agents Infer World Models? Evidence from Agentic Automata Learning
Can LLM agents infer world models? Evidence from automata learning
BD-LSC: a new benchmark dataset for lexical semantic change detection
Can LLM Coding Agents Reason About Time Series?
Can LLM coding agents reason about time series? A benchmark study
daVinci-kernel: an RL framework co-evolving skills for GPU kernel tuning