AI Agents × Safety & Evaluation

Hugging Face links agents to robots

Hugging Face links agents to robots

✎ Story body

Hugging Face bridged its Strands agents with LeRobot, letting models from the Hub drive physical robots. The signal starts from an official HF release and is carried by trade press like ITmedia—coverage emphasizing that something usable has shipped, rather than academic or community discussion. It reads as a shift from research toward implementation and distribution. The focus is not the model itself but how to connect existing agent frameworks to robot hardware. Open-source parts are composing toward embodiment, but real-world stability and the range of supported hardware are unknowns right after release.

▲ Official & Press
Official

From the Hugging Face Hub to robot hardware with Strands Agents and LeRobot

Hugging Face Blog ・ 2026-06-17 ・ 📌

From Hugging Face Hub to robot hardware with Strands Agents and LeRobot

Press

工数「76%」削減 味の素グループが「経理AIエージェント」導入で先陣を切れたワケ

ITmedia AI+ ・ 2026-06-18

Ajinomoto deploys autonomous accounting AI agent, cuts workload 76%

Press

「待ちの営業」はもう限界 ホンダがAIエージェントで挑む、商機を逃さない「濃い商談」の創出

ITmedia AI+ ・ 2026-06-18

Honda brings AI agents to car sales to drive higher-quality deals

Press

話題の「Claude Mythos」登場で変わるセキュリティ AIエージェント時代の防衛策

ITmedia AI+ ・ 2026-06-18

Claude Mythos reshapes security as AI attacks turn hourly

Community

Announcing Stack Overflow for Agents

Lobste.rs (AI tagged) ・ 2026-06-18

Announcing Stack Overflow for Agents

Press

かんぽ生命、AIで営業支援 “郵便局での一言”拾って保険提案へ 寸劇で分かる活用例

ITmedia AI+ ・ 2026-06-17

Japan Post Insurance adds AI agents to its sales workflow

Press

「ポケカ対戦AIエージェント」開発コンテスト開始 「不完全情報ゲーム」をどう制するか

ITmedia AI+ ・ 2026-06-17

Contest launches to build AI agents for Pokemon TCG, an imperfect-info game

Official

Agentic Resource Discovery: Let agents search

Hugging Face Blog ・ 2026-06-17

Hugging Face proposes agentic resource discovery via search

Academic (arxiv etc.) 42 ▾
Academic

Execution-State Capsules: Graph-Bound Execution-State Checkpoint and Restore for Low-Latency, Small-Batch, On-Device Physical-AI Serving

arXiv cs.LG (Machine Learning) ・ 2026-06-18

Execution-State Capsules: checkpoint/restore for on-device AI serving

Academic

LedgerAgent: Structured State for Policy-Adherent Tool-Calling Agents

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-18

LedgerAgent: structured state for policy-adherent tool-calling agents

Academic

Sovereign Execution Brokers: Enforcing Certificate-Bound Authority in Agentic Control Planes

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-18

Sovereign Execution Brokers for agentic control planes

Academic

Probe-and-Refine Tuning of Repository Guidance for Coding Agents

arXiv cs.LG (Machine Learning) ・ 2026-06-18

Probe-and-Refine: tuning repository guidance for coding agents

Academic

Efficient and Sound Probabilistic Verification for AI Agents

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-18

Efficient and sound probabilistic verification for AI agents

Academic

Contagion Networks: Evaluator Bias Propagation in Multi-Agent LLM Systems

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-18

Contagion Networks: evaluator bias propagation in multi-agent LLMs

Academic

Beyond Global Replanning: Hierarchical Recovery for Cross-Device Agent Systems

arXiv cs.CL (Computation and Language) ・ 2026-06-18

Hierarchical recovery for cross-device agent systems

Academic

Optimal Order of Multi-Agent and General Many-Body Systems

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-18

Optimal order of multi-agent and general many-body systems

Academic

UltraQuant: 4-bit KV Caching for Context-Heavy Agents

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-18

UltraQuant: 4-bit KV caching for context-heavy agents

Academic

Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-18

Analyzing defensive misdirection against attacks on agentic AI

Academic

LLM agent safety, multi-turn red-teaming, jailbreak benchmarks, adversarial robustness, safety-critical systems

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-18

Multi-turn red-teaming of LLM agents for safety-critical systems

Academic

CRAX: Fast Safe Reinforcement Learning Benchmarking

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-18

CRAX: fast benchmarking for safe reinforcement learning

Academic

AutoPass: Evidence-Guided LLM Agents for Compiler Performance Tuning

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-18

AutoPass: evidence-guided LLM agents for compiler performance tuning

Academic

Automating SKILL.md Generation for Computer-Using Agents via Interaction Trajectory Mining

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-18

Automating SKILL.md generation via interaction trajectory mining

Academic

A Model-Driven Approach for Developing Families of Reinforcement Learning Environments

arXiv cs.LG (Machine Learning) ・ 2026-06-18

A model-driven approach to building families of RL environments

Academic

ScholarQuest: A Taxonomy-Guided Benchmark for Agentic Academic Paper Search in Open Literature Environments

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-18

ScholarQuest: a taxonomy-guided benchmark for agentic paper search

Academic

Augmenting Game AI with Deep Reinforcement Learning

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-18

Augmenting game AI with deep reinforcement learning

Academic

FlowMaps: Modeling Long-Term Multimodal Object Dynamics with Flow Matching

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-18

FlowMaps: long-term multimodal object dynamics with flow matching

Academic

MedRLM: Recursive Multimodal Health Intelligence for Long-Context Clinical Reasoning, Sensor-Guided Screening, Evidence-Grounded Decision Support, and Community-to-Tertiary Referral Optimization

arXiv cs.CL (Computation and Language) ・ 2026-06-18

MedRLM: recursive multimodal AI for long-context clinical reasoning

Academic

When Lower Privileges Suffice: Investigating Over-Privileged Tool Selection in LLM Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-18

Investigating over-privileged tool selection in LLM agents

Academic

Connect the Dots: Training LLMs for Long-Lifecycle Agents with Cross-Domain Generalization Via Reinforcement Learning

arXiv cs.CL (Computation and Language) ・ 2026-06-18

Connect the Dots: RL training for long-lifecycle LLM agents

Academic

Multi-Agent Transactive Memory

arXiv cs.CL (Computation and Language) ・ 2026-06-18

Multi-agent transactive memory for sharing knowledge across agents

Academic

AtomMem: Building Simple and Effective Memory System for LLM Agents via Atomic Facts

arXiv cs.CL (Computation and Language) ・ 2026-06-18

AtomMem: an LLM-agent memory system built on atomic facts

Academic

JAMER: Project-Level Code Framework Dataset and Benchmark on Professional Game Engines

arXiv cs.CL (Computation and Language) ・ 2026-06-18

JAMER: a project-level code benchmark on game engines

Academic

AgentFinVQA: A Deployable Multi-Agent Pipeline for Auditable Financial Chart QA

arXiv cs.CL (Computation and Language) ・ 2026-06-18

AgentFinVQA: an auditable multi-agent pipeline for financial chart QA

Academic

Data Intelligence Agents: Interpreting, Modeling, and Querying Enterprise Data via Autonomous Coding Agents

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

Data Intelligence Agents query enterprise data autonomously

Academic

Enhancing Decision-Making with Large Language Models through Multi-Agent Fictitious Play

arXiv cs.CL (Computation and Language) ・ 2026-06-17

Multi-agent fictitious play boosts LLM decision-making

Academic

Optimal scenario design for climate emulation

arXiv cs.LG (Machine Learning) ・ 2026-06-17

Optimal scenario design improves climate emulation surrogates

Academic

Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models

arXiv cs.LG (Machine Learning) ・ 2026-06-17

Measuring commonsense and knowledge retention in VLA models

Academic

TxBench-PP: Analyzing AI Agent Performance on Small-Molecule Preclinical Pharmacology

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

TxBench-PP evaluates AI agents on preclinical pharmacology

Academic

Learning to Annotate Delayed and False AEB Events: A Practical System for Extreme Class Imbalance and Asymmetric Label Noise

arXiv cs.LG (Machine Learning) ・ 2026-06-17

Annotating rare delayed and false AEB events under class imbalance

Academic

AdsMind: A Physics-Grounded Multi-Agent System for Self-Correcting Discovery of Adsorption Configurations on Heterogeneous Catalyst Surfaces

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

AdsMind: physics-grounded multi-agent search for adsorption configs

Academic

A Technical Taxonomy of LLM Agent Communication Protocols

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

A technical taxonomy of LLM agent communication protocols

Academic

Towards an Agent-First Web: Redesigning the Web for AI Agents

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

Towards an agent-first web: redesigning the web for AI agents

Academic

RODS: Reward-Driven Online Data Synthesis for Multi-Turn Tool-Use Agents

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

RODS: reward-driven online data synthesis for tool-use agents

Academic

TRAP: Benchmark for Task-completion and Resistance to Active Privacy-extraction

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

TRAP benchmarks agents on task completion and privacy resistance

Academic

CAPRA: Scaling Feedback on Software Architecture Deliverables with a Multi-Agent LLM System

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

CAPRA: a multi-agent LLM system for software architecture feedback

Academic

RTSGameBench: An RTS Benchmark for Strategic Reasoning by Vision-Language Models

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

RTSGameBench: an RTS benchmark for strategic reasoning by VLMs

Academic

Decoupling Search from Reasoning: A Vendor-Agnostic Grounding Architecture for LLM Agents

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

Decoupling search from reasoning: a vendor-agnostic grounding architecture

Academic

Beyond Reward Engineering: A Data Recipe for Long-Context Reinforcement Learning

arXiv cs.CL (Computation and Language) ・ 2026-06-17

Beyond reward engineering: a data recipe for long-context RL

Academic

GateMem: Benchmarking Memory Governance in Multi-Principal Shared-Memory Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-17

GateMem: benchmarking memory governance in shared-memory agents

Academic

LegalWorld: A Life-Cycle Interactive Environment for Legal Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-17

LegalWorld: a life-cycle interactive environment for legal agents

← Story Archive