AI エージェント × 安全性・評価

HF、ロボ連携Strands+LeRobotエージェント公開

HF、ロボ連携Strands+LeRobotエージェント公開

✎ ストーリー本文

Hugging Face が Strands エージェントと LeRobot を橋渡しし、Hub のモデルからロボット実機を動かす連携を公開した。発生元は HF 公式が起点で、itmedia など専門報道が追う構成――学術やコミュニティより「使える形にまとまった」ことを伝える報道が厚く、研究段階から実装・配布段階へ寄った動きと読める。主眼はモデル単体ではなく、既存のエージェント基盤とロボットハードウェアをどうつなぐかにある。オープンソースの部品が組み合わさって身体性へ向かう流れだが、実運用での安定性や対応ハードの広がりは、公開直後の段階では未知数だ。

▲ 公式・報道
公式

From the Hugging Face Hub to robot hardware with Strands Agents and LeRobot

Hugging Face Blog ・ 2026-06-17 ・ 📌

Strands AgentsとLeRobot、Hugging Faceのモデルをロボット実機へ

報道

工数「76%」削減 味の素グループが「経理AIエージェント」導入で先陣を切れたワケ

ITmedia AI+ ・ 2026-06-18

味の素グループ、経理AIエージェント導入で承認業務を自律化し工数76%削減

報道

「待ちの営業」はもう限界 ホンダがAIエージェントで挑む、商機を逃さない「濃い商談」の創出

ITmedia AI+ ・ 2026-06-18

ホンダ、新車販売にAIエージェント導入で“濃い商談”を支援し成約創出

報道

話題の「Claude Mythos」登場で変わるセキュリティ AIエージェント時代の防衛策

ITmedia AI+ ・ 2026-06-18

Claude Mythos登場でAI攻撃が時間単位に、エージェント時代の新防衛策

コミュニティ

Announcing Stack Overflow for Agents

Lobste.rs (AI tagged) ・ 2026-06-18

「Stack Overflow for Agents」発表

報道

かんぽ生命、AIで営業支援 “郵便局での一言”拾って保険提案へ 寸劇で分かる活用例

ITmedia AI+ ・ 2026-06-17

かんぽ生命、AIエージェントで営業支援を本格化

報道

「ポケカ対戦AIエージェント」開発コンテスト開始 「不完全情報ゲーム」をどう制するか

ITmedia AI+ ・ 2026-06-17

ポケモンカード対戦AIエージェントの開発コンテスト開始、不完全情報ゲームに挑む

公式

Agentic Resource Discovery: Let agents search

Hugging Face Blog ・ 2026-06-17

Hugging Face、エージェントが自ら資源を探索する手法を提案

学術(arxiv ほか) 42本 ▾
学術

Execution-State Capsules: Graph-Bound Execution-State Checkpoint and Restore for Low-Latency, Small-Batch, On-Device Physical-AI Serving

arXiv cs.LG (Machine Learning) ・ 2026-06-18

小バッチ・低遅延なオンデバイスAI推論向けの実行状態保存復元機構

学術

LedgerAgent: Structured State for Policy-Adherent Tool-Calling Agents

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-18

方針順守のツール呼び出しエージェントに構造化状態を与えるLedgerAgent

学術

Sovereign Execution Brokers: Enforcing Certificate-Bound Authority in Agentic Control Planes

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-18

エージェント制御基盤で証明書束縛の権限を強制するSovereign Execution Brokers

学術

Probe-and-Refine Tuning of Repository Guidance for Coding Agents

arXiv cs.LG (Machine Learning) ・ 2026-06-18

コーディングエージェント向けにリポジトリ指示文を調整する手法

学術

Efficient and Sound Probabilistic Verification for AI Agents

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-18

AIエージェント向けの効率的で健全な確率的検証手法

学術

Contagion Networks: Evaluator Bias Propagation in Multi-Agent LLM Systems

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-18

マルチエージェントLLMで評価者バイアスが伝播する現象を分析

学術

Beyond Global Replanning: Hierarchical Recovery for Cross-Device Agent Systems

arXiv cs.CL (Computation and Language) ・ 2026-06-18

複数端末にまたがるエージェントの階層的な障害回復手法を提案

学術

Optimal Order of Multi-Agent and General Many-Body Systems

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-18

フィードバックを持つマルチエージェント・多体系の最適次数の枠組み

学術

UltraQuant: 4-bit KV Caching for Context-Heavy Agents

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-18

文脈の重いエージェント向け4ビットKVキャッシュUltraQuant

学術

Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-18

エージェントAIへのモデル誘導自動攻撃に対する防御的かく乱の分析

学術

LLM agent safety, multi-turn red-teaming, jailbreak benchmarks, adversarial robustness, safety-critical systems

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-18

安全重要系を監督するLLMエージェントの多ターン・レッドチーミング評価

学術

CRAX: Fast Safe Reinforcement Learning Benchmarking

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-18

安全な強化学習を高速にベンチマークするCRAX

学術

AutoPass: Evidence-Guided LLM Agents for Compiler Performance Tuning

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-18

コンパイラ性能調整を担う証拠誘導LLMエージェントAutoPass

学術

Automating SKILL.md Generation for Computer-Using Agents via Interaction Trajectory Mining

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-18

相互作用軌跡採掘でコンピュータ操作エージェントのSKILL.md生成を自動化

学術

A Model-Driven Approach for Developing Families of Reinforcement Learning Environments

arXiv cs.LG (Machine Learning) ・ 2026-06-18

強化学習環境のファミリーを開発するモデル駆動手法を提案

学術

ScholarQuest: A Taxonomy-Guided Benchmark for Agentic Academic Paper Search in Open Literature Environments

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-18

開かれた文献環境でのエージェント論文探索ベンチScholarQuest

学術

Augmenting Game AI with Deep Reinforcement Learning

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-18

深層強化学習でゲームAIを強化する研究

学術

FlowMaps: Modeling Long-Term Multimodal Object Dynamics with Flow Matching

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-18

フローマッチングで長期のマルチモーダル物体動態を予測するFlowMaps

学術

MedRLM: Recursive Multimodal Health Intelligence for Long-Context Clinical Reasoning, Sensor-Guided Screening, Evidence-Grounded Decision Support, and Community-to-Tertiary Referral Optimization

arXiv cs.CL (Computation and Language) ・ 2026-06-18

長文脈の臨床推論に向けた再帰的マルチモーダル医療AI「MedRLM」

学術

When Lower Privileges Suffice: Investigating Over-Privileged Tool Selection in LLM Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-18

LLMエージェントによる過剰権限のツール選択を調査

学術

Connect the Dots: Training LLMs for Long-Lifecycle Agents with Cross-Domain Generalization Via Reinforcement Learning

arXiv cs.CL (Computation and Language) ・ 2026-06-18

長期運用エージェント向けに強化学習でLLMを訓練するCoD

学術

Multi-Agent Transactive Memory

arXiv cs.CL (Computation and Language) ・ 2026-06-18

異種エージェント間で知識共有する多エージェント交換記憶

学術

AtomMem: Building Simple and Effective Memory System for LLM Agents via Atomic Facts

arXiv cs.CL (Computation and Language) ・ 2026-06-18

原子的事実でLLMエージェントの記憶システムを構築するAtomMem

学術

JAMER: Project-Level Code Framework Dataset and Benchmark on Professional Game Engines

arXiv cs.CL (Computation and Language) ・ 2026-06-18

プロ向けゲームエンジンのプロジェクト級コードデータセットJAMER

学術

AgentFinVQA: A Deployable Multi-Agent Pipeline for Auditable Financial Chart QA

arXiv cs.CL (Computation and Language) ・ 2026-06-18

監査可能な金融チャートQA向け多エージェントAgentFinVQA

学術

Data Intelligence Agents: Interpreting, Modeling, and Querying Enterprise Data via Autonomous Coding Agents

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

自律コーディングエージェントで企業データを解釈・問い合わせ

学術

Enhancing Decision-Making with Large Language Models through Multi-Agent Fictitious Play

arXiv cs.CL (Computation and Language) ・ 2026-06-17

マルチエージェントの架空プレイで LLM の意思決定を強化

学術

Optimal scenario design for climate emulation

arXiv cs.LG (Machine Learning) ・ 2026-06-17

気候エミュレーションの精度を高める最適シナリオ設計を提案

学術

Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models

arXiv cs.LG (Machine Learning) ・ 2026-06-17

VLA モデルは常識を保持しているか、知識保持度を測る研究

学術

TxBench-PP: Analyzing AI Agent Performance on Small-Molecule Preclinical Pharmacology

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

TxBench-PP、小分子前臨床薬理での AI エージェント性能を評価

学術

Learning to Annotate Delayed and False AEB Events: A Practical System for Extreme Class Imbalance and Asymmetric Label Noise

arXiv cs.LG (Machine Learning) ・ 2026-06-17

遅延・誤作動 AEB 事象を学習注釈、極端な不均衡に対応する実用系

学術

AdsMind: A Physics-Grounded Multi-Agent System for Self-Correcting Discovery of Adsorption Configurations on Heterogeneous Catalyst Surfaces

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

AdsMind、物理基盤のマルチエージェントで触媒の吸着配置を探索

学術

A Technical Taxonomy of LLM Agent Communication Protocols

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

LLM エージェントの通信プロトコルを技術的に分類・整理

学術

Towards an Agent-First Web: Redesigning the Web for AI Agents

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

エージェント優先の Web へ、AI 向けに Web を再設計する提案

学術

RODS: Reward-Driven Online Data Synthesis for Multi-Turn Tool-Use Agents

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

RODS、報酬駆動のオンラインデータ合成で多ターンツール利用を強化

学術

TRAP: Benchmark for Task-completion and Resistance to Active Privacy-extraction

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

TRAP、課題遂行とプライバシー抽出耐性を測るエージェント評価

学術

CAPRA: Scaling Feedback on Software Architecture Deliverables with a Multi-Agent LLM System

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

ソフト設計成果物の添削を自動化するマルチエージェントLLM「CAPRA」

学術

RTSGameBench: An RTS Benchmark for Strategic Reasoning by Vision-Language Models

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

視覚言語モデルの戦略的推論を測るRTSベンチマーク

学術

Decoupling Search from Reasoning: A Vendor-Agnostic Grounding Architecture for LLM Agents

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

検索と推論を分離するベンダー非依存のLLMエージェント基盤

学術

Beyond Reward Engineering: A Data Recipe for Long-Context Reinforcement Learning

arXiv cs.CL (Computation and Language) ・ 2026-06-17

長文脈強化学習のためのデータレシピ

学術

GateMem: Benchmarking Memory Governance in Multi-Principal Shared-Memory Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-17

多主体共有メモリのガバナンスを測る「GateMem」

学術

LegalWorld: A Life-Cycle Interactive Environment for Legal Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-17

法務エージェント向けライフサイクル型環境「LegalWorld」

← ストーリー アーカイブ