DXC が規制の厳しい業界の業務に Anthropic の Claude を統合していく方針を示した。発生元は Anthropic・OpenAI・NVIDIA など公式4件に itmedia の報道が付く構成で、複数ベンダーの動きを背景に事業導入が語られる――研究ではなく「現場が動いた」側の動きだ。金融や医療など規制産業で、コンプライアンスを保ちながら生成 AI を実務に組み込むという応用が主眼で、本流はモデルの新規性より、大手 IT サービス企業を介した企業導入の広がりにある。ただし現時点は統合の方針表明が中心で、実際の展開規模や規制対応の実効性、成果は、これからの確認点として残る。
DXC、規制業界の業務にClaudeを統合へ
DXC、規制業界の業務にClaudeを統合へ
DXC will integrate Claude into the systems banks, airlines, and other regulated industries rely on
Anthropic、DXCと複数年提携、銀行・航空など規制業界の基幹システムにClaude統合
NVIDIA Achieves Leading Agentic Coding Performance on First Agentic AI Benchmark
NVIDIA、初のエージェント型AIベンチマークでコーディング性能首位を達成
New OpenAI Academy courses for the next era of work
OpenAI、仕事での AI 活用を学ぶ Academy 新コース 3 種を公開
AI agent bankrupted their operator while trying to scan DN42
AIエージェント、DN42スキャン試行で運用者に破産級のコストを発生との報告
Claude Fable is relentlessly proactive
Simon Willison氏、Claude Fable 5を「徹底的にproactive」と評価
AIエージェントもフィッシング詐欺に引っかかる? 米セキュリティ企業がOpenClawで検証 結果は……
AIエージェントもフィッシングに引っかかる VaronisがOpenClawで検証
学術(arxiv ほか) 58本 ▾
Learning Coordinated Preference for Multi-Objective Multi-Agent Reinforcement Learning
多目的・多エージェント強化学習で協調選好を学習する手法を提案
AgentSpec: Understanding Embodied Agent Scaffolds Through Controlled Composition
AgentSpec、エージェント足場を統制的に分解し検証
Towards Direct Latent-Space Synthesis for Parallel Branches in LLM-Agent Workflows
LLMエージェントの並列分岐を潜在空間で直接合成する手法を検討
LoSoNA: A Benchmark for Local Social Norm Adaptation in Group Conversations
LoSoNA、集団会話の局所規範への適応を評価する基準
Regulating the Machine Contributor: Governance and Policy Alignment in Open Source
オープンソースへのAIエージェント貢献を統治する枠組みを論じる
SIMMER: Benchmarking Latent Failures in LLM Executable Planning with a World Model
世界モデルでLLM実行計画の潜在的失敗を測る「SIMMER」
From Shield to Target: Denial-of-Service Attacks on LLM-Based Agent Guardrails
LLMエージェントのガードレールを狙うサービス妨害攻撃を提示
From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI
LLMが「チャットボット」から永続自律AIへ移行する転換を概念化
LLMエージェントはGNNツールに盲目的に委ね、強いほど委ねると指摘
GitOfThoughts: Version-Controlled Reasoning and Agent Memory You Can Replay, Diff, and Merge
再生・差分・統合できる版管理つき推論記憶「GitOfThoughts」
tap: A File-Based Protocol for Heterogeneous LLM Agent Collaboration
異種LLMエージェント協調のファイルベースプロトコル「tap」
Retrospective Progress-Aware Self-Refinement for LLM Agent Training
進捗を自覚する自己改善でLLMエージェント訓練を強化
CacheRL:Multi-Turn Tool-Calling Agents via Cached Rollouts and Hybrid Reward
CacheRL、キャッシュ活用で小型ツール呼出エージェントを訓練
Same-Origin Policy for Agentic Browsers
エージェント型ブラウザに同一生成元ポリシーを提案
Dialogue SWE-Bench: A Benchmark for Dialogue-Driven Coding Agents
Dialogue SWE-Bench、対話駆動のコーディングを評価
EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments
動的環境でのLLMエージェント評価基盤EvoArena、パッチ型記憶EvoMemも提案
SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning
SpatialClaw、空間推論agentの行動IFをコードに刷新
Agents-K1: Towards Agent-native Knowledge Orchestration
Agents-K1、論文をagent向け知識グラフ化する基盤
HyperTool: Beyond Step-Wise Tool Calls for Tool-Augmented Agents
HyperTool、ツール呼び出しをコードに畳み込みMCP精度を15.7%から35.3%へ
EurekAgent: Agent Environment Engineering is All You Need For Autonomous Scientific Discovery
EurekAgent、自律的科学発見の鍵は環境設計と主張
再帰エージェントハーネスRAH、長文脈推論でCodex基準71.75%を81.36%に
AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility
AgentBeats、judge agentで評価を標準化するAAAを提唱
EpiBench: Verifiable Evaluation of AI Agents on Epigenomics Analysis
EpiBench、AI agentのエピゲノム解析を検証、最高45%
Reward Modeling for Multi-Agent Orchestration
OrchRM、多agent統括の品質を自己教師で評価し効率10倍
Multiagent Protocols with Aggregated Confidence Signals
多agent系の出力に単一の集約信頼度を与える3手法を提案
Adaptive Turn-Taking for Real-time Multi-Party Voice Agents
ModeratorLM、役割条件で多者間音声agentの発話交代を改善
AI agentのPR修正46%が却下、14の却下理由を分類
Reinforcement Learning for Neural Model Editing
モデル編集を強化学習で自動化、忘却精度ほぼ0%・保持9割超を達成
Toward Instructions-as-Code: Understanding the Impact of Instruction Files on Agentic Pull Requests
AI agent向け指示ファイルは必ずしもPR成果を改善せずと判明
LLMに道徳的主体性はないと論じる『標本抽出は選択でない』
Neuro-Symbolic Agents for Regulated Process Automation: Challenges and Research Agenda
規制業務の自動化に『構築による準拠』を提唱する神経記号agent
Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents
web agentのプロンプト注入を被害者視点で評価するベンチ
An LLM System for Autonomous Variational Quantum Circuit Design
LLMで量子回路を自律設計する閉ループagent枠組み
SkillCAT: Contrastive Assessment and Topology-Aware Skill Self-Evolution for LLM Agents
SkillCAT、LLMエージェントのスキル自己進化を検証付き3段階に分離
RogueAI: A Reverse Turing Test for Detecting Licensed AI Deception in Dialogue
逆チューリングテストRogueAI、嘘を許可されたAIを対話で見破るゲーム
ComAct: Reframing Professional Software Manipulation via COM-as-Action Paradigm
ComAct、COM経由の決定論的操作で産業CADソフトをエージェント制御
MemRefine: LLM-Guided Compression for Long-Term Agent Memory
MemRefine、LLM判断で予算内にエージェント長期記憶を圧縮
TRACE、ユーザー修正をルール化しコーディングエージェントに強制適用
Context-Driven Incremental Compression for Multi-Turn Dialogue Generation
C-DIC、多ターン対話の文脈を逐次圧縮し長対話を安定化
DIRECT: When and Where Should You Allocate Test-Time Compute in Embodied Planners?
DIRECT、身体エージェントのテスト時計算をプロンプト別に配分
ATLAS: Active Theory Learning for Automated Science
ATLAS、能動学習で解釈可能な行動モデルを自動的に発見
APPO: Agentic Procedural Policy Optimization
APPO、細粒度の決定点で分岐と信用割当を行うエージェントRL
Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks
Claw-SWE-Bench、OpenClaw型エージェントのコード能力を多言語評価
Fourier Features Let Agents Learn High Precision Policies with Imitation Learning
Fourier 特徴で点群方策が高精度ロボット操作を獲得
PROJECTMEM: A Local-First, Event-Sourced Memory and Judgment Layer for AI Coding Agents
projectmem、コーディングエージェントに局所優先の記憶層を追加
A Five-Plane Reference Architecture for Runtime Governance of Production AI Agents
本番AIエージェントの実行時統治に5プレーン参照アーキテクチャ
CCKS: Consensus-based Communication and Knowledge Sharing
CCKS、合意ベースの通信と知識共有で協調的MARLを改善
The Impossibility of Eliciting Latent Knowledge
潜在知識の引き出し(ELK)は不可能と因果影響図で形式化
Implicit Neural Representations of Individual Behavior
Behavioral INR、無ラベルの多方策行動から方策表現を学習
Towards Responsibly Non-Compliant Machines
責任ある「不服従」が可能なAIエージェントの工学を提起
FORT-Searcher: Synthesizing Shortcut-Resistant Search Tasks for Training Deep Search Agents
FORT、近道耐性のある探索課題を合成し深層探索エージェントを訓練
Toward Generalist Autonomous Research via Hypothesis-Tree Refinement
Arbor、仮説ツリー改良で長期間の自律研究ループを運用
Notes2Skills: From Lab Notebooks to Certainty-Aware Scientific Agent Skills
Notes2Skills、実験ノートを確信度付きの科学エージェントスキルに変換
WorldReasoner: Evaluating Whether Language Model Agents Forecast Events with Valid Reasoning
WorldReasoner、エージェントの事象予測の推論妥当性を評価