DXC signaled a plan to embed Anthropic's Claude across heavily regulated industries. The signal pairs four official vendor sources (Anthropic, OpenAI, NVIDIA) with ITmedia coverage—enterprise adoption framed against multiple vendors' moves, a 'the floor is moving' story rather than research. The applied focus is embedding generative AI into real work in regulated sectors like finance and healthcare while preserving compliance; the through-line is less model novelty than the spread of enterprise adoption via a major IT-services integrator. But this is mainly a statement of intent—actual deployment scale, the effectiveness of regulatory handling, and results remain to be confirmed.
DXC embeds Claude across industries
DXC embeds Claude across industries
DXC will integrate Claude into the systems banks, airlines, and other regulated industries rely on
Anthropic, DXC form alliance to bring Claude to regulated industries
NVIDIA Achieves Leading Agentic Coding Performance on First Agentic AI Benchmark
NVIDIA tops first agentic AI benchmark for agentic coding performance
New OpenAI Academy courses for the next era of work
OpenAI launches three Academy courses on practical AI skills at work
AI agent bankrupted their operator while trying to scan DN42
AI agent 'bankrupts' its operator while attempting to scan DN42
Claude Fable is relentlessly proactive
Simon Willison: Claude Fable 5 is 'relentlessly proactive'
AIエージェントもフィッシング詐欺に引っかかる? 米セキュリティ企業がOpenClawで検証 結果は……
Varonis: AI agents can fall for phishing, tested with OpenClaw
Academic (arxiv etc.) 58 ▾
Learning Coordinated Preference for Multi-Objective Multi-Agent Reinforcement Learning
Learning coordinated preferences for multi-objective multi-agent RL
AgentSpec: Understanding Embodied Agent Scaffolds Through Controlled Composition
AgentSpec dissects embodied agent scaffolds via controlled composition
Towards Direct Latent-Space Synthesis for Parallel Branches in LLM-Agent Workflows
Direct latent-space synthesis for parallel branches in LLM-agent workflows
LoSoNA: A Benchmark for Local Social Norm Adaptation in Group Conversations
LoSoNA benchmarks local social norm adaptation in group chats
Regulating the Machine Contributor: Governance and Policy Alignment in Open Source
Governance and policy alignment for AI contributors in open source
SIMMER: Benchmarking Latent Failures in LLM Executable Planning with a World Model
SIMMER: benchmarking latent failures in LLM executable planning
From Shield to Target: Denial-of-Service Attacks on LLM-Based Agent Guardrails
From shield to target: DoS attacks on LLM-based agent guardrails
From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI
From chatbot to digital colleague: the shift to persistent autonomous AI
LLM agents defer blindly to GNN tools — stronger backbones defer more
GitOfThoughts: Version-Controlled Reasoning and Agent Memory You Can Replay, Diff, and Merge
GitOfThoughts: version-controlled reasoning and agent memory
tap: A File-Based Protocol for Heterogeneous LLM Agent Collaboration
tap: a file-based protocol for heterogeneous LLM agent collaboration
Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments
Re-evaluating agent capabilities beyond familiar environments
Retrospective Progress-Aware Self-Refinement for LLM Agent Training
Progress-aware self-refinement for training LLM agents
CacheRL:Multi-Turn Tool-Calling Agents via Cached Rollouts and Hybrid Reward
CacheRL trains tool-calling agents via cached rollouts and hybrid reward
Same-Origin Policy for Agentic Browsers
A same-origin policy for agentic browsers
Dialogue SWE-Bench: A Benchmark for Dialogue-Driven Coding Agents
Dialogue SWE-Bench: a benchmark for dialogue-driven coding agents
EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments
EvoArena tests LLM agents in evolving environments with patch-based memory
SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning
SpatialClaw adopts code as the action interface for agentic spatial reasoning
Agents-K1: Towards Agent-native Knowledge Orchestration
Agents-K1 turns raw papers into agent-native scientific knowledge graphs
HyperTool: Beyond Step-Wise Tool Calls for Tool-Augmented Agents
HyperTool folds tool workflows into code, lifting MCP-Universe to 35.29%
EurekAgent: Agent Environment Engineering is All You Need For Autonomous Scientific Discovery
EurekAgent argues environment engineering drives autonomous scientific discovery
Recursive Agent Harnesses lift long-context coding accuracy to 81.36%
AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility
AgentBeats proposes agentified, protocol-standardized agent assessment (AAA)
EpiBench: Verifiable Evaluation of AI Agents on Epigenomics Analysis
EpiBench: no AI agent passes a majority of epigenomics analysis tasks
Reward Modeling for Multi-Agent Orchestration
OrchRM: self-supervised reward modeling for multi-agent orchestration
Multiagent Protocols with Aggregated Confidence Signals
Protocols aggregate multi-agent confidence into a single discriminative score
Adaptive Turn-Taking for Real-time Multi-Party Voice Agents
ModeratorLM conditions turn-taking on assigned roles for multi-party voice
46% of agentic PR fixes are rejected; study identifies 14 reasons
Reinforcement Learning for Neural Model Editing
RL agents learn to edit neural models for unlearning and debiasing
Toward Instructions-as-Code: Understanding the Impact of Instruction Files on Agentic Pull Requests
Instruction files don't necessarily improve agentic pull request outcomes
Paper argues LLMs lack moral agency: 'sampling is not choosing'
Neuro-Symbolic Agents for Regulated Process Automation: Challenges and Research Agenda
Neuro-symbolic agents propose 'compliance-by-construction' for regulated work
Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents
Stakeholder-centric benchmark attributes prompt-injection harm in web agents
An LLM System for Autonomous Variational Quantum Circuit Design
LLM agent framework autonomously designs variational quantum circuits
SkillCAT: Contrastive Assessment and Topology-Aware Skill Self-Evolution for LLM Agents
SkillCAT verifies and routes self-evolved skills for LLM agents
RogueAI: A Reverse Turing Test for Detecting Licensed AI Deception in Dialogue
RogueAI: a reverse Turing test for spotting licensed AI deception
ComAct: Reframing Professional Software Manipulation via COM-as-Action Paradigm
ComAct drives industrial CAD software via COM-as-Action paradigm
MemRefine: LLM-Guided Compression for Long-Term Agent Memory
MemRefine compresses agent memory to budget with LLM-judged merging
TRACE compiles user corrections into runtime checks for coding agents
Context-Driven Incremental Compression for Multi-Turn Dialogue Generation
C-DIC compresses multi-turn dialogue context incrementally for stability
DIRECT: When and Where Should You Allocate Test-Time Compute in Embodied Planners?
DIRECT allocates test-time compute per prompt for embodied planners
ATLAS: Active Theory Learning for Automated Science
ATLAS uses active learning to discover interpretable behavioral models
APPO: Agentic Procedural Policy Optimization
APPO branches and assigns credit at fine-grained decision points
Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks
Claw-SWE-Bench fairly benchmarks OpenClaw-style coding agent harnesses
Fourier Features Let Agents Learn High Precision Policies with Imitation Learning
Fourier features give point-cloud policies high-precision control
PROJECTMEM: A Local-First, Event-Sourced Memory and Judgment Layer for AI Coding Agents
projectmem adds a local-first memory layer for AI coding agents
A Five-Plane Reference Architecture for Runtime Governance of Production AI Agents
A five-plane reference architecture for runtime governance of AI agents
CCKS: Consensus-based Communication and Knowledge Sharing
CCKS improves cooperative MARL via consensus-based knowledge sharing
The Impossibility of Eliciting Latent Knowledge
Paper formalizes the impossibility of eliciting latent knowledge
Survey of automated embodied benchmark construction in five stages
Implicit Neural Representations of Individual Behavior
Behavioral INR learns policy representations from unlabeled multi-policy data
Survey of agentic environment engineering across its lifecycle
Towards Responsibly Non-Compliant Machines
Position paper sketches responsibly non-compliant intelligent machines
FORT-Searcher: Synthesizing Shortcut-Resistant Search Tasks for Training Deep Search Agents
FORT synthesizes shortcut-resistant tasks to train deep search agents
Toward Generalist Autonomous Research via Hypothesis-Tree Refinement
Arbor runs autonomous research via hypothesis-tree refinement
Language robustness in VLA models is a step-wise control problem
Notes2Skills: From Lab Notebooks to Certainty-Aware Scientific Agent Skills
Notes2Skills turns lab notebooks into certainty-aware agent skills
WorldReasoner: Evaluating Whether Language Model Agents Forecast Events with Valid Reasoning
WorldReasoner evaluates whether agents forecast events with valid reasoning