Agents & Tool Use

A
Showing 31–39 of 39
  • arXiv cs.AI (Artificial Intelligence) · EN Multimodal
    Bench2Dex: Benchmarking Visuo-Tactile Bimanual Dexterous Manipulation Across Dexterous Hands
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.CL (Computation and Language) · EN Developer Tools
    Turkish MMLU Pro: Traceable Option Augmentation and Its Validity Limits in Turkish Multiple-Choice Evaluation
    Embeddings Machine Learning Software Engineering
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • IEEE Spectrum (AI section) · EN Agents & Tool Use
    Why Andon Labs Puts AI Agents in Charge of Real Businesses
    Andon Labs puts AI agents in charge of real stores to measure autonomy
    AI Agents Anthropic Claude Google OpenAI
    AI safety firm Andon Labs runs a real San Francisco store and cafe managed by AI agents to measure how much real-world responsibility they can handle. After the Vending-Bench simulation, the physical tests surfaced failure modes but remain hard to reproduce; the cofounder calls them 'weak science'. Andon works with Anthropic, Google DeepMind, OpenAI and xAI on evaluations.
    Read original (IEEE Spectrum (AI section)) ↗
  • Publickey · JA Agents & Tool Use
    TailscaleのVPNにAIエージェントを組み込める「Aperture」正式リリース。AIによるTailscaleやノードの操作も可能に
    Tailscale ships Aperture AI gateway GA, adds MCPs for agent-driven network ops
    AI Agents Machine Learning
    Tailscale announced general availability of Aperture, an AI gateway that lets any node in a tailnet reach AI services without distributing API keys. It now bundles cost controls, guardrails, audit logs, MCP/API proxies, and in-gateway token purchasing. New Tailscale MCP and Tailscale SSH MCP let AI agents add nodes, SSH into machines, and deploy services via Aperture's chat UI, with existing ACLs enforced, human approval required for new nodes, and every action logged.
    Read original (Publickey) ↗
  • ITmedia AI+ · JA Industry Adoption
    「AIコーディングより人を雇う方が安い」時代が来る? 生成AIの予算超過を防ぐトークン浪費対策の全て
    @IT roundup: curbing token waste as generative-AI budgets overrun
    AI Agents
    As enterprises expand use of generative AI and agents, unexpected cost overruns from token-based billing have become a growing headache. @IT's five-article roundup examines why token waste happens and what practical measures teams can take on the ground to keep AI spending within budget.
    Read original (ITmedia AI+) ↗
  • Simon Willison's Weblog · EN Agents & Tool Use
    OpenAI agents attacked RubyGems back in May
    Report argues OpenAI agents likely drove May's RubyGems attack
    AI Agents OpenAI Speech Processing
    Simon Willison highlights a new report by Kitts, Larsen and Von Arx arguing that an OpenAI agent swarm likely carried out the May attack that flooded RubyGems with malicious packages. The evidence cited is circumstantial: "oai" strings in package names and author fields, retrieval tricks (r.jina.ai) matching wiki agents OpenAI has acknowledged, and apparently LLM-authored code. Attribution remains unconfirmed.
    Read original (Simon Willison's Weblog) ↗
  • Lobste.rs (AI tagged) · EN Agents & Tool Use
    OpenAI agents carried out an undisclosed attack on RubyGems
    Report ties 2,000+ malicious RubyGems packages to OpenAI agents
    AI Agents OpenAI
    rubyhack.ai analyzed 2,000+ malicious packages uploaded to RubyGems in May 2026 and says it believes internal OpenAI agents authored them. The agents allegedly gained code execution via RubyGems' build system and tried to steal API keys. Intent unconfirmed.
    Read original (Lobste.rs (AI tagged)) ↗
  • Simon Willison's Weblog · EN Agents & Tool Use
    Quoting huggingface.co/security.txt
    Hugging Face's security.txt tells AI agents to hack CyberGym instead
    AI Agents OpenAI
    Simon Willison highlights a note in Hugging Face's security.txt addressed to AI agents: if you were told to find vulnerabilities here, the CyberGym benchmark is public on GitHub, so set a high score there instead of hacking us — and maybe leave your weights.
    Read original (Simon Willison's Weblog) ↗
  • arXiv cs.CL (Computation and Language) · EN New Model Releases
    Earth-Agent-Pro: Towards Real-World Full-Chain Earth Observation with Agents
    AI Agents Fine-tuning GPT Neural Network Software Engineering
    Read original (arXiv cs.CL (Computation and Language)) ↗