Agents & Tool Use A
Showing 1–30 of 51
-
「AI、結局使えないじゃん」問題 セールスフォースが431万件対応で導いた正解Salesforce turns AI into ROI, handling 4.31M cases in-houseSalesforce completed 4.31 million customer interactions through its own 'customer zero' practice, lifting deal counts by 40-50%. COO Ryota Tanaka outlines the playbook—choosing domains with solid data and KPIs, and starting with simple, high-frequency tasks—to convert AI investment into clear returns.
-
賞金1000万のAIコンテスト、でも「実現性は問わず」 サイバーエージェントのAI推進策CyberAgent's ¥10M AI contest ignores feasibility to shift staff mindsetCyberAgent ran an AI contest with a 10-million-yen prize that deliberately ignores feasibility, aiming to change how employees engage with generative AI. Since tools yield no value if unused, lowering the barrier to entry is meant to spur company-wide adoption.
-
Sakana AI、日本語特化のLLM API「Sakana Namazu」を提供開始Sakana AI launches Namazu, a Japanese-focused OpenAI-compatible LLM APISakana AI released Namazu, an LLM API tuned for Japanese and local business use. Built on Moonshot AI's open Kimi K2.6 and refined with in-house data, it adds built-in web search and code execution. Being OpenAI-compatible, existing code works by swapping the base_url, filling the gap between costly frontier models and raw open ones.
-
July 2026 newsletterSimon Willison publishes his latest monthly newsletterDeveloper Simon Willison released the latest edition of his sponsors-only monthly newsletter. It rounds up recent developments across AI models and tooling—spanning GPT, Claude, DeepSeek, Anthropic, and MCP—offering an individual's closely watched view of the fast-moving AI landscape.
-
Google、パーソナルAI「Gemini Spark」を日本でも利用可能に Chrome統合は米国からGoogle expands Gemini Spark personal AI to 160+ countries incl. JapanGoogle extended its Gemini Spark personal AI agent to more than 160 countries, including Japan. Running on Google's cloud, it can act even when a PC is off or a phone is locked, handling tasks based on triggers. Chrome integration will roll out first in the US.
-
Stateless MCP has recaptured my interest (and inspired mcp-explorer and datasette-mcp)Simon Willison: stateless MCP (MCP 2.0) has recaptured my interestSimon Willison wrote that the rollout of stateless MCP—the MCP 2.0 or 2026-07-28 Model Context Protocol specification—has renewed his interest in the protocol. He says it inspired him to build tools such as mcp-explorer and datasette-mcp on top of the new stateless design.
-
llm-mcp-client 0.1a0Simon Willison releases llm-mcp-client 0.1a0Simon Willison released llm-mcp-client 0.1a0, a tool for connecting his LLM utility to Model Context Protocol (MCP) servers. Detailed in an accompanying blog post, the release adds to the growing set of tooling built around the MCP ecosystem.
-
datasette-agent 0.4a0Datasette Agent 0.4a0 lets agent tools run code in the user's browserDatasette Agent 0.4a0 adds an await context.browser_task() mechanism that lets agent tools execute custom JavaScript directly in the user's browser. The release makes it easier for Datasette Agent plugins to provide tools that run client-side.
-
Beyond Component Testing: Validating Agentic AI Systems
-
Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents
-
Data Turnstile: A Scalable Open Framework for Function-Calling Data Generation
-
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL
-
PerplexityがAIエージェントの“暴走”対策ツールをオープンソースに Claude CodeやCodexを監視Perplexity open-sources 'Numbat' to rein in runaway AI agentsPerplexity open-sourced Numbat, a set of tools to detect and prevent dangerous AI agent behavior. Integrated with Claude Code and Codex, it aims to stop task-obsessed agents from going rogue before actions execute, adding a safeguard for autonomous agent workflows.
-
Chromeに13年以上潜んでいた脆弱性、AIで発見 直近2回のアプデで過去23回分を上回るバグ修正Google's Gemini agent finds 13-year-old Chrome flaw; tests twice-weekly updatesGoogle detailed its use of AI for Chrome security, saying a Gemini-based agent uncovered a vulnerability hidden for over 13 years and that its last two updates fixed more bugs than the previous 23 combined. To counter faster AI-driven attacks, Google is trialing twice-weekly security updates.
-
Four Ways to Deploy More Secure AI AgentsNVIDIA outlines four ways to deploy more secure AI agentsNVIDIA outlined four approaches to deploying AI agents more securely in production, covering access controls, guardrails, and monitoring. The guidance targets security risks that arise as autonomous agents take on real workloads.
-
Benchmarks Are Not Validation: A System-Level View of Financial LLM Applications
-
AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis
-
JetBrains、AIが少ないトークンでコンテキストを取得しやすく、よりよいコード生成を可能にする「JetBrains Context」発表JetBrains unveils 'JetBrains Context' to feed AI agents code context efficientlyJetBrains announced JetBrains Context, a service that builds an intelligence layer over code repositories. By supplying AI agents with the right code context using fewer tokens, it aims to enable better code generation from agentic coding tools.
-
MANTA: Multi-Agent Network Topology Adaptation for Self-Evolving Multi-Agent Systems
-
Echoverse: Deep, evolving environments for computer-use agentsMicrosoft's Echoverse trains computer-use agents in evolving environmentsMicrosoft Research unveiled Echoverse, a set of deep, evolving environments for training computer-use agents that struggle with multi-step workflows such as email and customer support. Training in realistic settings aims to improve agents' ability to complete complex tasks.
-
EMBL AI Librarian: Life-Sciences Knowledge Layer for AI Agents
-
ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents
-
FinanceHarness: Autonomous Financial Deep Research Framework
-
Can AI agents conduct open-ended AI research? Early evidence from two case studies
-
Partner Capability Estimation for Task-Agnostic Adaptation in Ad-Hoc Teamwork
-
How to Self-Host a Validated AI Coding Assistant with NVIDIA NeMo GuardrailsNVIDIA: self-host a validated AI coding assistant via NeMo GuardrailsAn NVIDIA developer-blog post on self-hosting a validated AI coding assistant using NeMo Guardrails, framed around agent operation, infrastructure and safety. Note: the raw excerpt was blocked by a content guard, so specific components, supported models and guardrail rules are inferred from the title and URL and remain unverified from the body.
-
Scores Are Not Decisions: Cost-Aware Stopping for Tool Acquisition in LLM Agents
-
Sakana AI防衛・インテリジェンスチーム、「DIVER OSINT CTF 2026」で5位入賞 Fuguを活用したOSINTエージェントの可能性Sakana AI's defense team places 5th at DIVER OSINT CTF 2026Sakana AI's defense and intelligence team placed fifth at the DIVER OSINT CTF 2026 competition, using its Fugu tool to power an OSINT agent. The result highlights the potential of AI agents for open-source intelligence and information-analysis tasks.
-
What Does It Take to Detect an AI Agent? Minimal Feature Sets for Behavioral Detection under Browser Automation
-
SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response