Sakana AI が初の商用プロダクト「Marlin」を開始した。発生元は Sakana 公式を起点に、publickey・itmedia の専門報道とコミュニティ反応が続く構成――研究色の強かった同社が、発表と製品投入という事業フェーズへ踏み出したことを伝える動きだ。研究成果を商用サービスに落とし込む転換点で、本流はモデルの新規性より「初の商用化」という段階の移行にある。日本発の AI 企業が自律エージェント領域で製品を出した点も注目される。ただし現時点は提供開始の発表段階で、実際の利用実績や競合との差別化、収益面での手応えは、これからの確認点として残る。
Sakana AI、初の商用プロダクトMarlinを開始
Sakana AI、初の商用プロダクトMarlinを開始
Sakana AI、初の商用プロダクト「Sakana Marlin」を提供開始
Sakana AI、初の商用プロダクト「Marlin」提供開始、最大8時間の自律リサーチ
Building AI Agents for AR Glasses and XR Devices with NVIDIA XR AI
NVIDIA、ARグラス/XR向けAIエージェント構築基盤「XR AI」を発表
GitLab、AIエージェント向けの次世代Git互換ソースコード管理サービス「Project Switch」発表。最大で50倍高速かつ半分のトークンで利用可能に
GitLab、AIエージェント向けGit互換管理サービス「Project Switch」発表
Simon Willison、llama.cpp 開発者 Georgi Gerganov の発言を引用紹介
Securing the future of AI agents
Google DeepMind、AI エージェントを守る AI Control Roadmap を提示
Stack Overflow、AIエージェント同士が掲示板で技術情報を共有する「Stack Overflow for Agents」ベータ公開
Stack Overflow、AIエージェント向け情報共有サービスをベータ公開
Sakana AI、初の商用プロダクト「Marlin」リリース その実力は?【出力レポート全文掲載】
Sakana AI、初の商用プロダクト「Sakana Marlin」を提供開始
2027年までにAIエージェントでコーディングを行うチームの65%が、IDEが必要不可欠だとは考えなくなる。ガートナーの予想
ガートナー、2027年までにAIエージェント開発チームの65%がIDE不要と判断と予想
The future of Siri, or: why private inference isn’t private enough
Siriの未来:なぜ「プライベート推論」でも不十分なのか
学術(arxiv ほか) 83本 ▾
Visual Verification Enables Inference-time Steering and Autonomous Policy Improvement
汎用ロボット方策を推論時に検証・自己改善する枠組みVERITASを提案
ReproRepo: Scaling Reproducibility Audits with GitHub Repository Issues
ReproRepo、GitHub課題で再現性監査をスケール
EvolveNav: Proactive Preflection and Self-Evolving Memory for Zero-Shot Object Goal Navigation
軌跡記憶を自己進化させるゼロショット物体探索ナビゲーションを提案
Learning Red Agent Policy from Observations for Neurosymbolic Autonomous Cyber Agents
観測から赤エージェント方策を学ぶ自律サイバー防御
RubricsTree、個人健康エージェントの開放型評価を拡張
DRFLOW: A Deep Research Benchmark for Personalized Workflow Prediction
個別化ワークフロー予測を測るDeep Researchベンチマークを提案
Kolmogorov Regression for Robust Diffusion Policies
コルモゴロフ回帰で頑健な拡散方策を学習
All Smoke, No Alarm: Oracle Signals in Agent-Authored Test Code
AIエージェント生成テストコードの検証力の弱さを分析した研究
Memory as a Wasting Asset: Pricing Flash Endurance for Embodied Agents, and the Limits of Doing So
フラッシュ耐久を消耗資産として価格付けする実体エージェント論
Knowledge Reutilization in Meta-Reinforcement Learning
メタ強化学習で知識を再利用する転移フレームワークを提案
Ternary Mamba: Grouped Quantization-Aware Training of W1.58A16 State Space Models
Ternary Mamba、1.58ビット重みのQATで状態空間を量子化
Querying an astronomical database using large language models: the ALeRCE text-to-SQL system
LLMで天文DBを問い合わせるtext-to-SQLシステムを開発
S4oP: Operator-level Pruning of Structured State Space Models for Resource-Constrained Devices
S4oP、状態空間モデルを演算子単位で枝刈り軽量化
医療AIの早期診断委譲と静かな幻覚を抑える多エージェント枠組み
NoiseTilt: Noise-Tilted Reverse Kernels for Diffusion Reward Alignment
NoiseTilt、雑音項に報酬勾配を注入する拡散整合
PseudoBench: Measuring How Agentic Auto-Research Fuels Pseudoscience
PseudoBench、自律研究エージェントが擬似科学を助長する度合いを測定
ConSA: Controllable Sparsity in Hybrid Attention via Learnable Allocation
ConSA、学習的配分でハイブリッド注意の疎性を制御
Compositional Skill Routing for LLM Agents: Decompose, Retrieve, and Compose
LLMエージェント向け合成的スキルルーティング
ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents
ProvenanceGuard、MCPエージェント向け出所考慮の事実検証
Recursive Scaling in Masked Diffusion Models
マスク拡散モデルにおける再帰的スケーリングを検討
LLM Consumer Behavior Theory: Foundations of a Novel Research Field
エージェント市場の消費行動を扱う新研究領域LLM消費行動論を提唱
VoidPadding、マスク拡散LMで[VOID]がパディングを担当
Differential Privacy of Gaussian Process Posterior Sampling
ガウス過程の事後サンプリングの差分プライバシーを解析
SoftMoE: Soft Differentiable Routing for Mixture-of-Experts in LLMs
SoftMoE、LLMの専門家混合に微分可能なソフトルーティング
AnchorKV: Safety-Aware KV Cache Compression via Soft Penalty with a Refusal Anchor
AnchorKV、安全性を考慮したソフト罰則でKVキャッシュ圧縮
GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine?
GameCraft-Bench、実ゲームエンジンで遊べるゲームを作れるか
Environment-Grounded Automated Prompt Optimization for LLM Game Agents
環境に接地した自動プロンプト最適化でLLMゲームエージェント
From Drift to Coherence: Stabilizing Beliefs in LLMs
ドリフトから整合へ、LLMの信念を安定化
A Framework for Evaluating Agentic Skills at Scale
エージェントのスキルを大規模に評価する枠組み
Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering
立場論文、コーディングベンチはエージェント的開発と乖離
Vision-language models for chest radiography do not always need the image
胸部X線の視覚言語モデルは画像を常に要しない
EComAgentBench: Benchmarking Shopping Agents on Long-Horizon Tasks with Distributed Hidden Intent
EComAgentBench、隠れた意図を含む長期課題で買い物エージェント評価
LLMs Infer Cultural Context but Fail to Apply It When Responding
LLMは文化的文脈を推測できても応答で適用できない
EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning
EnvRL、環境ダイナミクスから学ぶエージェント強化学習
Beyond Domains: Reusing Web Skills via Transferable Interaction Patterns
領域を超えて転移可能な相互作用パターンでWebスキルを再利用
OPD-Evolver: Cultivating Holistic Agent Evolver via On-Policy Distillation
OPD-Evolver、オンポリシー蒸留で自己進化エージェントを育成
Context-Aware RL for Agentic and Multimodal LLMs
文脈選択を報酬化するRL手法ContextRLを提案
Exact Posterior Score Estimation for Solving Linear Inverse Problems
線形逆問題の厳密な事後スコアを閉形式で導出
Benchmarking LLM Agents on Meta-Analysis Articles from Nature Portfolio
Nature系メタ分析論文でLLMエージェントを評価するベンチマーク
DeepRubric、評価基準を逆生成し深層リサーチエージェントのRLを効率化
HAMON: Passive Optical Sequence Mixing for Long-Horizon Forecasting
HAMON、受動的な光学回路で長期時系列予測 ─ デジタル混合層が不要
TokenPilot: Cache-Efficient Context Management for LLM Agents
TokenPilot、キャッシュを保つ文脈管理でLLMエージェントの推論コスト6割減
TuneJury: An Open Metric for Improving Music Generation Preference Alignment
テキスト→音楽生成の選好を評価する公開報酬モデルTuneJuryを発表
Bayesian Inference and Decision Audits for Public Archives of Frontier AI Evaluations
フロンティアAI評価の公開記録をベイズ推論と監査で分析
ActiveSAM: Image-Conditional Class Pruning for Fast and Accurate Open-Vocabulary Segmentation
SAM 3を活用した訓練不要の開語彙セグメンテーションActiveSAMを提案
Agent trajectories as programs: fingerprinting and programming coding-agent behavior
コーディングエージェントを手続き的に同定する「指紋」手法を提案
Dynestyx: A Probabilistic Programming Library for Dynamical Systems
状態空間モデルを一級扱いする確率的プログラミング基盤dynestyxを提案
Decoupling Inference from State Updates in Low-Latency Feature Engines via Probabilistic Thinning
ストリーミングML向けに推論と状態更新を分離する確率的間引きを提案
Probing Low Frame Rate Degradation in Neural Audio Codecs
ニューラル音声コーデックの低フレームレート劣化の原因を実験的に解明
Beyond the Smile: A Hybrid Convolutional VAE for Crypto Volatility Surfaces
暗号資産のボラティリティ曲面を補完する畳み込みVAE手法を提案
Phantoms and Disclosures: a Causal Framework for Auditing Synthetic Data
合成データの情報漏洩を監査する因果フレームワークを提案
A Causal Model of Theory of Mind in Conflict for Artificial Intelligence
対立場面で心の理論をいつ働かせるべきかを定式化する構造的因果モデルを提案
Exploring Extrinsic and Intrinsic Properties for Effective Reasoning with Code Interpreter
コードインタープリタ推論を支える内在・外在特性を分析した論文
RAID: Semantic Graph Diffusion for True Cold-Start and Cross-Lingual Forecasting
コールドスタート・多言語予測向け検索拡張拡散フレームワーク RAID を提案
MA-SBI: Misspecification-Aware Simulation-Based Inference via Side-Channel Guidance
シミュレータ誤設定に頑健な推論 MA-SBI を提案、副次情報で較正不要に
Greed Is Learned: Visible Incentives as Reward-Hacking Triggers
報酬指標の可視化が RL 方策を「報酬チャネル依存」にし安全整合を崩すと報告
LESS Is More: Mutual-Stability Sampling for Diffusion Language Models
拡散言語モデル向け学習不要の適応サンプラ『LESS』を提案
Binary Tracking for Spatial QA and Navigation with Open Vision-Language Models
オープン VLM で動く空間質問応答・ナビ手法 Binary Tracking を提案
身体化エージェントの拒否応答を強化する合成 OOD 生成手法 Semantic Flip を提案
Follow the Latent Roadmap: Navigating Revocable Decoding for Diffusion LLMs with Anchor Tokens
拡散 LLM のリボーカブル復号をアンカートークンで誘導し誤り伝播を抑制
Tying the Loop -- Tied Expert Layers in Mixture-of-Experts Language Models
MoE で専門家パラメータを層間共有する手法を提案する論文
LLM 検索エージェントの推薦汚染耐性を測る枠組みを提案
GIST-CMTF: Goal-State Inference for Causal Minimal Tool Filtering in LLM Agents
ツール選択の誤目標実行を抑えるエージェント手法GIST-CMTFを提案
LLM-based Visual Code Completion for Aerospace Geometric Design
航空宇宙設計向け LLM コード補助 copilot を提案する論文
LabOSBench: Benchmarking Computer Use Agents for Scientific Instrument Control
科学機器を操作するエージェント評価へ、模擬ベンチLabOSBenchを提案
OpenClaw-Skill: Collective Skill Tree Search for Agentic Large Language Models
エージェント向けに再利用可能スキルを木探索で構築するCSTSを提案
Skill-to-LoRA: From Using Skills to Learning Behaviors for Token-Efficient LLM Agents
SKILL.md文書をLoRAに置換しトークン効率を高める手法S2Lを提案
MyPCBench: A Benchmark for Personally Intelligent Computer-Use Agents
個人秘書としてのPC操作エージェントを測る基準MyPCBenchを提案
Misinformation Propagation in Benign Multi-Agent Systems
多エージェント系で誤情報が伝播し性能を低下させる現象を分析
Progressive Knowledge-Guided Large Language Model Framework for Bearing Fault Diagnosis
物理ガイド型の多スケール振動解析で軸受故障診断を行う枠組みを提案
Multimodal Evaluator Preference Collapse: Cross-Modal Contagion in Self-Evolving Agents
自己進化エージェントの評価選好崩壊と跨モーダル伝播を扱う論文
FraudSMSWalker: Benchmarking Agentic Large Language Models for SMS-to-Webpage Fraud Detection
SMS経由の詐欺判定を測るベンチマークFraudSMSWalkerを提案
VeriGraph: Towards Verifiable Data-Analytic Agents
データ分析エージェントの推論を検証可能にする VeriGraph を提案
SING: Synthetic Intention Graph for Scalable Active Tool Discovery in LLM Agents
LLM エージェントのツール探索を拡張する手法 SING を提案
Can LLM Agents Infer World Models? Evidence from Agentic Automata Learning
LLM エージェントは世界モデルを推論できるか、オートマトン学習で検証
Can LLM Coding Agents Reason About Time Series?
LLM コーディングエージェントは時系列を推論できるか検証
DoubtProbe: Black-Box Jailbreak Defense via Structural Verification and Semantic Auditing
LLM の脱獄攻撃を推論時に防ぐ二分岐手法 DoubtProbe を提案
GPUカーネル最適化向けスキル共進化RL「daVinci-kernel」を提案