AI エージェント × 安全性・評価

Sakana AI、初の商用プロダクトMarlinを開始

Sakana AI、初の商用プロダクトMarlinを開始

✎ ストーリー本文

Sakana AI が初の商用プロダクト「Marlin」を開始した。発生元は Sakana 公式を起点に、publickey・itmedia の専門報道とコミュニティ反応が続く構成――研究色の強かった同社が、発表と製品投入という事業フェーズへ踏み出したことを伝える動きだ。研究成果を商用サービスに落とし込む転換点で、本流はモデルの新規性より「初の商用化」という段階の移行にある。日本発の AI 企業が自律エージェント領域で製品を出した点も注目される。ただし現時点は提供開始の発表段階で、実際の利用実績や競合との差別化、収益面での手応えは、これからの確認点として残る。

▲ 公式・報道
公式

Sakana AI、初の商用プロダクト「Sakana Marlin」を提供開始

Sakana AI Blog (ja) ・ 2026-06-14 ・ 📌

Sakana AI、初の商用プロダクト「Marlin」提供開始、最大8時間の自律リサーチ

公式

Building AI Agents for AR Glasses and XR Devices with NVIDIA XR AI

NVIDIA Developer Blog ・ 2026-06-16

NVIDIA、ARグラス/XR向けAIエージェント構築基盤「XR AI」を発表

報道

GitLab、AIエージェント向けの次世代Git互換ソースコード管理サービス「Project Switch」発表。最大で50倍高速かつ半分のトークンで利用可能に

Publickey ・ 2026-06-16

GitLab、AIエージェント向けGit互換管理サービス「Project Switch」発表

コミュニティ

Quoting Georgi Gerganov

Simon Willison's Weblog ・ 2026-06-16

Simon Willison、llama.cpp 開発者 Georgi Gerganov の発言を引用紹介

公式

Securing the future of AI agents

Google DeepMind Blog ・ 2026-06-16

Google DeepMind、AI エージェントを守る AI Control Roadmap を提示

報道

Stack Overflow、AIエージェント同士が掲示板で技術情報を共有する「Stack Overflow for Agents」ベータ公開

Publickey ・ 2026-06-15

Stack Overflow、AIエージェント向け情報共有サービスをベータ公開

報道

Sakana AI、初の商用プロダクト「Marlin」リリース その実力は?【出力レポート全文掲載】

ITmedia AI+ ・ 2026-06-15

Sakana AI、初の商用プロダクト「Sakana Marlin」を提供開始

報道

2027年までにAIエージェントでコーディングを行うチームの65%が、IDEが必要不可欠だとは考えなくなる。ガートナーの予想

Publickey ・ 2026-06-14

ガートナー、2027年までにAIエージェント開発チームの65%がIDE不要と判断と予想

コミュニティ

The future of Siri, or: why private inference isn’t private enough

Lobste.rs (AI tagged) ・ 2026-06-14

Siriの未来:なぜ「プライベート推論」でも不十分なのか

学術(arxiv ほか) 83本 ▾
学術

Visual Verification Enables Inference-time Steering and Autonomous Policy Improvement

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-16

汎用ロボット方策を推論時に検証・自己改善する枠組みVERITASを提案

学術

ReproRepo: Scaling Reproducibility Audits with GitHub Repository Issues

arXiv cs.CL (Computation and Language) ・ 2026-06-16

ReproRepo、GitHub課題で再現性監査をスケール

学術

EvolveNav: Proactive Preflection and Self-Evolving Memory for Zero-Shot Object Goal Navigation

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-16

軌跡記憶を自己進化させるゼロショット物体探索ナビゲーションを提案

学術

Learning Red Agent Policy from Observations for Neurosymbolic Autonomous Cyber Agents

arXiv cs.LG (Machine Learning) ・ 2026-06-16

観測から赤エージェント方策を学ぶ自律サイバー防御

学術

RubricsTree: Scalable and Evolving Open-Ended Evaluation of Personal Health Agents across Health Memory and Medical Skills

arXiv cs.CL (Computation and Language) ・ 2026-06-16

RubricsTree、個人健康エージェントの開放型評価を拡張

学術

DRFLOW: A Deep Research Benchmark for Personalized Workflow Prediction

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-16

個別化ワークフロー予測を測るDeep Researchベンチマークを提案

学術

Kolmogorov Regression for Robust Diffusion Policies

arXiv cs.LG (Machine Learning) ・ 2026-06-16

コルモゴロフ回帰で頑健な拡散方策を学習

学術

All Smoke, No Alarm: Oracle Signals in Agent-Authored Test Code

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-16

AIエージェント生成テストコードの検証力の弱さを分析した研究

学術

Memory as a Wasting Asset: Pricing Flash Endurance for Embodied Agents, and the Limits of Doing So

arXiv cs.LG (Machine Learning) ・ 2026-06-16

フラッシュ耐久を消耗資産として価格付けする実体エージェント論

学術

Your AI Travel Agent Would Book You a Bullfight: An Agentic Benchmark for Implicit Animal Welfare in Frontier AI Models

arXiv cs.CL (Computation and Language) ・ 2026-06-16

動物福祉の暗黙的配慮を測るエージェント型ベンチマーク

学術

Knowledge Reutilization in Meta-Reinforcement Learning

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-16

メタ強化学習で知識を再利用する転移フレームワークを提案

学術

Embedded Machine Learning for Microcontroller-Class Edge Devices: Data, Feature, Evaluation, and Deployment Pipelines

arXiv cs.LG (Machine Learning) ・ 2026-06-16

マイコン級エッジ向け組込み機械学習のパイプライン総説

学術

Ternary Mamba: Grouped Quantization-Aware Training of W1.58A16 State Space Models

arXiv cs.LG (Machine Learning) ・ 2026-06-16

Ternary Mamba、1.58ビット重みのQATで状態空間を量子化

学術

Querying an astronomical database using large language models: the ALeRCE text-to-SQL system

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-16

LLMで天文DBを問い合わせるtext-to-SQLシステムを開発

学術

S4oP: Operator-level Pruning of Structured State Space Models for Resource-Constrained Devices

arXiv cs.LG (Machine Learning) ・ 2026-06-16

S4oP、状態空間モデルを演算子単位で枝刈り軽量化

学術

Agentic AI-based Framework for Mitigating Premature Diagnostic Handoff and Silent Hallucination in Healthcare Applications

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-16

医療AIの早期診断委譲と静かな幻覚を抑える多エージェント枠組み

学術

NoiseTilt: Noise-Tilted Reverse Kernels for Diffusion Reward Alignment

arXiv cs.LG (Machine Learning) ・ 2026-06-16

NoiseTilt、雑音項に報酬勾配を注入する拡散整合

学術

PseudoBench: Measuring How Agentic Auto-Research Fuels Pseudoscience

arXiv cs.CL (Computation and Language) ・ 2026-06-16

PseudoBench、自律研究エージェントが擬似科学を助長する度合いを測定

学術

ConSA: Controllable Sparsity in Hybrid Attention via Learnable Allocation

arXiv cs.CL (Computation and Language) ・ 2026-06-16

ConSA、学習的配分でハイブリッド注意の疎性を制御

学術

Compositional Skill Routing for LLM Agents: Decompose, Retrieve, and Compose

arXiv cs.CL (Computation and Language) ・ 2026-06-16

LLMエージェント向け合成的スキルルーティング

学術

ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-16

ProvenanceGuard、MCPエージェント向け出所考慮の事実検証

学術

Recursive Scaling in Masked Diffusion Models

arXiv cs.LG (Machine Learning) ・ 2026-06-16

マスク拡散モデルにおける再帰的スケーリングを検討

学術

LLM Consumer Behavior Theory: Foundations of a Novel Research Field

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-16

エージェント市場の消費行動を扱う新研究領域LLM消費行動論を提唱

学術

Half a Link can Be Enough to Predict a Whole Link: Understanding Generalization in Knowledge Graph Foundation Models

arXiv cs.LG (Machine Learning) ・ 2026-06-16

半分のリンクで全リンク予測、KG基盤モデルの汎化を解明

学術

VoidPadding: Let [VOID] Handle Padding in Masked Diffusion Language Models so that [EOS] Can Focus on Semantic Termination

arXiv cs.CL (Computation and Language) ・ 2026-06-16

VoidPadding、マスク拡散LMで[VOID]がパディングを担当

学術

Differential Privacy of Gaussian Process Posterior Sampling

arXiv cs.LG (Machine Learning) ・ 2026-06-16

ガウス過程の事後サンプリングの差分プライバシーを解析

学術

SoftMoE: Soft Differentiable Routing for Mixture-of-Experts in LLMs

arXiv cs.LG (Machine Learning) ・ 2026-06-16

SoftMoE、LLMの専門家混合に微分可能なソフトルーティング

学術

Revisiting Structural Dependency in Autoregressive Multi-Task Table Recognition via Order-Independent Cell-Level Representations

arXiv cs.LG (Machine Learning) ・ 2026-06-16

順序非依存のセル表現で自己回帰的表認識を再考

学術

AnchorKV: Safety-Aware KV Cache Compression via Soft Penalty with a Refusal Anchor

arXiv cs.LG (Machine Learning) ・ 2026-06-16

AnchorKV、安全性を考慮したソフト罰則でKVキャッシュ圧縮

学術

GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine?

arXiv cs.CL (Computation and Language) ・ 2026-06-16

GameCraft-Bench、実ゲームエンジンで遊べるゲームを作れるか

学術

Environment-Grounded Automated Prompt Optimization for LLM Game Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-16

環境に接地した自動プロンプト最適化でLLMゲームエージェント

学術

From Drift to Coherence: Stabilizing Beliefs in LLMs

arXiv cs.LG (Machine Learning) ・ 2026-06-16

ドリフトから整合へ、LLMの信念を安定化

学術

Improving low-resource ASR using bilingual fine-tuning with language identification: a cross-linguistic evaluation

arXiv cs.CL (Computation and Language) ・ 2026-06-16

言語識別付き二言語微調整で低資源ASRを改善

学術

A Framework for Evaluating Agentic Skills at Scale

arXiv cs.CL (Computation and Language) ・ 2026-06-16

エージェントのスキルを大規模に評価する枠組み

学術

Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering

arXiv cs.CL (Computation and Language) ・ 2026-06-16

立場論文、コーディングベンチはエージェント的開発と乖離

学術

Vision-language models for chest radiography do not always need the image

arXiv cs.CL (Computation and Language) ・ 2026-06-16

胸部X線の視覚言語モデルは画像を常に要しない

学術

EComAgentBench: Benchmarking Shopping Agents on Long-Horizon Tasks with Distributed Hidden Intent

arXiv cs.CL (Computation and Language) ・ 2026-06-16

EComAgentBench、隠れた意図を含む長期課題で買い物エージェント評価

学術

LLMs Infer Cultural Context but Fail to Apply It When Responding

arXiv cs.CL (Computation and Language) ・ 2026-06-16

LLMは文化的文脈を推測できても応答で適用できない

学術

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning

arXiv cs.CL (Computation and Language) ・ 2026-06-16

EnvRL、環境ダイナミクスから学ぶエージェント強化学習

学術

Beyond Domains: Reusing Web Skills via Transferable Interaction Patterns

arXiv cs.CL (Computation and Language) ・ 2026-06-16

領域を超えて転移可能な相互作用パターンでWebスキルを再利用

学術

OPD-Evolver: Cultivating Holistic Agent Evolver via On-Policy Distillation

arXiv cs.CL (Computation and Language) ・ 2026-06-16

OPD-Evolver、オンポリシー蒸留で自己進化エージェントを育成

学術

Context-Aware RL for Agentic and Multimodal LLMs

arXiv cs.CL (Computation and Language) ・ 2026-06-15

文脈選択を報酬化するRL手法ContextRLを提案

学術

Exact Posterior Score Estimation for Solving Linear Inverse Problems

arXiv cs.LG (Machine Learning) ・ 2026-06-15

線形逆問題の厳密な事後スコアを閉形式で導出

学術

Benchmarking LLM Agents on Meta-Analysis Articles from Nature Portfolio

arXiv cs.CL (Computation and Language) ・ 2026-06-15

Nature系メタ分析論文でLLMエージェントを評価するベンチマーク

学術

DEEPRUBRIC: Evidence-Tree Rubric Supervision for Efficient Reinforcement Learning of Deep Research Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-15

DeepRubric、評価基準を逆生成し深層リサーチエージェントのRLを効率化

学術

HAMON: Passive Optical Sequence Mixing for Long-Horizon Forecasting

arXiv cs.LG (Machine Learning) ・ 2026-06-15

HAMON、受動的な光学回路で長期時系列予測 ─ デジタル混合層が不要

学術

TokenPilot: Cache-Efficient Context Management for LLM Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-15

TokenPilot、キャッシュを保つ文脈管理でLLMエージェントの推論コスト6割減

学術

TuneJury: An Open Metric for Improving Music Generation Preference Alignment

arXiv cs.LG (Machine Learning) ・ 2026-06-15

テキスト→音楽生成の選好を評価する公開報酬モデルTuneJuryを発表

学術

Bayesian Inference and Decision Audits for Public Archives of Frontier AI Evaluations

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-15

フロンティアAI評価の公開記録をベイズ推論と監査で分析

学術

ActiveSAM: Image-Conditional Class Pruning for Fast and Accurate Open-Vocabulary Segmentation

arXiv cs.LG (Machine Learning) ・ 2026-06-15

SAM 3を活用した訓練不要の開語彙セグメンテーションActiveSAMを提案

学術

Agent trajectories as programs: fingerprinting and programming coding-agent behavior

arXiv cs.LG (Machine Learning) ・ 2026-06-15

コーディングエージェントを手続き的に同定する「指紋」手法を提案

学術

Dynestyx: A Probabilistic Programming Library for Dynamical Systems

arXiv cs.LG (Machine Learning) ・ 2026-06-15

状態空間モデルを一級扱いする確率的プログラミング基盤dynestyxを提案

学術

Decoupling Inference from State Updates in Low-Latency Feature Engines via Probabilistic Thinning

arXiv cs.LG (Machine Learning) ・ 2026-06-15

ストリーミングML向けに推論と状態更新を分離する確率的間引きを提案

学術

Probing Low Frame Rate Degradation in Neural Audio Codecs

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-15

ニューラル音声コーデックの低フレームレート劣化の原因を実験的に解明

学術

Beyond the Smile: A Hybrid Convolutional VAE for Crypto Volatility Surfaces

arXiv cs.LG (Machine Learning) ・ 2026-06-15

暗号資産のボラティリティ曲面を補完する畳み込みVAE手法を提案

学術

Phantoms and Disclosures: a Causal Framework for Auditing Synthetic Data

arXiv cs.LG (Machine Learning) ・ 2026-06-15

合成データの情報漏洩を監査する因果フレームワークを提案

学術

A Causal Model of Theory of Mind in Conflict for Artificial Intelligence

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-15

対立場面で心の理論をいつ働かせるべきかを定式化する構造的因果モデルを提案

学術

Exploring Extrinsic and Intrinsic Properties for Effective Reasoning with Code Interpreter

arXiv cs.CL (Computation and Language) ・ 2026-06-15

コードインタープリタ推論を支える内在・外在特性を分析した論文

学術

RAID: Semantic Graph Diffusion for True Cold-Start and Cross-Lingual Forecasting

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-15

コールドスタート・多言語予測向け検索拡張拡散フレームワーク RAID を提案

学術

MA-SBI: Misspecification-Aware Simulation-Based Inference via Side-Channel Guidance

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-15

シミュレータ誤設定に頑健な推論 MA-SBI を提案、副次情報で較正不要に

学術

Greed Is Learned: Visible Incentives as Reward-Hacking Triggers

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-15

報酬指標の可視化が RL 方策を「報酬チャネル依存」にし安全整合を崩すと報告

学術

LESS Is More: Mutual-Stability Sampling for Diffusion Language Models

arXiv cs.CL (Computation and Language) ・ 2026-06-15

拡散言語モデル向け学習不要の適応サンプラ『LESS』を提案

学術

Binary Tracking for Spatial QA and Navigation with Open Vision-Language Models

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-15

オープン VLM で動く空間質問応答・ナビ手法 Binary Tracking を提案

学術

Semantic Flip: Synthetic OOD Generation for Robust Refusal in Embodied Question Answering and Spatial Localization

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-15

身体化エージェントの拒否応答を強化する合成 OOD 生成手法 Semantic Flip を提案

学術

Follow the Latent Roadmap: Navigating Revocable Decoding for Diffusion LLMs with Anchor Tokens

arXiv cs.CL (Computation and Language) ・ 2026-06-15

拡散 LLM のリボーカブル復号をアンカートークンで誘導し誤り伝播を抑制

学術

Tying the Loop -- Tied Expert Layers in Mixture-of-Experts Language Models

arXiv cs.CL (Computation and Language) ・ 2026-06-15

MoE で専門家パラメータを層間共有する手法を提案する論文

学術

How Much Can We Trust LLM Search Agents? Measuring Endorsement Vulnerability to Web Content Manipulation

arXiv cs.CL (Computation and Language) ・ 2026-06-15

LLM 検索エージェントの推薦汚染耐性を測る枠組みを提案

学術

GIST-CMTF: Goal-State Inference for Causal Minimal Tool Filtering in LLM Agents

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-15

ツール選択の誤目標実行を抑えるエージェント手法GIST-CMTFを提案

学術

LLM-based Visual Code Completion for Aerospace Geometric Design

arXiv cs.CL (Computation and Language) ・ 2026-06-15

航空宇宙設計向け LLM コード補助 copilot を提案する論文

学術

LabOSBench: Benchmarking Computer Use Agents for Scientific Instrument Control

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-15

科学機器を操作するエージェント評価へ、模擬ベンチLabOSBenchを提案

学術

OpenClaw-Skill: Collective Skill Tree Search for Agentic Large Language Models

arXiv cs.CL (Computation and Language) ・ 2026-06-15

エージェント向けに再利用可能スキルを木探索で構築するCSTSを提案

学術

Skill-to-LoRA: From Using Skills to Learning Behaviors for Token-Efficient LLM Agents

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-15

SKILL.md文書をLoRAに置換しトークン効率を高める手法S2Lを提案

学術

MyPCBench: A Benchmark for Personally Intelligent Computer-Use Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-15

個人秘書としてのPC操作エージェントを測る基準MyPCBenchを提案

学術

Misinformation Propagation in Benign Multi-Agent Systems

arXiv cs.CL (Computation and Language) ・ 2026-06-15

多エージェント系で誤情報が伝播し性能を低下させる現象を分析

学術

Progressive Knowledge-Guided Large Language Model Framework for Bearing Fault Diagnosis

arXiv cs.CL (Computation and Language) ・ 2026-06-15

物理ガイド型の多スケール振動解析で軸受故障診断を行う枠組みを提案

学術

Multimodal Evaluator Preference Collapse: Cross-Modal Contagion in Self-Evolving Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-15

自己進化エージェントの評価選好崩壊と跨モーダル伝播を扱う論文

学術

FraudSMSWalker: Benchmarking Agentic Large Language Models for SMS-to-Webpage Fraud Detection

arXiv cs.CL (Computation and Language) ・ 2026-06-15

SMS経由の詐欺判定を測るベンチマークFraudSMSWalkerを提案

学術

VeriGraph: Towards Verifiable Data-Analytic Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-15

データ分析エージェントの推論を検証可能にする VeriGraph を提案

学術

SING: Synthetic Intention Graph for Scalable Active Tool Discovery in LLM Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-15

LLM エージェントのツール探索を拡張する手法 SING を提案

学術

Can LLM Agents Infer World Models? Evidence from Agentic Automata Learning

arXiv cs.CL (Computation and Language) ・ 2026-06-15

LLM エージェントは世界モデルを推論できるか、オートマトン学習で検証

学術

Can LLM Coding Agents Reason About Time Series?

arXiv cs.CL (Computation and Language) ・ 2026-06-15

LLM コーディングエージェントは時系列を推論できるか検証

学術

DoubtProbe: Black-Box Jailbreak Defense via Structural Verification and Semantic Auditing

arXiv cs.CL (Computation and Language) ・ 2026-06-15

LLM の脱獄攻撃を推論時に防ぐ二分岐手法 DoubtProbe を提案

学術

daVinci-kernel: Co-Evolving Skill Selection, Summarization, and Utilization via RL for GPU Kernel Optimization

arXiv cs.CL (Computation and Language) ・ 2026-06-15

GPUカーネル最適化向けスキル共進化RL「daVinci-kernel」を提案

← ストーリー アーカイブ