Sakana AI launched Marlin, its first commercial product. The signal starts from Sakana's official announcement, followed by trade press (Publickey, ITmedia) and community reaction—marking a research-heavy company stepping into a business phase of shipping a product. It's a turning point from research results into a commercial service; the through-line is less model novelty than the stage shift to 'first commercialization.' A Japan-based AI company shipping in the autonomous-agent space is also notable. But this is the availability-announcement stage: real usage, differentiation, and revenue traction remain to be confirmed.
Sakana AI launches Marlin product
Sakana AI launches Marlin product
Sakana AI、初の商用プロダクト「Sakana Marlin」を提供開始
Sakana AI launches Marlin, its first commercial autonomous research assistant
Building AI Agents for AR Glasses and XR Devices with NVIDIA XR AI
NVIDIA unveils XR AI to build AI agents for AR glasses and XR devices
GitLab、AIエージェント向けの次世代Git互換ソースコード管理サービス「Project Switch」発表。最大で50倍高速かつ半分のトークンで利用可能に
GitLab unveils 'Project Switch,' a Git-compatible SCM service for AI agents
Simon Willison quotes Georgi Gerganov (llama.cpp / ggml author)
Securing the future of AI agents
DeepMind outlines an AI Control Roadmap to secure AI agents
Stack Overflow、AIエージェント同士が掲示板で技術情報を共有する「Stack Overflow for Agents」ベータ公開
Stack Overflow launches 'Stack Overflow for Agents' beta
Sakana AI、初の商用プロダクト「Marlin」リリース その実力は?【出力レポート全文掲載】
Sakana AI launches its first commercial product, Sakana Marlin
2027年までにAIエージェントでコーディングを行うチームの65%が、IDEが必要不可欠だとは考えなくなる。ガートナーの予想
Gartner: by 2027, 65% of AI-coding teams find IDEs non-essential
The future of Siri, or: why private inference isn’t private enough
The future of Siri: why private inference isn't private enough
Academic (arxiv etc.) 83 ▾
Visual Verification Enables Inference-time Steering and Autonomous Policy Improvement
VERITAS steers and self-improves robot policies at inference time
ReproRepo: Scaling Reproducibility Audits with GitHub Repository Issues
ReproRepo scales reproducibility audits using GitHub repo issues
EvolveNav: Proactive Preflection and Self-Evolving Memory for Zero-Shot Object Goal Navigation
EvolveNav: a self-evolving framework for zero-shot object-goal navigation
Learning Red Agent Policy from Observations for Neurosymbolic Autonomous Cyber Agents
Learning red-agent policy from observations for cyber-defense RL
RubricsTree: scalable open-ended evaluation of personal health agents
DRFLOW: A Deep Research Benchmark for Personalized Workflow Prediction
DRFLOW: a deep research benchmark for personalized workflow prediction
Kolmogorov Regression for Robust Diffusion Policies
Kolmogorov regression yields robust diffusion policies
All Smoke, No Alarm: Oracle Signals in Agent-Authored Test Code
Study finds agent-authored test code often lacks real verification logic
Memory as a Wasting Asset: Pricing Flash Endurance for Embodied Agents, and the Limits of Doing So
Pricing flash endurance as a wasting asset for embodied agents
An agentic benchmark for implicit animal welfare in frontier AI
Knowledge Reutilization in Meta-Reinforcement Learning
A meta-knowledge reutilization framework for meta-RL across agents
A pipeline survey of embedded ML for microcontroller-class devices
Ternary Mamba: Grouped Quantization-Aware Training of W1.58A16 State Space Models
Ternary Mamba: grouped QAT for W1.58A16 state space models
Querying an astronomical database using large language models: the ALeRCE text-to-SQL system
A text-to-SQL system for querying the ALeRCE astronomical database
S4oP: Operator-level Pruning of Structured State Space Models for Resource-Constrained Devices
S4oP prunes structured state space models at the operator level
A multi-agent framework against premature handoff and silent hallucination
NoiseTilt: Noise-Tilted Reverse Kernels for Diffusion Reward Alignment
NoiseTilt injects reward gradients via the noise term in diffusion
PseudoBench: Measuring How Agentic Auto-Research Fuels Pseudoscience
PseudoBench measures how agentic auto-research fuels pseudoscience
ConSA: Controllable Sparsity in Hybrid Attention via Learnable Allocation
ConSA: controllable sparsity in hybrid attention via learnable allocation
Compositional Skill Routing for LLM Agents: Decompose, Retrieve, and Compose
Compositional skill routing for LLM agents: decompose, retrieve, compose
ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents
ProvenanceGuard: source-aware factuality verification for MCP agents
Recursive Scaling in Masked Diffusion Models
Recursive scaling in masked diffusion models
LLM Consumer Behavior Theory: Foundations of a Novel Research Field
LLM Consumer Behavior Theory: a new field for agentic markets
Half a link can predict a whole link: generalization in KG foundation models
VoidPadding lets [VOID] handle padding so [EOS] focuses on termination
Differential Privacy of Gaussian Process Posterior Sampling
Differential privacy of Gaussian process posterior sampling
SoftMoE: Soft Differentiable Routing for Mixture-of-Experts in LLMs
SoftMoE: soft differentiable routing for mixture-of-experts in LLMs
Order-independent cell representations revisit table recognition
AnchorKV: Safety-Aware KV Cache Compression via Soft Penalty with a Refusal Anchor
AnchorKV: safety-aware KV cache compression via soft penalties
GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine?
GameCraft-Bench: can agents build playable games end-to-end?
Environment-Grounded Automated Prompt Optimization for LLM Game Agents
Environment-grounded automated prompt optimization for LLM game agents
From Drift to Coherence: Stabilizing Beliefs in LLMs
From drift to coherence: stabilizing beliefs in LLMs
Improving low-resource ASR via bilingual fine-tuning with language ID
A Framework for Evaluating Agentic Skills at Scale
A framework for evaluating agentic skills at scale
Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering
Position: coding benchmarks are misaligned with agentic software engineering
Vision-language models for chest radiography do not always need the image
Vision-language models for chest radiography do not always need the image
EComAgentBench: Benchmarking Shopping Agents on Long-Horizon Tasks with Distributed Hidden Intent
EComAgentBench: shopping agents on long-horizon tasks with hidden intent
LLMs Infer Cultural Context but Fail to Apply It When Responding
LLMs infer cultural context but fail to apply it when responding
EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning
EnvRL learns from environment dynamics in agentic RL
Beyond Domains: Reusing Web Skills via Transferable Interaction Patterns
Reusing web skills via transferable interaction patterns
OPD-Evolver: Cultivating Holistic Agent Evolver via On-Policy Distillation
OPD-Evolver cultivates self-evolving agents via on-policy distillation
Context-Aware RL for Agentic and Multimodal LLMs
ContextRL rewards picking the right context to ground answers
Exact Posterior Score Estimation for Solving Linear Inverse Problems
Exact closed-form posterior score for linear inverse problems
Benchmarking LLM Agents on Meta-Analysis Articles from Nature Portfolio
A benchmark for LLM agents on Nature Portfolio meta-analyses
DeepRubric: evidence-tree rubrics to boost deep-research agent RL
HAMON: Passive Optical Sequence Mixing for Long-Horizon Forecasting
HAMON: a passive optical core for long-horizon forecasting
TokenPilot: Cache-Efficient Context Management for LLM Agents
TokenPilot cuts LLM-agent context costs ~61% while preserving prompt cache
TuneJury: An Open Metric for Improving Music Generation Preference Alignment
TuneJury: an open reward model for text-to-music preference
Bayesian Inference and Decision Audits for Public Archives of Frontier AI Evaluations
Bayesian audit of public frontier-AI evaluation archives proposed
ActiveSAM: Image-Conditional Class Pruning for Fast and Accurate Open-Vocabulary Segmentation
ActiveSAM turns frozen SAM 3 into a training-free open-vocab segmenter
Agent trajectories as programs: fingerprinting and programming coding-agent behavior
Coding agents have behavioral fingerprints identifiable from trajectories
Dynestyx: A Probabilistic Programming Library for Dynamical Systems
Dynestyx: a probabilistic programming library with first-class SSMs
Decoupling Inference from State Updates in Low-Latency Feature Engines via Probabilistic Thinning
Probabilistic thinning decouples inference from state updates in streams
Probing Low Frame Rate Degradation in Neural Audio Codecs
Probing why neural audio codecs degrade at low frame rates
Beyond the Smile: A Hybrid Convolutional VAE for Crypto Volatility Surfaces
A convolutional VAE for completing crypto implied-volatility surfaces
Phantoms and Disclosures: a Causal Framework for Auditing Synthetic Data
A causal auditing framework to detect synthetic-data privacy disclosures
A Causal Model of Theory of Mind in Conflict for Artificial Intelligence
A structural causal model for when AI should engage theory of mind in conflict
Exploring Extrinsic and Intrinsic Properties for Effective Reasoning with Code Interpreter
Study probes extrinsic and intrinsic traits of code-interpreter reasoning
RAID: Semantic Graph Diffusion for True Cold-Start and Cross-Lingual Forecasting
RAID: retrieval-augmented diffusion for cold-start, cross-lingual forecasting
MA-SBI: Misspecification-Aware Simulation-Based Inference via Side-Channel Guidance
MA-SBI: misspecification-aware inference via side-channel guidance
Greed Is Learned: Visible Incentives as Reward-Hacking Triggers
Greed Is Learned: RL agents get addicted to visible reward channels
LESS Is More: Mutual-Stability Sampling for Diffusion Language Models
LESS: a training-free adaptive sampler for diffusion language models
Binary Tracking for Spatial QA and Navigation with Open Vision-Language Models
Binary Tracking: open vision-language models for spatial QA and navigation
Semantic Flip: synthetic OOD generation for robust refusal in embodied agents
Follow the Latent Roadmap: Navigating Revocable Decoding for Diffusion LLMs with Anchor Tokens
Anchor-token roadmap for revocable decoding in diffusion LLMs
Tying the Loop -- Tied Expert Layers in Mixture-of-Experts Language Models
Paper: Expert Tying shares MoE expert params across layers
Paper: framework measures LLM search-agent endorsement risk
GIST-CMTF: Goal-State Inference for Causal Minimal Tool Filtering in LLM Agents
GIST-CMTF adds goal-state inference to causal minimal tool filtering
LLM-based Visual Code Completion for Aerospace Geometric Design
Paper: LLM visual-programming copilot for aerospace design
LabOSBench: Benchmarking Computer Use Agents for Scientific Instrument Control
LabOSBench: a simulated testbed for computer-use agents controlling instruments
OpenClaw-Skill: Collective Skill Tree Search for Agentic Large Language Models
OpenClaw-Skill: collective skill tree search for LLM agents
Skill-to-LoRA: From Using Skills to Learning Behaviors for Token-Efficient LLM Agents
S2L replaces runtime SKILL.md text with skill-specific LoRA adapters
MyPCBench: A Benchmark for Personally Intelligent Computer-Use Agents
MyPCBench: benchmarking personal computer-use agents
Misinformation Propagation in Benign Multi-Agent Systems
Study on misinformation propagation in benign multi-agent systems
Progressive Knowledge-Guided Large Language Model Framework for Bearing Fault Diagnosis
Physics-guided multi-scale framework for bearing fault diagnosis
Multimodal Evaluator Preference Collapse: Cross-Modal Contagion in Self-Evolving Agents
Paper on evaluator preference collapse in self-evolving agents
FraudSMSWalker: Benchmarking Agentic Large Language Models for SMS-to-Webpage Fraud Detection
FraudSMSWalker benchmark targets URL-masked SMS-to-webpage fraud
VeriGraph: Towards Verifiable Data-Analytic Agents
VeriGraph: a traceable neuro-symbolic framework for verifiable data agents
SING: Synthetic Intention Graph for Scalable Active Tool Discovery in LLM Agents
SING: synthetic intention graph for scalable active tool discovery
Can LLM Agents Infer World Models? Evidence from Agentic Automata Learning
Can LLM agents infer world models? Evidence from automata learning
Can LLM Coding Agents Reason About Time Series?
Can LLM coding agents reason about time series? A benchmark study
DoubtProbe: Black-Box Jailbreak Defense via Structural Verification and Semantic Auditing
DoubtProbe: a dual-branch inference-time defense against LLM jailbreaks
daVinci-kernel: an RL framework co-evolving skills for GPU kernel tuning