AI エージェント × 安全性・評価

xAI、コーディングエージェントGrok Build公開

xAI、コーディングエージェントGrok Build公開

✎ ストーリー本文

xAI がコーディング支援エージェント「Grok Build」を公開した。発生元は arXiv 4件に itmedia の報道1件が付く学術寄りの構成で、公式発表というより研究と報道が起点――技術的背景を伴う公開だ。コードの生成にとどまらず、計画・実装・修正を自律的に回すエージェント型の開発支援が主眼で、本流はモデルの生成品質そのものより、開発工程を任せられるエージェントの競争が各社に広がっている点にある。OpenAI や NVIDIA が先行するこの領域に xAI が参入した構図として読める。ただし公開直後で、実際のコード品質や既存エージェントとの優劣、開発現場での採用は、今後の確認点として残る。

▲ 公式・報道
報道

xAIがコーディングエージェント「Grok Build」ベータ公開。サブエージェントを並列に実行可能など

Publickey ・ 2026-05-25 ・ 📌

xAI、コーディングエージェント「Grok Build」ベータ公開、サブ agent 並列実行

学術(arxiv ほか) 16本 ▾
学術

MobileMoE: Scaling On-Device Mixture of Experts

arXiv cs.LG (Machine Learning) ・ 2026-05-26

MobileMoE、端末向け1B未満MoE言語モデルで密モデル比2-4倍の効率を実現すると主張

学術

Causal Risk Minimization for High-Dimensional Treatments

arXiv cs.LG (Machine Learning) ・ 2026-05-26

高次元介入空間に対応する因果リスク最小化を学習問題として定式化

学術

Learning When to Think While Listening in Large Audio-Language Models

arXiv cs.LG (Machine Learning) ・ 2026-05-26

Wait-Think-Answer、音声-言語モデル向けストリーミング思考制御を学習する手法

学術

MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent Research

arXiv cs.AI (Artificial Intelligence) ・ 2026-05-25

MobileGym、検証可能な判定と並列実行に対応するブラウザ製モバイル GUI エージェント研究基盤を公開。

学術

Beyond Summaries: Structure-Aware Labeling of Code Changes with Large Language Models

arXiv cs.AI (Artificial Intelligence) ・ 2026-05-25

論文:LLM でコード変更にラベル付けする 2 段階手法、recall 84% を達成。

学術

Automated Benchmark Auditing for AI Agents and Large Language Models

arXiv cs.CL (Computation and Language) ・ 2026-05-25

Auto Benchmark Audit、168のAIベンチマークから設計欠陥を体系検出

学術

Confidence and Calibration of Activation Oracles for Reliable Interpretation of Language Model Internals

arXiv cs.AI (Artificial Intelligence) ・ 2026-05-25

Activation Oracle信頼度推定、bootstrap mode frequencyが最高較正

学術

Peak-Then-Collapse and the Four Interface Channels of Knowledge-Graph Tool Use

arXiv cs.CL (Computation and Language) ・ 2026-05-25

RLVRツール利用は「ピーク後崩壊」、4種の失敗モードを観察

学術

CausaLab: A Scalable Environment for Interactive Causal Discovery Toward AI Scientists

arXiv cs.AI (Artificial Intelligence) ・ 2026-05-25

CausaLab、LLMエージェントの対話的因果発見を仮想ラボで評価

学術

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models

arXiv cs.CL (Computation and Language) ・ 2026-05-25

STORMS、テキストCoTを介さず内部化した時空間推論で動画-言語モデルを高速化

学術

MAGIC: Multimodal Alignment & Grounding-aware Instruction Coreset for Vision-Language Models

arXiv cs.CL (Computation and Language) ・ 2026-05-25

MAGIC、学習不要の前向きシグナルで多モーダル指示Coresetを構築

学術

What Makes a Medical Checker Trainable? Diagnosing Signal Collapse and Reward Hacking in Checker-Guided RAG for Biomedical QA

arXiv cs.CL (Computation and Language) ・ 2026-05-25

医療RAGのNLI報酬は分布次第、log-probはシグナル崩壊と指摘

学術

Neural Scalable Symbolic Search Framework for Complex Logical Queries with Multiple Free Variables

arXiv cs.AI (Artificial Intelligence) ・ 2026-05-25

NeSyS-k、k自由変数を持つ複雑論理クエリ応答を共同ランクで近似

学術

When Do LLM Agents Treat Surface Noise Differently from Semantic Noise? A 68-Cell Measurement Study with a Held-Out Trace-Level Validation

arXiv cs.CL (Computation and Language) ・ 2026-05-25

LLMエージェント、意味的擾乱に表層擾乱より約20pt多く回答が変わる

学術

Creative Quality Alignment: Expert Tacit Knowledge Transfer via Chain-of-Thought Fine-Tuning

arXiv cs.AI (Artificial Intelligence) ・ 2026-05-25

Calibrated Surprise指標の実装検証、低資源/小モデル条件で実現可と報告

学術

From Latent Space to Training Data: Explainable Specialization in Minimal MLPs

arXiv cs.AI (Artificial Intelligence) ・ 2026-05-25

最小MLPの隠れニューロン特化が学習データのプロトタイプ復元を改善

← ストーリー アーカイブ