AI Agents × Safety & Evaluation

xAI ships Grok Build coding agent

xAI ships Grok Build coding agent

✎ Story body

xAI released 'Grok Build,' a coding-assistant agent. The composition leans academic—four arXiv papers plus one ITmedia report—research and press rather than an official announcement, a release with technical background. Beyond code generation, the focus is agentic development support that autonomously runs plan, implement, and fix; the through-line is less generation quality itself than the spread across vendors of a race over agents you can hand development workflows to. It reads as xAI entering a space where OpenAI and NVIDIA are ahead. But it's right after release—actual code quality, how it compares with existing agents, and adoption in real development remain to confirm.

▲ Official & Press
Press

xAIがコーディングエージェント「Grok Build」ベータ公開。サブエージェントを並列に実行可能など

Publickey ・ 2026-05-25 ・ 📌

xAI opens "Grok Build" coding agent beta, runs subagents in parallel

Academic (arxiv etc.) 16 ▾
Academic

MobileMoE: Scaling On-Device Mixture of Experts

arXiv cs.LG (Machine Learning) ・ 2026-05-26

MobileMoE: sub-billion on-device MoE LMs claim a new on-device Pareto frontier

Academic

Causal Risk Minimization for High-Dimensional Treatments

arXiv cs.LG (Machine Learning) ・ 2026-05-26

Causal risk minimization for high-dimensional treatment spaces

Academic

Learning When to Think While Listening in Large Audio-Language Models

arXiv cs.LG (Machine Learning) ・ 2026-05-26

Learnable wait-think-answer control balances latency and reasoning in audio LMs

Academic

MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent Research

arXiv cs.AI (Artificial Intelligence) ・ 2026-05-25

MobileGym launches verifiable, browser-hosted mobile GUI agent simulator.

Academic

Beyond Summaries: Structure-Aware Labeling of Code Changes with Large Language Models

arXiv cs.AI (Artificial Intelligence) ・ 2026-05-25

Paper: LLMs label code changes via 2-stage pipeline, 84% recall achieved.

Academic

Automated Benchmark Auditing for AI Agents and Large Language Models

arXiv cs.CL (Computation and Language) ・ 2026-05-25

Auto Benchmark Audit flags critical issues across 168 LLM and agent benchmarks

Academic

Confidence and Calibration of Activation Oracles for Reliable Interpretation of Language Model Internals

arXiv cs.AI (Artificial Intelligence) ・ 2026-05-25

Bootstrap mode frequency: best-calibrated confidence for activation oracles

Academic

Peak-Then-Collapse and the Four Interface Channels of Knowledge-Graph Tool Use

arXiv cs.CL (Computation and Language) ・ 2026-05-25

Peak-then-collapse: minimal KG tool RLVR yields four recurring failure modes

Academic

CausaLab: A Scalable Environment for Interactive Causal Discovery Toward AI Scientists

arXiv cs.AI (Artificial Intelligence) ・ 2026-05-25

CausaLab evaluates interactive causal discovery by LLM agents in a virtual lab

Academic

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models

arXiv cs.CL (Computation and Language) ・ 2026-05-25

STORMS internalizes spatial-temporal reasoning in video-language models

Academic

MAGIC: Multimodal Alignment & Grounding-aware Instruction Coreset for Vision-Language Models

arXiv cs.CL (Computation and Language) ・ 2026-05-25

MAGIC: training-free, forward-only coreset for multimodal instruction tuning

Academic

What Makes a Medical Checker Trainable? Diagnosing Signal Collapse and Reward Hacking in Checker-Guided RAG for Biomedical QA

arXiv cs.CL (Computation and Language) ・ 2026-05-25

Output distribution decides whether NLI checkers train medical RAG agents

Academic

Neural Scalable Symbolic Search Framework for Complex Logical Queries with Multiple Free Variables

arXiv cs.AI (Artificial Intelligence) ・ 2026-05-25

NeSyS-k: neural symbolic search for multi-variable complex KG query answering

Academic

When Do LLM Agents Treat Surface Noise Differently from Semantic Noise? A 68-Cell Measurement Study with a Held-Out Trace-Level Validation

arXiv cs.CL (Computation and Language) ・ 2026-05-25

LLM agents treat semantic noise far more disruptively than surface noise

Academic

Creative Quality Alignment: Expert Tacit Knowledge Transfer via Chain-of-Thought Fine-Tuning

arXiv cs.AI (Artificial Intelligence) ・ 2026-05-25

Empirical study confirms Calibrated Surprise metric at engineering scale

Academic

From Latent Space to Training Data: Explainable Specialization in Minimal MLPs

arXiv cs.AI (Artificial Intelligence) ・ 2026-05-25

Coverage regularization in minimal MLPs aids dataset reconstruction

← Story Archive