NVIDIA × マルチモーダル

NVIDIA、世界行動モデル(WAM)の到来を概観

NVIDIA、世界行動モデル(WAM)の到来を概観

✎ ストーリー本文

NVIDIA が「世界行動モデル(World-Action Models, WAM)」の到来を概観する見解を示した。発生元は5件すべて NVIDIA 公式という一社発の構成――外部の研究やコミュニティの反応を伴わない、ベンダー主導のビジョン提示だ。世界モデルに行動生成を統合し、シミュレーション内で計画・行動できるエージェントへ向かう方向性で、本流はモデルの分類論というより「次に何が来るか」の地ならしにある。ロボティクスや自律システムの土台としての位置づけが強調される。ただし内容は概観・展望が中心で、具体的なモデルや実装、第三者による検証は伴っておらず、実体化の度合いは今後の発表待ちだ。

▲ 公式・報道
公式

Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models

NVIDIA Developer Blog ・ 2026-06-15 ・ 📌

NVIDIA、ロボット制御の新潮流「World-Action Model」を解説

公式

Building AI Agents for AR Glasses and XR Devices with NVIDIA XR AI

NVIDIA Developer Blog ・ 2026-06-16

NVIDIA、ARグラス/XR向けAIエージェント構築基盤「XR AI」を発表

公式

Build Your Own Transaction Foundation Model for Financial Intelligence

NVIDIA Developer Blog ・ 2026-06-16

NVIDIA、金融分析向け「取引基盤モデル」の自作手法を解説

報道

生成AI×自動運転で注目のTesla・Waymo・NVIDIA 各社が目指す「フィジカルAI」は何が違うのか

ITmedia AI+ ・ 2026-06-16

Tesla・Waymo・NVIDIA、自動運転で目指す「フィジカルAI」の違いを整理

公式

Build On-Device AI Companions with the NVIDIA ACE Game Agent SDK and Unreal Engine 5 Plugins

NVIDIA Developer Blog ・ 2026-06-16

NVIDIA、オンデバイス AI 向け ACE Game Agent SDK と UE5 プラグインを発表

公式

How to Optimize Transformer-Based Models for Low-Precision Training

NVIDIA Developer Blog ・ 2026-06-16

NVIDIA、Transformer モデルの低精度学習を最適化する解説記事を公開

公式

NVIDIA Blackwell Tops MLPerf Training 6.0 with Industry-Leading Scale and Performance

NVIDIA Developer Blog ・ 2026-06-16

NVIDIA、Blackwell が MLPerf Training 6.0 で首位と発表

コミュニティ

The Fable 5 Export Controls Harm US Cyber Defense

Simon Willison's Weblog ・ 2026-06-16

Simon Willison、Fable 5の輸出規制は米サイバー防衛を損なうと批判

公式

Fine-Tuning Biological Foundation Models with LoRA Using NVIDIA BioNeMo Recipes

NVIDIA Developer Blog ・ 2026-06-15

NVIDIA、BioNeMo Recipes で生物基盤モデルの LoRA fine-tuning 手法を解説

公式

Boosting MoE Training Throughput with Advanced Fusion Kernels

NVIDIA Developer Blog ・ 2026-06-15

NVIDIA、融合カーネルでMoE学習スループットを向上

学術(arxiv ほか) 16本 ▾
学術

Trust the Right Teacher: Quality-Aware Self-Distillation for GUI Grounding

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-16

GUI接地向けの品質考慮型自己蒸留手法を提案

学術

Uncertainty Quantification for Flow-Based Vision-Language-Action Models

arXiv cs.LG (Machine Learning) ・ 2026-06-16

フローベースVLAモデルの不確実性定量化

学術

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

arXiv cs.LG (Machine Learning) ・ 2026-06-16

Qwen-RobotManip、整合がロボット操作基盤モデルの規模化を解放

学術

Vision-language models for chest radiography do not always need the image

arXiv cs.CL (Computation and Language) ・ 2026-06-16

胸部X線の視覚言語モデルは画像を常に要しない

学術

Geometric Action Model for Robot Policy Learning

arXiv cs.LG (Machine Learning) ・ 2026-06-15

幾何基盤モデルを再利用するロボット操作方策GAMを提案

学術

Learning the Geometry of Data: A Mathematical Review of Shape Space Analysis

arXiv cs.LG (Machine Learning) ・ 2026-06-15

データの幾何学を学ぶ ─ 形状空間解析の数理を体系的にレビュー

学術

FusionRS: A Large-Scale RGB-Infrared Remote Sensing Dataset for Dual-Modal Vision-Language Foundation Models

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-15

RGB・赤外を対応づけたリモセン向けデータセットFusionRSを公開

学術

ROVE: Unlocking Human Interventions for Humanoid Manipulation via Reinforcement Learning

arXiv cs.LG (Machine Learning) ・ 2026-06-15

ROVE、不完全な人手介入から学ぶヒューマノイド操作のRL枠組み

学術

Functional Gradient Descent with Adaptive Representations

arXiv cs.LG (Machine Learning) ・ 2026-06-15

関数空間で勾配降下するFGDを適応的表現で実用化する手法を提案

学術

Binary Tracking for Spatial QA and Navigation with Open Vision-Language Models

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-15

オープン VLM で動く空間質問応答・ナビ手法 Binary Tracking を提案

学術

Semantic Flip: Synthetic OOD Generation for Robust Refusal in Embodied Question Answering and Spatial Localization

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-15

身体化エージェントの拒否応答を強化する合成 OOD 生成手法 Semantic Flip を提案

学術

A Perception vs. Distortion Perspective on Score-Based Generative Channel Estimation

arXiv cs.LG (Machine Learning) ・ 2026-06-15

スコアベース通信路推定を知覚-歪みトレードオフの観点で理論解析

学術

LabOSBench: Benchmarking Computer Use Agents for Scientific Instrument Control

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-15

科学機器を操作するエージェント評価へ、模擬ベンチLabOSBenchを提案

学術

Decoupling Semantics from Distortions: Multi-Scale Two-Stream Vision-Language Alignment for AI-Generated Image Quality Assessment

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-15

AI生成画像の品質評価へ、意味と歪みを分離する二系統手法MST-CLIPIQAを提案

学術

Decision-Weighted Flow Matching for Contextual Stochastic Optimization

arXiv cs.LG (Machine Learning) ・ 2026-06-15

DW-FM: 下流の意思決定の後悔に整合する重み付きフローマッチングを提案

学術

Uncertainty Is Not a Safety Net for Clinical VQA, but Can It Anticipate Model Failure?

arXiv cs.CL (Computation and Language) ・ 2026-06-15

臨床 VQA で不確実性推定は安全網にならないと検証

← ストーリー アーカイブ