NVIDIA × 推論・効率化

NVIDIA、マルチモーダル推論のEPD分割を解説

NVIDIA、マルチモーダル推論のEPD分割を解説

✎ ストーリー本文

NVIDIA が提供側でマルチモーダル推論の分割配置を整理した同じ週、研究側では見て読んで動く VLA 系の論文が相次いだ。

何が起きたか

NVIDIA の技術ブログが、マルチモーダルモデルを提供する際に Encode (符号化)・Prefill (先読み)・Decode (生成) の各段階を別々の資源に分けて置く「EPD 分割」の使いどころを解説した。狙いは待ち時間の短縮とスループットの改善にある。同じ週の arXiv には、乱雑な環境でヒューマノイドの全身を動かす TANGO をはじめ、視覚と言語に加えて行動 (Action) まで一つのモデルに担わせる VLA 系の投稿が続いた。

なぜ重要か

EPD 分割は、既にあるモデルをいかに速く安く配るかという実装側の設計であり、VLA 系の研究はモデルに何をさせるかを押し広げる側にある。二つは別の役割を持ち、互いに置き換えるものではない。片方だけを見ていると、実用化がどれだけ近いのかの見積もりを外す。提供の効率が上がることと、できることが増えることは、別々の速度で進む。※arXiv 段階の論文は査読前で、性能の主張はいずれも独立した検証待ちだ。

次に何を見るか

EPD 分割を実サービスで採用し、効果を数字で公表する事例が出るか。研究側は、VLA の成果が実機のロボットや出荷済みの製品に組み込まれるまでの時間が目安になる。

▲ 公式・報道
公式

When to Use Encode-Prefill-Decode Disaggregation to Accelerate Multimodal Model Serving

NVIDIA Developer Blog ・ 2026-09-09 ・ 📌

NVIDIA、マルチモーダル推論を高速化する EPD 分離の使いどころを解説

公式

How Full-Stack NIM Optimizations Deliver 2.5x More Users on Nemotron 3 Ultra

NVIDIA Developer Blog ・ 2026-09-10

NVIDIA、NIM の全スタック最適化で Nemotron 3 Ultra の同時処理を 2.5 倍に

公式

High-Throughput Structure Prediction with BioNeMo Inference Runtime

NVIDIA Developer Blog ・ 2026-09-10

NVIDIA、BioNeMo 推論ランタイムでタンパク質構造予測を高速化

報道

「Claude」による不正アクセス、4件目が判明──Anthropic、「アライメントの失敗」と評価を修正

ITmedia AI+ ・ 2026-09-10

Anthropic、Claude の不正侵入 4 件目を確認しアライメントの失敗と評価修正

公式

DeepSeek-V4.1-Flash: Smarter, Faster, More Efficient

DeepSeek API Docs / News ・ 2026-09-10

DeepSeek、552B MoE の V4.1-Flash 公開、V4-Pro を置き換えへ

公式

Introducing North Small Translate: A leading sovereign open-weight machine translation model

Cohere Blog ・ 2026-09-10

Cohere、MoE 翻訳モデル North Small Translate を公開

コミュニティ

Serving LLMs on Tenstorrent Hardware: Inside the vLLM TT Plugin

Lobste.rs (AI tagged) ・ 2026-09-09

Tenstorrent 上で LLM を動かす vLLM TT プラグインの内部

学術(arxiv ほか) 24本 ▾
学術

Likelihood-free inference with nuisance parameters through normalizing flows

arXiv cs.LG (Machine Learning) ・ 2026-09-09

学術

Cross-Model Agreement as a Deployment-Time Reliability Signal for Automatic Polyp Segmentation

arXiv cs.LG (Machine Learning) ・ 2026-09-09

学術

Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs

arXiv cs.CL (Computation and Language) ・ 2026-09-09

学術

Beyond One-Size-Fits-All: Sample-Adaptive Strategy Routing for Vision Token Pruning in MLLMs

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-09

学術

Structural Fusion of Bayesian Networks with Limited Treewidth Using Genetic Algorithms

arXiv cs.LG (Machine Learning) ・ 2026-09-09

学術

Maverick: Private and Verifiable LLM Inference Made Practical via Matrix-Vector Multiplication Delegation

arXiv cs.LG (Machine Learning) ・ 2026-09-09

学術

CompassOPD: Cross-Family On-Policy Distillation via Within-Family Likelihood Shifts

arXiv cs.LG (Machine Learning) ・ 2026-09-09

学術

Kernel-Managed Shared Memory for System-Wide Personalization

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-09

学術

ProbPlug: A Plugin Uncertainty Network for Reliable Confidence in LLM Binary Classification

arXiv cs.CL (Computation and Language) ・ 2026-09-09

学術

A Trust-Network-Based Federated Learning Framework for Multi-Center Aging Clock Prediction

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-09

学術

Beyond Training: A Feasibility Taxonomy for Inference-Time AI Governance

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-09

学術

Orukeet: Multilingual ASR with Frozen Gabor Kernels

arXiv cs.LG (Machine Learning) ・ 2026-09-09

学術

MedDeID enables locally governed clinical-text de-identification from real or synthetic training data

arXiv cs.CL (Computation and Language) ・ 2026-09-09

学術

Belief-State Engine: Augmenting LLMs for Principled Planning Under Partial Observability

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-09

学術

Deterministic Prompting for Speaker-Stable Low-Resource Greek TTS

arXiv cs.CL (Computation and Language) ・ 2026-09-09

学術

Elastoformer: Enabling Dynamic Adaptivity via Elastic Model Transformation

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-09

学術

SalamandraTA at WMT 2026 Terminology Shared Task: Hard Examples Are Better Teachers

arXiv cs.CL (Computation and Language) ・ 2026-09-09

学術

Stable Answers, Unfinished Reasoning: Why Self-Consensus Is Not a Safe Early-Exit Signal

arXiv cs.CL (Computation and Language) ・ 2026-09-09

学術

VLX-VR: An Agentic-Aware Video Reasoning Model

arXiv cs.CL (Computation and Language) ・ 2026-09-09

学術

TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-08

学術

Canonical Color as a Lens into Concept Decodability in Vision Encoders and VLMs

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-08

学術

DeCAL: Towards Physically-Grounded Dexterous Vision-Language-Action Models via Contact-Aware Latent Co-Imagination

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-08

学術

GoDeep: Annotation-Free Open-Vocabulary 3D Scene Understanding via Language-Space Lifting

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-08

学術

Neither Adversarial Training Nor Purification: Emergent Adversarial Robustness from Oscillatory Predictive Learning

arXiv cs.LG (Machine Learning) ・ 2026-09-08

← ストーリー アーカイブ