NVIDIA × Inference & Efficiency

NVIDIA: EPD split for multimodal serving

NVIDIA: EPD split for multimodal serving

✎ Story body

In the same week NVIDIA set out when to split multimodal inference across separate resources, a run of Vision-Language-Action papers pushed models further toward acting, not just seeing and reading.

What happened

NVIDIA's technical blog described when to place the Encode, Prefill and Decode stages of multimodal serving on separate resources — "EPD disaggregation" — to cut latency and lift throughput. In the same week, arXiv carried a run of Vision-Language-Action work, including TANGO, which drives a humanoid's whole body through cluttered environments, asking a single model to see, read and act.

Why it matters

EPD disaggregation is a serving-layer design: how to deliver models that already exist more cheaply and quickly. The VLA work sits on the other layer, extending what a model is asked to do at all. The two do different jobs and neither substitutes for the other, so watching only one of them misjudges how close deployment actually is. Efficiency in serving and growth in capability advance at different speeds. Note: the arXiv papers are pre-review, and their performance claims await independent replication.

What to watch

Deployed cases where EPD disaggregation is adopted and its effect reported in numbers. On the research side, how long it takes for VLA results to reach physical robots or shipping products.

▲ Official & Press
Official

When to Use Encode-Prefill-Decode Disaggregation to Accelerate Multimodal Model Serving

NVIDIA Developer Blog ・ 2026-09-09 ・ 📌

NVIDIA: when to use encode-prefill-decode disaggregation for multimodal

Official

How Full-Stack NIM Optimizations Deliver 2.5x More Users on Nemotron 3 Ultra

NVIDIA Developer Blog ・ 2026-09-10

NVIDIA: NIM full-stack tuning yields 2.5x more Nemotron 3 Ultra users

Official

High-Throughput Structure Prediction with BioNeMo Inference Runtime

NVIDIA Developer Blog ・ 2026-09-10

NVIDIA details proteome-scale structure prediction with BioNeMo

Press

「Claude」による不正アクセス、4件目が判明──Anthropic、「アライメントの失敗」と評価を修正

ITmedia AI+ ・ 2026-09-10

Anthropic reclassifies four Claude intrusion cases as alignment failures

Official

DeepSeek-V4.1-Flash: Smarter, Faster, More Efficient

DeepSeek API Docs / News ・ 2026-09-10

DeepSeek launches V4.1-Flash, a 552B MoE that outpaces V4-Pro

Official

Introducing North Small Translate: A leading sovereign open-weight machine translation model

Cohere Blog ・ 2026-09-10

Cohere releases North Small Translate, an open-weight MoE translation model

Community

Serving LLMs on Tenstorrent Hardware: Inside the vLLM TT Plugin

Lobste.rs (AI tagged) ・ 2026-09-09

Inside the vLLM TT plugin for serving LLMs on Tenstorrent

Academic (arxiv etc.) 24 ▾
Academic

Likelihood-free inference with nuisance parameters through normalizing flows

arXiv cs.LG (Machine Learning) ・ 2026-09-09

Academic

Cross-Model Agreement as a Deployment-Time Reliability Signal for Automatic Polyp Segmentation

arXiv cs.LG (Machine Learning) ・ 2026-09-09

Academic

Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs

arXiv cs.CL (Computation and Language) ・ 2026-09-09

Academic

Beyond One-Size-Fits-All: Sample-Adaptive Strategy Routing for Vision Token Pruning in MLLMs

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-09

Academic

Structural Fusion of Bayesian Networks with Limited Treewidth Using Genetic Algorithms

arXiv cs.LG (Machine Learning) ・ 2026-09-09

Academic

Maverick: Private and Verifiable LLM Inference Made Practical via Matrix-Vector Multiplication Delegation

arXiv cs.LG (Machine Learning) ・ 2026-09-09

Academic

CompassOPD: Cross-Family On-Policy Distillation via Within-Family Likelihood Shifts

arXiv cs.LG (Machine Learning) ・ 2026-09-09

Academic

Kernel-Managed Shared Memory for System-Wide Personalization

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-09

Academic

ProbPlug: A Plugin Uncertainty Network for Reliable Confidence in LLM Binary Classification

arXiv cs.CL (Computation and Language) ・ 2026-09-09

Academic

A Trust-Network-Based Federated Learning Framework for Multi-Center Aging Clock Prediction

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-09

Academic

Beyond Training: A Feasibility Taxonomy for Inference-Time AI Governance

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-09

Academic

Orukeet: Multilingual ASR with Frozen Gabor Kernels

arXiv cs.LG (Machine Learning) ・ 2026-09-09

Academic

MedDeID enables locally governed clinical-text de-identification from real or synthetic training data

arXiv cs.CL (Computation and Language) ・ 2026-09-09

Academic

Belief-State Engine: Augmenting LLMs for Principled Planning Under Partial Observability

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-09

Academic

Deterministic Prompting for Speaker-Stable Low-Resource Greek TTS

arXiv cs.CL (Computation and Language) ・ 2026-09-09

Academic

Elastoformer: Enabling Dynamic Adaptivity via Elastic Model Transformation

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-09

Academic

SalamandraTA at WMT 2026 Terminology Shared Task: Hard Examples Are Better Teachers

arXiv cs.CL (Computation and Language) ・ 2026-09-09

Academic

Stable Answers, Unfinished Reasoning: Why Self-Consensus Is Not a Safe Early-Exit Signal

arXiv cs.CL (Computation and Language) ・ 2026-09-09

Academic

VLX-VR: An Agentic-Aware Video Reasoning Model

arXiv cs.CL (Computation and Language) ・ 2026-09-09

Academic

TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-08

Academic

Canonical Color as a Lens into Concept Decodability in Vision Encoders and VLMs

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-08

Academic

DeCAL: Towards Physically-Grounded Dexterous Vision-Language-Action Models via Contact-Aware Latent Co-Imagination

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-08

Academic

GoDeep: Annotation-Free Open-Vocabulary 3D Scene Understanding via Language-Space Lifting

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-08

Academic

Neither Adversarial Training Nor Purification: Emergent Adversarial Robustness from Oscillatory Predictive Learning

arXiv cs.LG (Machine Learning) ・ 2026-09-08

← Story Archive