NVIDIA × 推論・効率化

NVIDIA、推論の効率と耐障害性の指針を公開

NVIDIA、推論の効率と耐障害性の指針を公開

✎ ストーリー本文

NVIDIA が同じ日に公開した三つの技術解説は、モデル・実行・配線と扱う層がばらばらだ。どれも計算を速くする話ではない。

何が起きたか

一つ目は、密結合型と Mixture-of-Experts の選び分けだ。例に挙がる Nemotron 3.5 Lightning は、総パラメータ 30B のうち、トークン 1 つの処理に 3B だけを働かせる。必要な部分だけを選んで動かせば、全体を毎回動かさずに大きな容量を使える。二つ目は Vera Rubin 上で動く Groq 3 LPX で、実行のタイミングをあらかじめ確定させて無駄な待ちを減らす。三つ目は NVLink 6 で、リンク層からシステム層まで冗長と回復の段を重ね、障害が起きても学習を続けさせる。

なぜ重要か

三つは別々の技術だが、並べると同じ方向を向いて見える。演算器を速くするのではなく、限られた資源をどう使い切り、どう止めないかの話だ。とくに三つ目が効くのは、大規模な学習ではクラスタ内の全 GPU が足並みを揃えて動くため、1 台が止まれば学習全体が止まるからだ。ただし電力を制約として明示しているのは後ろの二つで、モデルの選び分けはあくまで処理量の話として書かれている。

次に何を見るか

公開されたのは設計思想であって、実測値ではない。電力あたりでどれだけ処理できたか、障害が起きたとき学習が実際どこまで続いたかを導入側が自分の環境で出せるようになったとき、この三層が効いているかが分かる。

▲ 公式・報道
公式

Dense vs. MoE Models: Active Parameters, Throughput, and When to Choose Each

NVIDIA Developer Blog ・ 2026-09-15 ・ 📌

NVIDIA、Dense と MoE の使い分けを活性パラメータ視点で解説

公式

Bypassing inference bottlenecks: Accelerating complex AI search with Retrieve-for-Train

Google Research Blog ・ 2026-09-15

Google、推論時思考を回避する検索枠組み Retrieve-for-Train

公式

How NVIDIA Groq 3 LPX Deterministic Execution Drives Power-Efficient High-Interactivity Inference on NVIDIA Vera Rubin

NVIDIA Developer Blog ・ 2026-09-15

NVIDIA、Groq 3 LPX の決定的実行で電力効率と応答性を両立

公式

How NVIDIA NVLink 6 Delivers Multi-Layer Resiliency for AI Factories

NVIDIA Developer Blog ・ 2026-09-15

NVIDIA、NVLink 6 の多層冗長で AI ファクトリの稼働を確保

報道

How OpenAI Used Its Own LLMs to Design Its Jalapeño Chip

IEEE Spectrum (AI section) ・ 2026-09-14

OpenAI、自社LLMで初の自社チップJalapeñoを設計、RTLからテープアウトまで9カ月

学術(arxiv ほか) 22本 ▾
学術

Stellar Colosseum: A Many-Agent Harness for Long-Horizon Research in Mathematics and Theoretical Computer Science

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-14

学術

Bridging Control, Inference, Transport, and Thermodynamics: From Theory to Applications in Learning

arXiv cs.LG (Machine Learning) ・ 2026-09-14

学術

Anatomical Grounding and Leakage-Aware Multimodal Contrastive Learning for Alzheimer's Disease Classification from Structural MRI

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-14

学術

Learning to Coach for Experiential Learning

arXiv cs.CL (Computation and Language) ・ 2026-09-14

学術

Proportional-Fair Resource Allocation and Dual-Threshold Early-Exit Inference for Secure Cooperative Multi-Layer Edge Intelligence

arXiv cs.LG (Machine Learning) ・ 2026-09-14

学術

Enabling Streaming User Transcription in Full-Duplex Speech-to-Speech Models

arXiv cs.CL (Computation and Language) ・ 2026-09-14

学術

Look Before You Leap: Factual Decoding with Internal Attribution Signals

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-14

学術

Merging the Knowledge of LLMs for Automatic Speech Recognition

arXiv cs.CL (Computation and Language) ・ 2026-09-14

学術

Backward SDEs-based Diffusion for Physics-Constrained Generation

arXiv cs.LG (Machine Learning) ・ 2026-09-14

学術

More Than Just Access: Generative AI as Communication Intermediary for Blind and Low-Vision Users

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-14

学術

Don't Send What You Don't Need: Question-Guided Token Pruning as a Privacy Defense for Vision-Language Models

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-14

学術

Circuit-MLLM: Topological Logic-Guided Latent-Space Visual Reasoning for Circuit Schematic Understanding

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-14

学術

Through the Eyes of the Beholder: Biometric and Demographic Conditioning for Multimodal Sexism Detection

arXiv cs.CL (Computation and Language) ・ 2026-09-14

学術

Option-Aware Retrieval and Task-Specific VLM Adaptation for Medical VQA

arXiv cs.CL (Computation and Language) ・ 2026-09-14

学術

To Each Language Its Tokenizer: Modular Tokenizers for Efficient Multilingual LLMs

arXiv cs.CL (Computation and Language) ・ 2026-09-14

学術

Beyond Noise: Understanding and Overcoming Temperature Effects in Analog DNN Inference

arXiv cs.LG (Machine Learning) ・ 2026-09-14

学術

How Lossless Is Lossless Speculative Decoding? The Role of Numerical Precision in Orthrus

arXiv cs.CL (Computation and Language) ・ 2026-09-14

学術

MarKey: Marginal Utility Guided Greedy Keyframe Selection for Long Video Understanding

arXiv cs.CL (Computation and Language) ・ 2026-09-14

学術

Dynamic Semantic Compression for Efficient Latent-Space Inference in Large Language Models

arXiv cs.CL (Computation and Language) ・ 2026-09-14

学術

Reason What Matters: Retrieval-Grounded Reasoning for Universal Multimodal Embeddings

arXiv cs.CL (Computation and Language) ・ 2026-09-14

学術

Semiotic Relations and Proof Methods: A Cross-Genre Study of Argument Structure with Large Language Models

arXiv cs.CL (Computation and Language) ・ 2026-09-14

学術

EMR: Self-Evolving Medical Multi-Agent System via Experience Mining and Reuse

arXiv cs.CL (Computation and Language) ・ 2026-09-14

← ストーリー アーカイブ