NVIDIA published a step-by-step recipe for creating an NVFP4 checkpoint of Nemotron 3 Ultra using the Model Optimizer toolkit. Coverage was official-only and technical. The theme isn't a new model but a sharpening of the how-to: getting the same model to run thinner. As long-running agents make inference cost the real constraint, quantization playbooks land harder than model launches. Next signal: adoption reports from teams actually shipping FP4 inference at scale.
NVIDIA × Inference & Efficiency
NVIDIA optimizes Nemotron 3 Ultra
NVIDIA optimizes Nemotron 3 Ultra
✎ Story body
▲ Official & Press
Official
Creating the NVIDIA Nemotron 3 Ultra NVFP4 Checkpoint with NVIDIA Model Optimizer
NVIDIA builds an NVFP4 Nemotron 3 Ultra checkpoint via Model Optimizer
Academic (arxiv etc.) 7 ▾
Academic