What Is NVFP4? Faster LLM Inference Without Losing Quality

NVIDIA Developer
2,564 views July 23, 2026

Learn what NVFP4 is, why it helps you run bigger LLMs on less GPU memory without a big quality hit, and how to create an NVFP4‑quantized Nemotron 3 Ultra checkpoint using NVIDIA Model Optimizer. For more details, see these blogs on Nemotron 3 Ultra NVFP4 quantization and the NVFP4 format. 🔗 https://developer.nvidia.com/blog/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer/ 🔗 https://developer.nvidia.com/blog/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer/ 🔗 https://developer.nvidia.com/blog/introducing-nvfp4-for-efficient-and-accurate-low-precision-inference/

Keyboard shortcuts

On. Switch them off if they collide with your assistive tools; ? still opens this sheet.

Go to

Press g then the letter.

  • gh Latest
  • gs Sources
  • gm Media
  • gv Videos
  • gp Podcasts
  • gc Calendar
  • gd Decoder
  • gz Dataviz
  • ga Datasets
  • gb Blog
  • gk Markets
  • gj Careers
  • gn Prompt Notebook

On this page

  • / Focus search, where there is one
  • t Back to top
  • ? This list
  • Esc Close