Run NVIDIA Nemotron 3.5 Lightning on DGX Spark

NVIDIA Developer
6,380 views August 11, 2026

See how to deploy NVIDIA Nemotron 3.5 Lightning on DGX Spark and use it for fast, high-volume agentic workloads. In this walkthrough, we deploy Nemotron 3.5 Lightning with vLLM and DSpark speculative decoding, connect it to OpenCode, and put the model to work across 16 concurrent tasks, generating more than 500 tokens per second on a single DGX Spark. Nemotron 3.5 Lightning is an open 30B MoE model with 3B active parameters, built to serve as the fast execution model in always-on agent systems. It supports up to 1M tokens of context and is designed to work with agent harnesses including OpenCode, OpenClaw and Hermes Agent. You’ll learn how to: • Deploy Nemotron 3.5 Lightning on DGX Spark • Configure DSpark for speculative decoding • Connect the model to OpenCode • Run multiple agent tasks concurrently • Get started with customization and other deployment options 📝 Tech blog: https://nvda.ws/4xn0kMd 📥 Download: https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4

Keyboard shortcuts

On. Switch them off if they collide with your assistive tools; ? still opens this sheet.

Go to

Press g then the letter.

  • gh Latest
  • gs Sources
  • gm Media
  • gv Videos
  • gp Podcasts
  • gc Calendar
  • gd Decoder
  • gz Dataviz
  • ga Datasets
  • gb Blog
  • gk Markets
  • gj Careers
  • gn Prompt Notebook

On this page

  • / Focus search, where there is one
  • t Back to top
  • ? This list
  • Esc Close