Run NVIDIA Nemotron 3.5 Lightning on DGX Spark
See how to deploy NVIDIA Nemotron 3.5 Lightning on DGX Spark and use it for fast, high-volume agentic workloads. In this walkthrough, we deploy Nemotron 3.5 Lightning with vLLM and DSpark speculative decoding, connect it to OpenCode, and put the model to work across 16 concurrent tasks, generating more than 500 tokens per second on a single DGX Spark. Nemotron 3.5 Lightning is an open 30B MoE model with 3B active parameters, built to serve as the fast execution model in always-on agent systems. It supports up to 1M tokens of context and is designed to work with agent harnesses including OpenCode, OpenClaw and Hermes Agent. You’ll learn how to: • Deploy Nemotron 3.5 Lightning on DGX Spark • Configure DSpark for speculative decoding • Connect the model to OpenCode • Run multiple agent tasks concurrently • Get started with customization and other deployment options 📝 Tech blog: https://nvda.ws/4xn0kMd 📥 Download: https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4