NVIDIA Developer

Local AI: Running Nemotron on DGX Spark & Station | Nemotron Labs

Published 2026-09-24 · Duration 56:06

Summary

The session details the growing viability of local AI, emphasizing that running Large Language Models (LLMs) on dedicated local hardware like DGX Spark and DGX Station offers significant advantages in sovereignty, privacy, and cost efficiency compared to cloud-only inference. Key technical advancements include the 'Exo backend,' a unified software stack that enables consistent model deployment across diverse hardware (Nvidia and Apple silicon). The discussion also covered advanced scaling techniques, such as Process/Data (PD) Disaggregation, allowing compute-intensive parts of inference to run on one device while memory-intensive parts run on another, maximizing utilization of heterogeneous local compute resources.

Download summary

Key takeaways

  1. Local AI Value Proposition 2:50

    Local AI provides sovereignty, privacy, and cost efficiency by eliminating per-token cloud costs. For single-user, low-concurrency workloads, local inference can achieve high tokens per second, matching or exceeding the performance of cloud serving setups (Transcript, 0:02:50).

  2. Unified Local AI Stack (Exo Backend) 7:20

    The Exo backend addresses the fragmentation problem in local AI by providing a unified stack that runs across different hardware targets, including Nvidia hardware and Apple silicon (M3 Ultra). This ensures that models can be run consistently regardless of the underlying device (Transcript, 0:07:20).

  3. Advanced Scaling and Hardware Utilization 11:17

    DGX Spark and DGX Station support both horizontal (clustering) and vertical scaling. DGX Station, with up to 750 GB of unified VRAM, is recommended for running very large open-source models (e.g., GLM 5.3) that exceed the capacity of a single Spark unit (Transcript, 1:17:00).

Technical details

  • Inference Economics (Local vs. Cloud) 130s

    Cloud serving optimizes for high throughput per dollar via batching, often sacrificing individual user interactivity. Local AI, by contrast, has a marginal cost of generating a token close to zero (only power cost), making it economically superior for self-contained workloads (Transcript, 0:02:10).

  • Process/Data (PD) Disaggregation 340s

    PD Disaggregation allows splitting inference tasks between different accelerators (e.g., Mac and Spark) by running compute-bound parts (like prefill) on one device and memory-bound parts (like decode) on another. This requires a unified stack to maintain consistent KV cache layouts (Transcript, 0:11:30).

  • Hardware Comparison (Spark vs. Mac) 360s

    The DGX Spark offers high compute (e.g., 400 Tflops) and dedicated compute units (e.g., 4-bit units), while the Mac (M3 Ultra) offers high memory bandwidth (e.g., 800 GB/s). Optimal performance is achieved by combining these strengths (Transcript, 0:11:50).

Mentioned resources

  • Nemotron 3.5 Lightning (LLM Model)
  • DGX Spark (Hardware Accelerator)
  • DGX Station (Hardware Accelerator)
  • Exo Backend (Software Stack)
  • local.ai (Resource/Tool)
  • Brev Connect (Service/Tool)

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.