# Local AI: Running Nemotron on DGX Spark & Station | Nemotron Labs

## Executive summary

The session details the growing viability of local AI, emphasizing that running Large Language Models (LLMs) on dedicated local hardware like DGX Spark and DGX Station offers significant advantages in sovereignty, privacy, and cost efficiency compared to cloud-only inference. Key technical advancements include the 'Exo backend,' a unified software stack that enables consistent model deployment across diverse hardware (Nvidia and Apple silicon). The discussion also covered advanced scaling techniques, such as Process/Data (PD) Disaggregation, allowing compute-intensive parts of inference to run on one device while memory-intensive parts run on another, maximizing utilization of heterogeneous local compute resources.

## Key takeaways

- Local AI Value Proposition: Local AI provides sovereignty, privacy, and cost efficiency by eliminating per-token cloud costs. For single-user, low-concurrency workloads, local inference can achieve high tokens per second, matching or exceeding the performance of cloud serving setups (Transcript, 0:02:50).
- Unified Local AI Stack (Exo Backend): The Exo backend addresses the fragmentation problem in local AI by providing a unified stack that runs across different hardware targets, including Nvidia hardware and Apple silicon (M3 Ultra). This ensures that models can be run consistently regardless of the underlying device (Transcript, 0:07:20).
- Advanced Scaling and Hardware Utilization: DGX Spark and DGX Station support both horizontal (clustering) and vertical scaling. DGX Station, with up to 750 GB of unified VRAM, is recommended for running very large open-source models (e.g., GLM 5.3) that exceed the capacity of a single Spark unit (Transcript, 1:17:00).

## Technical details

- Inference Economics (Local vs. Cloud): Cloud serving optimizes for high throughput per dollar via batching, often sacrificing individual user interactivity. Local AI, by contrast, has a marginal cost of generating a token close to zero (only power cost), making it economically superior for self-contained workloads (Transcript, 0:02:10).
- Process/Data (PD) Disaggregation: PD Disaggregation allows splitting inference tasks between different accelerators (e.g., Mac and Spark) by running compute-bound parts (like prefill) on one device and memory-bound parts (like decode) on another. This requires a unified stack to maintain consistent KV cache layouts (Transcript, 0:11:30).
- Hardware Comparison (Spark vs. Mac): The DGX Spark offers high compute (e.g., 400 Tflops) and dedicated compute units (e.g., 4-bit units), while the Mac (M3 Ultra) offers high memory bandwidth (e.g., 800 GB/s). Optimal performance is achieved by combining these strengths (Transcript, 0:11:50).

## Practical implications

- For engineering teams, the choice between DGX Spark and DGX Station depends on the model size and required VRAM capacity; Station is preferred for the largest open-source models.
- When designing local AI workflows, utilize unified stacks (like Exo) to ensure portability and reduce development complexity across different hardware types (e.g., Mac, Spark).
- For high-performance local deployments, consider combining compute-bound and memory-bound tasks using PD Disaggregation to maximize the utilization of heterogeneous compute resources.
- The ability to cluster multiple DGX units (Sparks or Stations) allows for scaling beyond the limits of a single device, enabling the deployment of frontier-level models in private, air-gapped environments.

## Topics

Local AI, LLM Inference, Edge Computing, Hardware Acceleration, Distributed Systems, Nemotron 3.5 Lightning, DGX Spark, DGX Station, Exo Backend, Brev Connect

Source: https://www.youtube.com/watch?v=WWp_RFBn34M
