Key takeaways
- NVIDIA's DGX Spark ships with 128GB unified memory and runs inference on models up to 200 billion parameters entirely offline.
- Priced at $4,699, it's the first single-box desktop system to bring data-center-class AI compute to individual developers.
- The real significance is structural: for the first time, serious LLM workloads don't require a cloud API call.
- Apple Silicon and NVIDIA are pursuing fundamentally different visions of what the AI computer looks like one optimized for efficiency, one for raw model scale.
- Privacy, API cost, and agent latency are the three forces pulling enterprise AI toward local inference.
When NVIDIA announced DGX Spark at GTC in March 2025 originally codenamed Project DIGITS most coverage framed it as a gaming or consumer hardware story. That's the wrong frame. DGX Spark is not a faster graphics card. It's a new category of device: a personal AI supercomputer built from the same Blackwell architecture that powers NVIDIA's data center empire, shrunk into a box smaller than a textbook.
The key specification to understand is not clock speed or gaming FPS. It's memory. DGX Spark carries 128GB of unified CPU and GPU memory the same pool, accessible at once. That single number determines what you can run. With 128GB, you can load a full 70B parameter model at BF16 precision with no quantization, or push to 405B models at Q2–Q3 quantization. You can run two 30B models simultaneously. You can do things that, until late 2025, required a rack of A100s.
NVIDIA describes it plainly: "1 petaFLOP of AI performance in a compact desktop form factor." That's 1,000 trillion operations per second. Delivered by a GB10 Grace Blackwell Superchip with fifth-generation Tensor Cores, FP4 support, and a 20-core ARM CPU.


