Nvidia announced Project DIGITS at CES in January 2025 and brought it to market as the Nvidia DGX Spark: a palm-sized desktop computer built around the GB10 Grace Blackwell superchip with 128 GB of unified memory. It went on sale in mid-October 2025 at $3,999, according to Engadget. This guide explains what the hardware is, what size of AI model it can hold, why its memory speed matters more than its headline compute figure, and who should buy one. Every specification below comes from Nvidia’s DGX Spark page; speed estimates are labelled as arithmetic, not benchmarks.
Key takeaways
- DGX Spark is the shipping name of Project DIGITS. GB10 Grace Blackwell superchip, 20-core Arm CPU, 128 GB of LPDDR5x unified memory, 4 TB SSD, in a 150 x 150 x 50.5 mm box.
- Capacity is the selling point. Nvidia rates one unit for inference on models up to 200 billion parameters and fine-tuning up to 70 billion; up to four linked units handle models up to 700 billion.
- Memory bandwidth is the limit. At 273 GB/s it holds far larger models than a 32 GB gaming GPU, but generates tokens on large dense models much more slowly than a discrete card could.
- Price has moved. $3,999 at launch in October 2025, with reports of $4,699 by February 2026 as memory costs rose. Partner versions from ASUS, Acer, Dell and MSI use the same chip.
What DGX Spark actually is
Nvidia’s published specification:
“Unified memory” means the CPU and GPU share one 128 GB pool, the same idea Apple uses in its M-series Macs. A gaming PC splits memory between system RAM and the graphics card’s own VRAM, and a model that does not fit in VRAM either will not load or slows to a crawl. On DGX Spark the whole pool is available to the model.
What size of model fits
A model’s memory footprint is roughly its parameter count multiplied by the bytes stored per parameter, plus room for the context (the KV cache) and runtime overhead. At 4-bit quantization each parameter takes about half a byte:
| Model size | Weights at 4-bit (approx.) | One DGX Spark (128 GB) | RTX 5090 (32 GB) |
|---|---|---|---|
| 8B | ~5 GB | Fits easily | Fits easily |
| 32B | ~18 GB | Fits easily | Fits, limited context |
| 70B | ~40 GB | Fits with room for long context | Does not fit |
| 120B | ~65 GB | Fits | Does not fit |
| 200B | ~110 GB | Fits (Nvidia’s rated maximum) | Does not fit |
| 405B | ~220 GB | Needs linked units | Does not fit |
Approximate weight sizes; exact figures depend on the quantization format. Our VRAM calculator works out the memory for a specific model, quantization and context length, and the VRAM requirements guide lists popular models.
Why memory bandwidth decides the speed
When a language model writes each token, the hardware has to read the model’s active weights out of memory. For large models that read, not raw compute, is usually the bottleneck. That makes memory bandwidth a useful way to estimate the upper limit of generation speed: divide the bandwidth by the gigabytes read per token.
| Hardware | Memory bandwidth | Ceiling for a dense 70B model at 4-bit (~40 GB) | Ceiling for a dense 8B model at 4-bit (~5 GB) |
|---|---|---|---|
| Nvidia DGX Spark | 273 GB/s | ~7 tokens/s | ~55 tokens/s |
| Apple M4 Max | 546 GB/s | ~14 tokens/s | ~110 tokens/s |
| RTX 5090 | ~1,792 GB/s | Model does not fit in 32 GB | ~360 tokens/s |
These are arithmetic ceilings from published bandwidth figures, not measured results. Real-world speeds come in lower once compute, software and context length are accounted for.
Two practical conclusions follow. First, for models small enough to fit on a gaming GPU, a discrete card is much faster. Second, DGX Spark is at its best with mixture-of-experts (MoE) models, which store many parameters but activate only a fraction for each token. A large MoE model needs the Spark’s 128 GB to load, yet reads only its active weights per token, so it can generate far faster than a dense model of the same total size.
Who DGX Spark is for
- Developers who need big models locally. If your work involves models in the 70B to 200B range and your data cannot leave the building, a single box that holds them is the core appeal.
- Teams building for Nvidia’s data-centre stack. DGX OS ships with Nvidia’s AI software, so code developed on a Spark moves to larger DGX and cloud systems with little change.
- Fine-tuning experiments. Nvidia rates it for fine-tuning models up to 70 billion parameters, which is out of reach for a single consumer graphics card.
Who should look elsewhere
- People running models that fit in 24 to 32 GB. A desktop with a high-end graphics card will generate tokens considerably faster. Compare cards in our best GPUs for local LLMs guide.
- Anyone wanting a general-purpose computer. DGX Spark is designed as an AI development machine, not a home or gaming PC.
- Occasional users. If you only need a large model now and then, renting cloud GPUs or paying per token through an API is usually cheaper. The self-hosting vs API calculator shows where the break-even sits.
Pros and cons
DGX Spark strengths
- 128 GB unified memory holds models up to 200B parameters (Nvidia rating)
- Nvidia AI software stack on DGX OS
- Low power next to multi-GPU workstations: 140 W chip TDP, 240 W supply
- Up to four units can be linked for models up to 700B parameters
- Small and light at 1.2 kg
DGX Spark weaknesses
- 273 GB/s bandwidth limits speed on large dense models
- Slower than a discrete GPU for models that fit in its VRAM
- Memory is soldered; no upgrades
- Price has risen since launch along with memory costs
What it actually costs to own one
The sticker price is only the start of the decision, and it has been a moving target. Nvidia’s Founders Edition launched at $3,999 in October 2025, and was reported at $4,699 by February 2026 as the global DRAM shortage made its 128 GB of LPDDR5x more expensive to build. Because so much of the cost is soldered-on memory, DGX Spark pricing tracks the memory market and can move again, so check the Nvidia Marketplace for the current figure.
You are not limited to the gold Founders box. Partner units from ASUS (Ascent GX10), Acer (Veriton GN100), Dell (Pro Max GB10) and MSI (EdgeXpert) use the same GB10 superchip and the same 128 GB of unified memory. The trade-off is usually storage: partner machines often ship with a 1 TB SSD instead of the Founders Edition’s 4 TB, which is how the cheaper ones come in below Nvidia’s price.
Then add the running costs people forget:
- Storage you will outgrow. Model weights, datasets and container images are large. On a 1 TB unit, budget for fast external NVMe almost immediately.
- Electricity. It draws little power next to a tower workstation, but an always-on inference box still runs up a bill.
- The second unit. Linking two units for models around 405 billion parameters roughly doubles the hardware cost.
The honest comparison is against the cloud, not against nothing. A rented high-end GPU instance bills by the hour and never depreciates in your cupboard. DGX Spark only wins on total cost when it stays busy: continuous local inference, regular fine-tuning, or workloads where keeping data on-premises is the point.
Verify these figures at the source
Specifications and prices change without notice. These are the pages the figures above come from:
FAQ
Is Nvidia DGX Spark the same as Project DIGITS?
Yes. Project DIGITS was the name Nvidia used when it announced the GB10-based personal AI computer at CES 2025. It went on sale as DGX Spark in October 2025.
How much does DGX Spark cost?
It launched at $3,999 in October 2025, and reports put the Founders Edition at $4,699 by February 2026 as memory prices rose. Partner models from ASUS, Acer, Dell and MSI with smaller SSDs can cost less.
What is the largest model DGX Spark can run?
Nvidia rates a single unit for inference on models up to 200 billion parameters and fine-tuning up to 70 billion. Linking up to four units raises the inference limit to models of up to 700 billion parameters.
What is the memory bandwidth, and why does it cap performance?
DGX Spark’s 128 GB of unified LPDDR5x runs at 273 GB/s. That figure, not the headline petaFLOP of FP4 compute, limits token generation speed on large language models, because each token requires reading the active weights from memory. It is enough to load very large models that will not fit on a 24 to 32 GB gaming GPU, but well below the bandwidth of a discrete card like the RTX 5090.
Should I buy the Founders Edition or a partner unit like the ASUS Ascent GX10?
The compute is the same, so the choice comes down to storage and price. The Founders Edition includes a 4 TB SSD; many partner units ship with 1 TB and cost less. If you will keep many large models and datasets locally, the extra storage can be worth paying for.
Can I use DGX Spark as a normal desktop computer?
It runs Nvidia DGX OS, which is based on Ubuntu Linux, so everyday desktop software works. It is designed and priced as an AI development machine, though, and a conventional PC is far better value if you do not need its memory capacity.
