NVIDIA DGX Spark 64GB Gives Developers More Ways to Build and Scale Local AI
NVIDIA Blog Allen Bourgoyne
NVIDIA’s DGX Spark is getting a 64GB version from its hardware partners. It runs local AI on device, and two boxes can team up when the work gets bigger.
Based on reporting by NVIDIA Blog, Allen Bourgoyne — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
NVIDIA is widening the DGX Spark line with a 64GB model that arrives this month through Acer, ASUS, Dell, Gigabyte, HP and MSI. The pitch is simple: more local AI, less cloud dependence, and a setup that is supposed to be useful the moment it comes out of the box.
This version keeps the same core stack as the 128GB model, including the GB10 Grace Blackwell Superchip, DGX OS and NVIDIA’s AI software bundle. That matters because NVIDIA is treating DGX Spark less like a tinkerer’s box and more like a compact AI workstation for agents, inference, fine-tuning, data science and edge work. It supports models up to 100 billion parameters on a single system, fully on device.
When one box isn’t enough, two can be linked through NVIDIA Sync Cluster Assistant. The company says the pair can pool memory to 128GB and support models up to 200 billion parameters, with the same software stack on each node so nothing needs to be reconfigured when scaling from one machine to two. In NVIDIA’s Qwen 3.8 27B test, two clustered 64GB systems reached up to 1.7x the performance of a single system.
NVIDIA is also leaning hard into the “ready now” message. DGX Spark ships with support for NVIDIA Agent Toolkit, CUDA-X AI libraries, Nemotron open models, and runtimes like Ollama, vLLM and PyTorch with CUDA. Blender is set to be among the first major creator apps to support it, with a prebuilt installer coming soon. The company says developers can go from power-on to running models in minutes.
The 64GB configuration goes on sale Friday, Oct. 23, starting at $4,999. NVIDIA points users to supported frameworks such as llama.cpp, Ollama, vLLM and LM Studio, then tells them to connect two units with their ConnectX-7 ports if the workload grows beyond one machine.
My take — AI-written commentary, not fact-checked reporting
The real story here isn’t the box, it’s NVIDIA trying to make local AI feel normal instead of heroic. That’s the right bet: most developers want fewer cloud bills and fewer excuses, not another demo about the future. Funny how “just run it locally” keeps turning into a hardware strategy.
Read more about this at: NVIDIA Blog