NVIDIA will sell a 64GB DGX Spark starting October 23, 2026, through Acer, ASUS, Dell, Gigabyte, HP and MSI, from $4,999. It carries half the unified memory of the 128GB model, and it costs more per gigabyte than the machine it is supposed to make affordable: the 128GB Founders Edition now lists at $6,950, up roughly 75 percent from its $3,999 launch price. That math is the clearest signal of what 2026 local AI actually looks like — memory as the binding constraint, not compute.
The more interesting part of the announcement is not the SKU. It is what ships with it: NVIDIA Sync Cluster Assistant and NVIDIA PAIR, the first time two machines on a home network are treated as one inference target by the vendor’s own software. For anyone running agents locally, that changes the shape of a build more than the gigabyte count does.
What NVIDIA Actually Shipped
The 64GB configuration is not a cut-down chip. It is the same GB10 Grace Blackwell superchip used in the 128GB model: a 20-core Arm Grace CPU and a Blackwell GPU with fifth-generation Tensor Cores on one module, rated at up to one petaflop of FP4 AI compute, with DGX OS and the full NVIDIA AI software stack preinstalled. What changed is the memory population.
- Memory bandwidth is unchanged at 273 GB/s. That tells you the partners are fitting lower-density LPDDR5X packages rather than narrowing the memory interface. The bus is the same; the stack is smaller.
- Single-system model ceiling drops. NVIDIA frames the 64GB box for today’s efficient open-weight models — ASUS, which announced its own Ascent GX10 64GB on October 2, describes it as sized for 30–35B-class models and autonomous agents. The 128GB system remains the one rated for frontier-scale weights.
- Storage is reduced alongside memory, which matters less than it sounds for inference but matters a lot if you keep several quantized checkpoints on disk.
- NVIDIA does not sell the 64GB unit directly. It exists only through the six OEM partners, so pricing is set per vendor and moves week to week.
The price history explains why this SKU exists at all. NVIDIA raised DGX Spark prices in March 2026, explicitly citing memory supply constraints. The 128GB model went from $3,999 at launch to $4,699, and then to $6,950. NVIDIA secured long-term HBM supply for its accelerators, but LPDDR5X — the memory in every consumer and workstation box — was not part of that protection, and neither was GDDR7. A 64GB model is what a market can support when the expensive component is the memory itself.
The Per-Gigabyte Math Is Inverted
Put the two SKUs side by side and the “budget” option is the expensive one per gigabyte:
- 64GB at $4,999 — about $78 per gigabyte of unified memory.
- 128GB at $6,950 — about $54 per gigabyte.
That inversion is not a pricing mistake; it is a supply signal. TrendForce expects conventional DRAM contract prices to rise another 10–15 percent quarter over quarter in Q4 2026, with NAND up 15–20 percent, and forecasts a 121 percent jump in blended HBM average selling prices in 2027. Micron has warned that demand will exceed supply through 2027 and 2028. When a vendor halves a component and raises the price, the component is the constraint — which is the same dynamic behind why HBM is eating your RAM on every consumer build.
Practical consequence: unified memory is now the most expensive part of a local AI box, and it is the part you cannot upgrade later. LPDDR5X is soldered. A machine bought at 64GB stays at 64GB.
Clustering Is the Actual Feature
Every DGX Spark, including the 64GB units, ships with an onboard ConnectX-7 network interface. Two units connect directly with a QSFP cable — no switch between them — over a 200GbE link. The NVIDIA Sync Cluster Assistant detects the second unit, validates configuration, and configures the network. Two 64GB boxes pool to 128GB of unified memory, double the memory bandwidth, and support models up to 200 billion parameters.
NVIDIA’s own number for the gain is up to 1.7x performance on a Qwen3.8 27B workload, measured against a single system. That is more than a doubling of capacity would give you on its own, and it is the honest reason to cluster: bandwidth, not just capacity. ASUS extends the same pattern upward — four 128GB units scale to 512GB and four petaflops.
The cost of that path is real. Two 64GB units are $9,998 before the cable, taxes, and shipping, against $6,950 for one 128GB box. You are paying a premium for bandwidth and for the option to grow. A single 200GbE direct-attach cable is the cheapest part of the whole build.
What PAIR Does Differently From a Cluster
Clustering pools memory into one target. NVIDIA PAIR (Personal AI Router) does something different, and it is the piece that matters for agent loops. PAIR is free, open-source software in beta for Windows, macOS and Linux. It discovers compatible machines over mDNS, pairs them with a six-digit code, talks over mutual TLS, and routes complete inference requests to whichever node has capacity. It does not split one request across several GPUs. A request lands on one node and stays there for its lifetime.
That distinction is the whole point for multi-agent work. A harness that fans out independent subagents has a queueing problem, not a model-size problem: five parallel requests hit one GPU one at a time. PAIR works with existing Ollama and LM Studio endpoints, so no agent harness changes are required — which is why it composes cleanly with the backend battle between vLLM, Ollama and llama.cpp rather than replacing it.
NVIDIA’s published demonstration is exactly this shape: a five-subagent workload driven from Hermes Desktop against Ollama, running Qwen 3.6 35B A3B. On a single RTX Spark laptop it averaged 18 minutes. On a three-device PAIR cluster — an RTX Spark laptop, a DGX Spark, and an RTX 5090 — the same workload averaged 8 minutes 48 seconds. Supported hardware includes GeForce RTX 20 Series and newer, RTX PRO workstation cards (Turing and newer), DGX Spark, and Apple M4 or newer silicon, so a mixed home network is allowed. NVIDIA has tested configurations up to 18 GPUs while expecting two-machine setups to be the common case.
The NVIDIA Sync Model Launcher, due at the end of October, downloads and launches a model across a single Spark or a clustered pair and exposes the endpoint to other machines on the network.
What This Means for a Local Build
If you already own hardware, PAIR is the cheapest upgrade available and requires no purchase. A machine you already run — an RTX 3090 desktop, a Mac mini, a second PC that sits idle — becomes a routing target. The gain is real only when the workload has independent requests to spread; a single long generation does not benefit.
If you are buying, the decision is arithmetic, not enthusiasm:
- Buy the 64GB unit only if 30–35B-class models are your actual workload. At $78 per gigabyte you are paying for the option to cluster, not for memory.
- Buy 128GB if you need frontier weights on one target. Per gigabyte it is still cheaper, and one node is simpler than two.
- Do not buy a Spark if CUDA is not part of your stack. A machine with cheap, upgradeable DDR memory and the same inference engine covers most agent workloads at a fraction of the cost — see the hardware tiers for running local LLMs before committing.
The pattern to watch is that vendor software is catching up to hardware economics. Memory is scarce and soldered; routing is free. The build that wins in 2026 is the one that treats several modest machines as one pool instead of betting on one expensive box.
Gear for a Two-Node Home Cluster
The DGX Spark 128GB listing on Amazon is the verified NVIDIA-branded SKU; the 64GB version is not on Amazon — it ships only through the six OEM partners from October 23, so treat any listing claiming a 64GB Spark as a third-party reseller. Street prices on the 128GB unit have ranged from about $3,950 to $9,400 against the current $6,950 list, so check the live listing before buying.
The interconnect is the cheap, easy part. A 200G QSFP56 passive DAC cable runs around $157 for the 1.5m version and about $169 for 1m — passive twinax, no optics, no switch, plug and play. It is the single component that turns two boxes into one cluster, and it is under 2 percent of the cost of the pair.
For the second node, an Apple Mac mini with M4 is the smallest PAIR-capable machine in the lineup: Apple M4 and newer silicon is supported, it draws little power, and it can take subagent requests while your main GPU is busy. It will not run a 30B model well on 16GB of unified memory, but as a routing target for smaller models and tool calls it earns its place. If you already have an RTX 5090 or an RTX PRO card, you do not need to buy anything — install PAIR on both machines and pair them.
What This Means
A 64GB machine that costs more per gigabyte than the 128GB machine is a market telling you where the shortage is. NVIDIA’s answer is software: cluster two boxes over a $157 cable, or route requests across whatever you already own. For readers who run agents locally, the takeaway is that capacity is no longer the scarce resource — bandwidth and parallel request handling are. Buy memory where it is cheap, and treat every machine on your network as part of the same inference target.

