NVIDIA RTX Spark superchip powering a compact Windows AI desktop with 128GB unified memory

NVIDIA RTX Spark: The 1-Petaflop Windows AI Superchip Ships This Month

NVIDIA’s RTX Spark PCs start reaching store shelves this month, and Microsoft is putting the spotlight on them at a joint Windows and Surface event in San Francisco on October 7. The platform’s headline claim is the reason this matters for anyone who runs AI on their own hardware: a single machine, roughly the size of a Mac mini or a slim laptop, with up to 128GB of unified memory, a petaflop of FP4 AI compute, native CUDA on Windows, and the ability to run 120-billion-parameter language models with million-token context entirely on the device. No cloud. No API bill. No rate limits.

What RTX Spark Actually Is

RTX Spark is NVIDIA’s first-generation N-series consumer superchip, built on the same Grace Blackwell architecture that powers the company’s data center racks, scaled down to a laptop and mini-PC package. It fuses an Arm-based Grace CPU and a Blackwell RTX GPU on one die pair, connected by the NVLink-C2C chip-to-chip interconnect at 600 GB/s, and shares one pool of LPDDR5X unified memory between them. NVIDIA published two official N1X configurations on September 3, 2026, and confirmed at IFA that OEM shipments begin in October, rolling out region by region:

  • N1X high-end: 20-core Grace CPU, 6,144-core Blackwell RTX GPU with fifth-generation FP4 Tensor Cores, unified memory from 24GB up to 128GB, ships in laptops and compact desktops, 45-80W power envelope, rated up to 1 petaflop of FP4 AI compute.
  • N1X second tier: 18-core Grace CPU, 5,120-core Blackwell GPU, unified memory capped at 24GB or 32GB, laptop-only at launch, with compact desktops expected later.
  • RTX Spark N1 (the budget variant): 10 or 12-core Grace CPU, up to 2,048 CUDA cores, up to 64GB unified memory, 18-45W, aimed at thinner and cheaper laptops that still carry the RTX Spark name.

All of it runs Windows 11 on Arm, and CUDA — the software stack that accelerates most of the world’s AI — runs natively on the chip. The launch roster is ASUS, Dell, HP, Lenovo, Microsoft Surface, and MSI, with Acer and GIGABYTE confirmed to follow; NVIDIA’s own site points to availability from those six OEMs starting Friday, October 23, alongside a new NVIDIA Sync Cluster Assistant for scaling workloads across machines.

Why Unified Memory Beats VRAM for Local Models

The bottleneck for running large models on consumer hardware has never been compute — it has always been memory. A consumer GeForce tops out at 16GB of VRAM on the RTX 5070 Ti, 24GB on the 5080, and once a model’s weights exceed that pool, you either shard across multiple GPUs, quantize aggressively, or you don’t run the model at all. Unified memory changes the arithmetic. A 120B-parameter model quantized to 4-bit needs roughly 60GB for weights; with 128GB of unified LPDDR5X, that model fits with room left over for a long context window, the KV cache, and the OS. NVIDIA’s official claim is 120B-parameter LLMs with up to 1 million tokens of context using agents locally, and its developer documentation goes further: two RTX Spark units linked over NVLink-C2C share memory into an effective 256GB pool, putting 200B+ parameter models — server-rack territory — on two desktop boxes.

The Realistic Performance Picture

One honest caveat before the hype: unified LPDDR5X has far less memory bandwidth than GDDR7 on a discrete GPU, and token generation speed on large models is bandwidth-bound. Expect fast prefill (the GPU compute is genuinely petaflop-class) and modest decode on the biggest models — single-digit to low-double-digit tokens per second on a 120B model is the realistic band, based on measurements from the GB10-based DGX Spark, which uses the same memory architecture. Small models (7B-8B) run fast. The machine is a model host, not a token cannon, and the right expectation is agentic workloads — long-running, background, tool-using agents — where a few tokens per second sustained is exactly what you need.

CUDA on Windows Is the Real Headline

Spec sheets aside, the strategically important line in NVIDIA’s announcement is that CUDA runs natively on RTX Spark. The local-AI ecosystem has spent the last two years splitting into two camps: the CUDA world (vLLM, TensorRT, CUDA-optimized llama.cpp builds, quantization kernels, GPU-accelerated agent tooling) and the Apple Metal world (MLX, Ollama on Mac, Apple Intelligence). Windows machines with NVIDIA discrete GPUs covered part of the CUDA camp, but only with small VRAM. A Windows machine with 128GB unified memory and native CUDA closes the last structural gap: the same tooling that runs on data center hardware now runs, unmodified, on a consumer desktop. TensorRT, RTX, DLSS, OptiX, Reflex, and G-SYNC come along for the ride, which is NVIDIA’s pitch that one box does AI development, 12K video editing, 4K AI video generation, and AAA gaming at 1440p.

The Microsoft Angle: Surface Laptop Ultra and the Dev Box

Microsoft’s October 7 event is the consumer-facing half of the story. Two products are in the window:

  • Surface Laptop Ultra: announced May 31, Microsoft’s most powerful Surface Laptop ever — a 15-inch mini-LED PixelSense Ultra touchscreen with up to 2,000 nits peak HDR, under 18mm thick, under 4.5 pounds, a replaceable SSD, up to 128GB unified memory, full CUDA support, and HDMI, USB-C, USB-A, SD card, and headphone ports. It has sat in pre-release status with no price and no ship date; October 7 is the obvious moment to fix that.
  • Surface RTX Spark Dev Box: shown at Build on June 2, a compact Windows 11 Pro desktop with the same superchip, 128GB of unified memory, and a 100W thermal envelope, pitched at developers prototyping, fine-tuning, and running their own models locally. Microsoft says it will sell in the US exclusively on Microsoft.com, and it has not yet cleared FCC authorization — a sign the launch is genuinely tight.

Windows Central reports the entry configuration pairs the 18-core CPU with 5,120 GPU cores and 24GB or 32GB of memory, while higher-end configurations get the 20-core/6,144-core chip with 32GB, 48GB, 64GB, or 128GB, shipping with Windows 11 Pro. The outlet expects starting prices well above $2,000.

Pricing: What the Numbers Say So Far

NVIDIA and its OEMs have not published official MSRPs, and that silence is the biggest open question of the launch. The concrete signals so far:

  • A Morgan Stanley analyst note pegs the high-end N1X configuration at roughly $2,899, inside the $2,500-$3,000 band analysts have floated.
  • The N1 tier is estimated at $1,799 and up, and Tom’s Guide speculates $1,999-$2,499 for lower-tier systems — explicitly guesses, not NVIDIA numbers.
  • Channel reports say ASUS and MSI’s first production batches are already allocated to channel partners, so early stock will be scarce.
  • The reference point is the existing DGX Spark — the GB10 Grace Blackwell developer workstation that launched at about $3,000 with Linux. RTX Spark is its consumer Windows derivative, and the price band lands in the same neighborhood.

Hardware You Can Buy Today for Local AI

RTX Spark machines are not on retail shelves yet, so there is nothing to link for them — if you want a 128GB-class CUDA box, you are on a pre-order list, not a purchase. If you want to run local models and agents this week, these are the proven lanes, all verified Amazon listings:

The Apple Mac mini with M4 Pro (24GB unified, 512GB SSD) is the closest current analog to the RTX Spark design philosophy — shared unified memory, low power, silent — and at $1,449.99 list (often around $1,279 with promos) it undercuts the rumored Spark price by half; the M4 Pro line scales to 64GB, which handles 30B-class models at 4-bit, and the same listing family carries higher storage tiers. For a tighter budget, the Beelink SER9 mini PC (Ryzen AI 9 HX 370, 32GB LPDDR5X, 1TB) runs at $609 against a $749 list price and runs 8B-14B models comfortably through Ollama or llama.cpp today. If you build your own desktop, the discrete-GPU route is the ASUS Prime RTX 5070 Ti with 16GB GDDR7 — small VRAM, but fast decode for models that fit — and pairing it with the G.SKILL Trident Z5 Neo 64GB (2x32GB) DDR5-6000 kit gives you a 64GB system RAM pool for CPU-offloaded inference on larger quantized models.

What This Means

The data center side of the industry is spending gigawatts — Meta and OpenAI have each committed up to 6GW to AMD’s Instinct racks, Anthropic up to 2GW, and hyperscaler power contracts now read like utility infrastructure deals. The edge side is where individuals live, and RTX Spark is the first Windows machine that takes local model hosting seriously: CUDA native, 128GB of unified memory, million-token context, and a petaflop of FP4 in a box that fits on a desk. The practical read: if your workloads are 7B-30B models, buy today — the Mac mini and mini-PC lanes deliver most of the value at a third of the price. If you need 100B+ models with long context on one machine, the October 7 pricing reveal is the decision point, and the $2,899 estimate means these machines will be a serious purchase, not an impulse one. Watch the event for real MSRPs, then decide whether you buy the shipping hardware or the shipping window.

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *