Microsoft Puts Local AI on Windows From $2,599, and the Memory Tier Sets What Runs

Compute & Hardware

Microsoft's RTX Spark hardware removes the Linux step from local artificial intelligence, or AI. The memory tier you buy and the model your agent picks still need an owner.

By Shashi Bellamkonda · October 8, 2026
$2,599
Starting price of the entry configuration, which carries 24 gigabytes of memory (Windows Central, 2026)
53
Gigabytes for Microsoft's compressed local coding model (Microsoft, 2026)
75.5
Gigabytes of peak memory at a 256,000-token context window (Microsoft, 2026)
$5,999
Price of the 128-gigabyte Dev Box, shipping in November (The Gadgeteer, 2026)
Key Takeaway
Windows removes the Linux step, and agent installers now handle model setup. The memory tier decides which models fit. Fleet administration of the model and the sandbox policy is the open item.

Linux is no longer the price of entry for local AI on an NVIDIA desktop, and Microsoft put numbers on that on October 7. The Surface Laptop Ultra starts at $2,599 and ships October 16, while the Surface RTX Spark Dev Box costs $5,999 and ships in November (The Gadgeteer, 2026). The Dev Box boots into Windows 11 Pro with Visual Studio Code and GitHub Copilot already installed, and Windows Subsystem for Linux sits one layer down for anyone who wants it.

NVIDIA's DGX Spark arrives configured differently. It ships with DGX OS, which NVIDIA builds on Linux, plus NVIDIA's development libraries and container tools, so a developer works inside that Linux environment from the first boot. DGX Station for Windows, which NVIDIA lists for the fourth quarter, reaches Linux workloads through the same Windows Subsystem for Linux layer.

Nothing in the launch material I reviewed lists a language model among the software that comes preinstalled.

Agent Installers Now Set the Model, and Each One Brings Its Own

Setup friction moved to the application layer. Hermes Agent from Nous Research offers one-click local setup on Windows: it detects the installed NVIDIA graphics processing unit, or GPU, and picks a model and configuration to run through an integrated inference engine, according to IT Brief UK's September report. OpenClaw added a simplified Windows setup for local models on RTX GPUs with at least 24 gigabytes of video memory, and Microsoft's new Get Started experience for agents on Windows features Hermes and OpenClaw among its agents.

GitHub Copilot adds a second route by the end of October. Its Auto mode will decide task by task whether a local or a cloud model does the work, and a developer can also select a local model by hand through Windows ML or a local endpoint of their own.

Both routes remove a manual step, and each puts a different company in charge of the model on the device.

In a one-click install, the agent chooses the model, and your team inherits the choice.
Key Takeaway
Each agent installer brings its own model and its own update schedule. Ask every vendor in your pilot for both in writing before you standardize on one.

The Model in Microsoft's Demo Needs a Memory Tier Above the Entry Price

Microsoft built a local version of MAI Code 1.1 Flash, a coding model with 137 billion total parameters, 6.8 billion of them active for each token it generates. Compressing the model to lower numerical precision cut its size by about 80 percent, to 53 gigabytes, and it peaked at 75.5 gigabytes of memory at a 256,000-token context window (Microsoft, 2026). In Microsoft's October 5 test, the compressed model scored 70.8 percent on SWE-Bench Verified, a software engineering benchmark, against 72.6 percent for the cloud version (Microsoft, 2026).

Memory sets the limit. Unified memory is one pool, so Windows, your applications, and the working memory an agent builds as it reads files draw from the same capacity as the model weights. Laid against the laptop configurations Windows Central lists, Microsoft's figures sort the tiers this way, and the fit notes are my arithmetic rather than a Microsoft statement:

  • 24 gigabytes, $2,599 (Windows Central, 2026): well below the 53-gigabyte size of the model weights.
  • 32 gigabytes, $3,699.99 (Windows Central, 2026): below it as well.
  • 128 gigabytes, $5,899.99 for the laptop and $5,999 for the Dev Box (Windows Central, 2026; The Gadgeteer, 2026): clears the 75.5-gigabyte full-context peak with 52 gigabytes to spare.

The $2,599 price buys the platform. On the configurations I could confirm, the only one that holds Microsoft's demonstrated coding model at full context is the 128-gigabyte tier at $5,899.99, so price that tier before you count seats. The Dev Box lists at $100 more than that laptop and puts the same 128 gigabytes on a desk (The Gadgeteer, 2026).

Software Arrives in Stages After the October 16 Ship Date

The rollout is staged, and the sequence matters for pilot planning:

  • October 15: NVIDIA's new Nemotron model is due (Windows Report, 2026).
  • October 16: Surface Laptop Ultra ships.
  • End of October: GitHub Copilot's local-versus-cloud routing arrives (Microsoft, 2026).
  • November: the Surface RTX Spark Dev Box ships.
  • Fourth quarter: DGX Station for Windows, with no price in NVIDIA's announcement.

A laptop pilot that starts the week of October 16 exercises the agent installers. Copilot's routing layer arrives two weeks later, so a scorecard written in the first week covers an incomplete stack.

Execution Containers Name Manageability, and the Console Question Is Open

Microsoft Execution Containers, or MXC, is an open-source library from the Windows team that translates policy into native operating system controls for AI agents. Microsoft frames it around three principles: containment, identity, and manageability (Windows Report, 2026).

GitHub Copilot exposes MXC through a "Sandbox new sessions" toggle that a developer sets per project and through a /sandbox command in its command line tool, so sandbox settings begin at the developer's keyboard. NVIDIA says DGX Station for Windows extends existing Windows fleet management to the machine, with agents governed through familiar Microsoft tools.

I raised the same question about NVIDIA OpenShell in June, when I asked whether its policy management would fold into Microsoft Intune.

Manageability is a stated principle, and in the material I reviewed, no page names the console that sets MXC policy across a fleet, mentions Intune or Entra ID for it, or says whether a developer can override it. That answer may sit in documentation I did not reach, which is why the question belongs in your vendor review.

CIO/CTO Viability Question
Before you approve a purchase order for RTX Spark laptops, ask Microsoft to name in writing the console that sets the local model and the MXC sandbox policy for your fleet, and whether a developer can override either. Then price the memory: the coding model Microsoft demonstrated needs the 128-gigabyte tier, listed at $5,899.99 for the laptop. If the answer is that each agent vendor manages its own model, who in your organization owns that model on the day it changes?
Sources

Hunt, Cale. "Surface Laptop Ultra with RTX Spark Preorders Are Live: Here's What You Need to Know about Pricing, Availability, and More." Windows Central, 7 Oct. 2026, windowscentral.com.

Microsoft. "Windows and Surface October 2026 News." Microsoft News, 7 Oct. 2026, news.microsoft.com.

Windows Report. "Microsoft October 2026 Surface and Windows Event: All the Biggest Announcements." Windows Report, 7 Oct. 2026, windowsreport.com.

Mitchell, Sean. "Nvidia & Microsoft Push Local AI Agents on Windows." IT Brief UK, 4 Sept. 2026, itbrief.co.uk.

Nguyen, Vincent. "Surface Laptop Ultra Starts at $2,599 as Microsoft Puts More AI Work on the PC." The Gadgeteer, 7 Oct. 2026, the-gadgeteer.com.

Nikoletich, Patrick, and Stuart Schaefer. "Bringing Local Models and Sandboxed Tools to Windows and GitHub Copilot." Command Line, Microsoft, 7 Oct. 2026, commandline.microsoft.com.

NVIDIA. "NVIDIA DGX Station for Windows Puts a Trillion-Parameter AI Supercomputer on Every Enterprise Desk." NVIDIA Newsroom, 31 May 2026, nvidianews.nvidia.com.

NVIDIA. "System Overview." DGX Spark User Guide, NVIDIA Documentation, docs.nvidia.com.

Disclaimer: This blog reflects my personal views only. Content does not represent the views of my employer, Info-Tech Research Group. AI tools may have been used for brevity, structure, or research support. Please independently verify any information before relying on it.