Surface Laptop Ultra Starts at $2,599, but Battery Life Is Still the Big Unknown

Microsoft’s RTX Spark Surface Laptop Ultra can run 120B-parameter AI models locally, but its $2,599 price, Arm compatibility and undisclosed battery life deserve scrutiny.
Microsoft Surface Laptop Ultra laptop with NVIDIA RTX Spark

Microsoft has opened preorders for the Surface Laptop Ultra, a $2,599.99 Windows notebook built around Nvidia’s new RTX Spark system-on-chip. The 15-inch machine ships October 16 and targets developers, creators and researchers who want to run large AI models locally instead of paying for every inference in the cloud.

The headline specifications are unusual for a portable PC: an Arm-based Nvidia Grace CPU, a Blackwell RTX GPU, unified memory that can scale to 128GB, and up to one petaflop of theoretical FP4 AI performance. Microsoft says the top configuration can run models exceeding 120 billion parameters on the device. Those numbers give the Ultra a credible use case that most “AI PCs” do not have, but they do not settle the most practical buying questions. Microsoft has not yet published a battery-life estimate, and its performance comparisons come from vendor-controlled tests on preproduction hardware.

Surface Laptop Ultra price, release date and specifications

Preorders began October 7, with retail availability scheduled for October 16. The entry configuration starts at $2,599.99 and uses an 18-core RTX Spark chip with 5,120 GPU cores and 24GB of unified memory. Higher configurations move to a 20-core Grace CPU, 6,144-core Blackwell GPU and as much as 128GB of shared memory. Storage tops out at 2TB, and Microsoft designed the SSD to be removable and replaceable.

The unified-memory design is central to the pitch. In a conventional performance laptop, system RAM and dedicated GPU memory are separate pools. RTX Spark lets the CPU and GPU draw from one pool, so a large model is not constrained by the fixed VRAM capacity of a discrete mobile GPU. The operating system and applications still consume part of that memory, meaning a 128GB configuration cannot devote the entire 128GB to a model.

Microsoft Surface Laptop Ultra open on a desk showing its 15-inch display
Surface Laptop Ultra has a 15-inch touchscreen, three USB-C ports, USB-A, HDMI, an SD card reader and a headphone jack. Image: Microsoft.

The aluminum chassis weighs less than 4.5 pounds and measures under 18mm thick. Its 15-inch PixelSense Ultra touchscreen runs at up to 120Hz, covers 100% of the DCI-P3 color space and reaches a claimed 2,000 nits of peak HDR brightness in a 10% window. That last figure is not sustained full-screen brightness, an important distinction for anyone comparing displays for editing or outdoor use.

Port selection includes three USB-C connections, USB-A, HDMI, a full-size SD card reader and a 3.5mm audio jack. One USB-C port doubles as Microsoft’s new Magnetic Connect charging interface: the included cable detaches if pulled, while the port still accepts ordinary USB-C data, video and charging accessories. The laptop can drive as many as three 4K external displays.

What local AI performance can mean in practice

Nvidia rates RTX Spark at up to one petaflop of FP4 compute, but that is a low-precision theoretical figure rather than a promise that every model or application will reach that throughput. The practical advantages are memory capacity, CUDA support and the ability to keep prompts, source code and proprietary data on the PC.

Microsoft is adapting its software stack around those capabilities. Windows ML is adding llama.cpp support, while GitHub’s HydraFusion routing will be able to send some Copilot tasks to a local model and others to the cloud. An experimental preview for the GitHub Copilot app, Copilot CLI and Visual Studio Code is due later in October. Microsoft also plans local support for a 3-bit version of its 137-billion-parameter MAI Code 1.1 Flash model, with a 256K context window, alongside optimized Nvidia Nemotron and DeepSeek models.

That can reduce cloud API spending during repetitive coding, evaluation and content-generation work. It can also help teams experiment with sensitive data without uploading every input. Local processing is not automatically private, however: the application, extensions, model files and any connected tools still determine where data travels.

Microsoft reports that selected 64GB RTX Spark PCs produced up to 2.1 times faster time to first token than a 64GB MacBook Pro with M5 Pro in a llama.cpp test using a quantized Qwen 3.5 27B model. Nvidia reported 4.3 times faster AI image generation with FLUX.2 Klein 4B and 6.2 times faster video generation with LTX 2.3. These are narrow, vendor-commissioned tests using specific precision settings and preproduction Windows systems. They should not be read as a general verdict on application speed, efficiency or battery life.

Windows on Arm remains part of the buying decision

RTX Spark brings Nvidia’s CUDA software ecosystem to Windows on Arm, which matters for developers whose AI and creative tools already depend on CUDA libraries. It does not make every existing Windows application native to the architecture. Buyers should check the exact versions of their editors, plug-ins, device drivers, virtual private network clients and specialist engineering tools before ordering.

Emulation can cover many x86 applications, but compatibility and performance can vary, especially with low-level drivers and older plug-ins. The same caution applies to games. Nvidia says RTX Spark systems can exceed 100 frames per second at 1440p in supported titles using DLSS, Reflex and G-Sync, and Microsoft demonstrated Gears of War: E-Day. That is evidence that gaming is possible, not proof that the Ultra behaves like a conventional GeForce gaming laptop across a broad library.

The unanswered battery-life question

Microsoft has emphasized sustained performance, saying the redesigned cooling system offers up to 2.5 times the thermal capacity of the current 15-inch Surface Laptop. The Ultra also ships with a 140-watt adapter. Yet the company has not disclosed an estimated runtime and says battery details will arrive closer to availability.

That omission matters more here than on a desktop-replacement gaming machine. Local inference can keep the GPU and memory subsystem busy for long periods, and a laptop bought specifically to avoid the cloud needs to remain useful away from an outlet. Independent tests should measure ordinary web and office use, sustained token generation, image generation, fan noise and performance on battery. A single “up to” figure will not capture those different workloads.

Who should preorder, and who should wait

The clearest early audience is a developer or creator who already knows why 64GB or 128GB of shared memory and CUDA support are valuable. Teams prototyping local agents, evaluating large open-weight models, handling sensitive data or moving the same workflow between a laptop and Nvidia servers may be able to justify the price.

Most buyers should wait for reviews. The $2,599 base model’s 24GB memory capacity is far less distinctive than the higher configurations, while the cost of those upgrades, real battery life and broad application compatibility will determine whether this is a practical mobile workstation or a transportable AI development system. The hardware is one of the first Windows laptops with enough memory and GPU support to make serious local-model work plausible. Whether it is a good laptop as well remains the part Microsoft has not yet demonstrated.

Sources

Leave a Reply

Your email address will not be published. Required fields are marked *

Previous Post
OpenAI product page showing GPT-6 Intelligent UI examples in ChatGPT

ChatGPT’s GPT-6 Intelligent UI Turns Answers Into Interactive Tools

Next Post
A person holds a smartphone displaying a music streaming app beside a pair of headphones.

AI Music Bot Fraud Gets First U.S. Prison Sentence as Copyright Office Opens Inquiry

Related Posts
OpenAI CEO Sam Altman and Broadcom CEO Hock Tan holding a display with the Jalapeño inference chip wafer

OpenAI’s Jalapeño Benchmarks Turn Inference Power Into the AI Chip Fight

OpenAI published first benchmark results for Jalapeño, its Broadcom-built inference chip, claiming higher performance per watt and lower latency than leading commercial AI systems. The useful question is not whether it replaces Nvidia immediately, but whether custom inference silicon can turn power limits into a product advantage for ChatGPT, Codex, and agentic AI workloads.
Read More