
Microsoft Corporate Vice President, Surface Devices Brett Ostrum speaks during the Microsoft May 20 Briefing event at Microsoft in Redmond, Washington, on May 20, 2024. JASON REDMOND/AFP via Getty Images
Months before any official review embargo lifts, a tech enthusiast who says he found a pre-production Microsoft Surface Laptop Ultra lying on the side of a road near Microsoft's Redmond, Washington campus has published the first independent, hands-on benchmark data for NVIDIA's RTX Spark platform — and the results reveal a significant gap between the chip's headline AI computing promise and what the pre-release software can currently deliver in his month-long hands-on review.
The tester, a TechPowerUp forum member known as Fouquin, says he picked up the device in mid-June 2026, assuming it was discarded hardware. It turned out to be an EV1.5 engineering sample — the designation applied to pre-production validation units — carrying the top-tier RTX Spark N1X configuration with a 20-core Arm CPU, 6,144 CUDA cores, 24GB of unified memory, and a 512GB SSD, as confirmed by TechSpot's independent report. After roughly one month of testing under both old and new pre-release drivers, Fouquin posted a detailed breakdown covering performance, power management, thermals, and build quality. The write-up has since been reported by Tom's Hardware, TechSpot, VideoCardz, and Notebookcheck as the most extensive independent hands-on with RTX Spark hardware yet conducted, as detailed in Tom's Hardware's coverage.
The findings arrive at a moment when buyers evaluating premium AI-focused laptops for the fall 2026 window need to understand what RTX Spark has and has not yet demonstrated in independent testing. The headline from this prototype: the silicon is capable, the build quality is strong, and the battery performance is genuinely impressive — but the platform's defining selling point, the ability to run large AI models locally using CUDA, did not work on either of the two pre-release driver versions Fouquin tested.
NVIDIA's RTX Spark marketing centers on a single number: 1 petaflop of FP4 AI performance, enough to run 120-billion-parameter language models entirely on-device without a cloud subscription. That claim is the reason developers and AI researchers have been watching this platform since its announcement at Computex 2026 in Taipei, and it is the area where the Fouquin data is most consequential.
In practice, CUDA-based tests in the Phoronix AI benchmark suite failed to execute correctly on both driver versions tested, as Tom's Hardware's coverage confirms. The older 591.33 drivers, which shipped with the device from November 2025, lacked mature CUDA support. The newer 616.00 developer preview — which does include CUDA 13.4 — still could not complete AI workloads through the CUDA path. Fouquin was forced to fall back to CPU-based and Vulkan Compute paths, which are significantly slower alternatives, according to TechSpot's reporting.
To be precise about what this finding does and does not mean: it says nothing about whether a retail Surface Laptop Ultra in October or November 2026 will successfully run CUDA AI workloads. Driver development for a new platform is iterative, and EV1.5 engineering samples are explicitly not running production software. What the data does establish is that no independent party has yet been able to confirm NVIDIA's AI performance claims on actual RTX Spark hardware, because the CUDA stack needed to run those workloads has not yet been validated outside NVIDIA's own controlled demonstrations.
For buyers deciding whether to wait for this platform or purchase a current AI-capable laptop now, the honest summary is: trust the NVIDIA AI performance claims only when independent reviewers with retail units can verify them.
Read more: NVIDIA RTX Spark Brings Blackwell AI To Windows Laptops This Fall: Intel And AMD Shares Slide
To understand what RTX Spark is trying to accomplish, a quick look at the engineering is essential. The chip is not a single monolithic die like Apple's M-series — it is a 2.5D chiplet package built on TSMC's 3nm process, pairing a MediaTek-designed Grace CPU die with an NVIDIA Blackwell GPU die, connected via NVIDIA's NVLink chip-to-chip interconnect (NVLink C2C) at 300 gigabytes per second of bidirectional bandwidth, per a technical architecture breakdown.
The CPU portion uses 10 Arm Cortex-X925 performance cores paired with 10 Arm Cortex-A725 efficiency cores in a 5p5e configuration, reaching peak clock speeds of approximately 4.1 GHz, as detailed in this CPU core configuration guide. The GPU carries 6,144 CUDA cores across 48 streaming multiprocessors with fifth-generation Tensor Cores capable of processing the FP4 number format — a lower-precision format that trades some numerical accuracy for more AI operations per second, which is what produces the 1 petaflop headline figure.
What makes the platform architecturally distinctive is the unified memory pool: up to 128GB of LPDDR5X at 300 GB/s bandwidth, shared dynamically between the CPU and GPU, per NVIDIA's unified memory specifications. This eliminates the memory transfer bottleneck that exists in conventional laptops, where moving data between system DRAM and discrete GPU VRAM costs both latency and bandwidth. A 120-billion-parameter language model that would require a separate $4,699 DGX Spark workstation to run comfortably can theoretically fit inside a laptop's memory pool when that pool scales to 128GB.
The tradeoff worth noting: Apple's M4 Max, the most direct competitor in the unified memory space, offers approximately 540 GB/s of memory bandwidth — nearly double what RTX Spark provides — from a monolithic SoC design where the CPU and GPU share the same die rather than communicating across a chip-to-chip interconnect. RTX Spark counters with a dramatically larger maximum memory capacity (128GB vs. 128GB at the high end, but the base M4 Pro MacBook Pro ships with 48GB) and a vastly more powerful GPU compute stack for CUDA-heavy workloads, where Apple's Metal platform has no equivalent.
With the default power limits raised to 80W PL1 and 95W PL2 via a hidden High Performance profile Fouquin discovered, the prototype's CPU produced Cinebench 2024 multicore and single-core scores of 1,386 and 123 points respectively, per VideoCardz's Cinebench results. In Cinebench 2026, the same configuration produced 5,771 points in multicore and 540 points in single-core, with a multi-to-single-core performance ratio of 10.68x — measured over the benchmark's 10-minute throttling mode, as reported by VideoCardz.
The Cinebench 2024 multicore score trails the 12-core MacBook Pro with M4 Pro, while the Cinebench 2026 multicore score falls noticeably behind the Apple M4 Max, per Tom's Hardware's analysis. The earlier Clang compiler benchmark from Computex hands-on sessions told a different story: in code compilation, RTX Spark outperformed the standard 10-core M5 by approximately 54%, coming within about 7% of the 15-core M5 Pro — suggesting the platform is stronger in highly threaded developer workloads than in the mixed workloads Cinebench measures, as shown in the Clang benchmark comparison.
One important artifact of the pre-production state: Cinebench currently misidentifies the processor as "JMJWOA-Generic-CPU" and reports a fixed 1.01 GHz clock, which is a benchmark identification quirk of the pre-release platform rather than any reflection of the chip's actual operating frequency, as VideoCardz confirmed.
The GPU benchmarks from this prototype are the numbers most likely to be misread, and the context is critical: NVIDIA's stated performance target for the RTX Spark N1X GPU is roughly equivalent to a mobile RTX 5070 — matching the RTX 5070's CUDA core count but operating at laptop-class power limits rather than a desktop card's 250W thermal budget.
What the Fouquin prototype produced under both driver versions was substantially less. GPU clock speeds fluctuated between 1.5 and 2.3 GHz during gaming, accompanied by periodic stutter and frequent crashes in Helldivers 2, as Tom's Hardware reported. In 3DMark tests, the best results came with the older 591.33 drivers rather than the newer 616.00 preview: Port Royal produced 26-31 FPS, Nomad achieved 20-21 FPS, and Unigine Superposition reached 56-58 FPS. Compared against desktop reference cards, those results roughly correspond to an RTX 3060 or RTX 3060 Ti, not the mobile RTX 5070-class performance the platform is targeting.
The counterintuitive finding that updating to newer drivers actually hurt performance in some tests is a clear sign of software immaturity rather than hardware limitation. One particularly notable data point: power consumption behavior deteriorated with the newer drivers, with idle power draw rising from 7-8 watts on the old drivers to 12-23 watts depending on power profile — a regression that added heat and battery draw without measurable performance benefit, as TechSpot's reporting confirms.
Kingdom Come: Deliverance II ran properly and served as one of the few examples of a game that works under the Windows on Arm emulation layer. The broader gaming picture reflects the persistent challenge of Windows on Arm: most game libraries are not natively compiled for Arm architecture, and while Microsoft's Prism emulation layer has improved substantially since its introduction, games with kernel-level anti-cheat systems still require native Arm builds from each developer individually, as covered in this Windows on Arm compatibility overview.
Read more: Nvidia RTX Spark Superchip: Windows PC Chip With Full CUDA Stack Targets Dell, Microsoft This Fall
At sustained workloads, the prototype's CPU peaked at 98-100°C (208-212°F) and the GPU hit 80-90°C (176-194°F), as documented by Notebookcheck's detailed analysis. Those temperatures are elevated for a thin chassis, though pre-production thermal tuning is consistently among the last items finalized before retail builds ship, and Microsoft has publicly confirmed that the Surface Laptop Ultra's cooling system was designed specifically for sustained workloads with 2.5x the thermal headroom of the Surface Laptop 7 15-inch.
One result stands out as potentially significant for retail buyers regardless of driver state: Fouquin confirmed that plugging the laptop into AC power produced no measurable performance difference compared to running on battery. NVIDIA made this claim at Computex — that RTX Spark machines would deliver full performance on battery — and the prototype data supports it, even under pre-release software, per Tom's Hardware's reporting. For professionals who need reliable workstation-level performance away from a desk, that result matters.
Not everything in Fouquin's month-long hands-on was a cautionary note. The physical machine drew consistent praise. He described the keyboard as the best he had used outside of a MacBook, characterized by punchy key travel, smooth resistance, and zero switch wobble. The milled aluminum chassis mirrors the 16-inch MacBook Pro in physical footprint — fitting, given that Microsoft is explicitly positioning the Surface Laptop Ultra as its answer to Apple's professional laptop line, as TechSpot noted.
The touchpad is similarly MacBook-caliber, and the hinge mechanism drew high marks. A proximity-sensing backlight that activates when the laptop's camera detects a user nearby is a useful quality-of-life feature. The 15-inch mini-LED display's physical quality impressed despite unfinished firmware: local dimming zone behavior, G-Sync support, touch input calibration, and HDR tuning all showed signs of incomplete software — issues consistent with an engineering sample running pre-retail firmware, according to TechSpot's hands-on report.
The chassis weighs under 4.5 lbs (approximately 2 kg) and measures less than 18mm (0.71 inches) thick, which matches the MacBook Pro 16 in both weight and profile, per Microsoft's official product page. Microsoft has also confirmed the Surface Laptop Ultra features a user-replaceable SSD — a meaningful durability and repairability advantage over the MacBook Pro, where storage is soldered.
One caveat Fouquin noted on repairability: while the device is technically serviceable, thin aluminum snap-on panels cover all primary components including the SSD, meaning practical repair attempts are likely to require careful technique or risk bending and breaking panels.
The available data is incomplete and prototype-specific, but worth summarizing honestly. In Cinebench 2026 multicore, the prototype trailed the M4 Max — though early Clang developer benchmarks from Computex showed RTX Spark beating the standard M5 by 54% and coming within 7% of the 15-core M5 Pro, suggesting the gap is workload-dependent, per the Clang benchmark data. In GPU rasterization under pre-release drivers, the prototype produced RTX 3060-level results against NVIDIA's own mobile RTX 5070-class target — a gap that is almost certainly driver-driven rather than hardware-driven. In CUDA AI workloads, no comparison is currently possible because the CUDA path did not function, as Tom's Hardware confirmed.
Apple's M4 Max offers roughly 540 GB/s of memory bandwidth compared to RTX Spark's 300 GB/s — but RTX Spark's advantage in raw CUDA compute, full DLSS support, and a memory capacity ceiling of 128GB potentially changes which platform is preferable for AI-intensive workloads, when the CUDA stack is working, according to NVIDIA's unified memory specifications.
The honest conclusion from the available data: CPU parity with Apple's M-series is plausible for the retail product, GPU performance at the mobile RTX 5070-class target will require driver maturation the prototype has not yet demonstrated, and AI workload capability remains unvalidated by any independent party.
Microsoft has not officially announced pricing for the Surface Laptop Ultra. A Morgan Stanley analysis of channel checks at Computex placed N1X-class flagship configurations at approximately $2,899 and lower N1 configurations near $1,799, as covered in TechTimes' pricing analysis. Fouquin's own assessment after testing the device is more expansive: he predicted the Surface Laptop Ultra would start no lower than $2,800, with the top-end 128GB unified memory configuration reaching approximately $4,500 — territory directly competitive with Apple's flagship MacBook Pro M5 Max configurations at $3,599-$3,899, per TechSpot's report.
Six OEMs have confirmed RTX Spark devices for a fall 2026 window: ASUS, Dell, HP, Lenovo, Microsoft Surface, and MSI, with Acer and Gigabyte expected to follow. Pricing on third-party N1X configurations may differ from Microsoft's Surface Laptop Ultra positioning.
For anyone targeting this platform specifically for local AI workloads — the use case NVIDIA has marketed most aggressively — the only responsible answer is: wait for retail units and independent benchmarks. The CUDA pipeline failure on the prototype is a pre-release software issue that will almost certainly be resolved, but "almost certainly" is not the same as "confirmed working," and a $2,899-plus purchase decision deserves verified performance data.
For buyers interested in RTX Spark primarily for creative work, development, and gaming, the picture is similar but less urgent: the CPU performance appears real and competitive, the build quality is genuinely strong, and the GPU will almost certainly perform at a higher level with production drivers than the prototype showed. The gaming situation depends heavily on which titles are in your library — Windows on Arm compatibility has improved substantially, but the emulation layer still introduces friction that native x86 Windows does not.
NVIDIA and Microsoft have until fall to close the software gap. The hardware, at least, appears to be ready.
This capability has not yet been independently verified on actual RTX Spark hardware. NVIDIA's 1 petaflop FP4 figure comes from the company's own Computex demonstrations. In the only independent hands-on with RTX Spark hardware conducted to date, CUDA-based AI workloads failed to complete on either of the two pre-release driver versions tested. The CUDA stack required for those workloads is still being finalized ahead of the fall 2026 retail launch. Independent reviewers with retail units will provide the first verifiable performance data.
Both platforms eliminate the DRAM/VRAM split that exists in conventional laptops, letting CPU and GPU work from the same memory pool. The key differences: Apple Silicon uses a monolithic chip design with approximately 540 GB/s of memory bandwidth (M4 Max), while RTX Spark uses a chiplet design connecting separate CPU and GPU dies via NVLink C2C at 300 GB/s — lower bandwidth, but NVIDIA's platform scales to 128GB unified memory and brings a full CUDA ecosystem for AI and GPU compute workloads that Apple's Metal platform does not support. Which matters more depends on the workload: Apple wins on per-GB bandwidth; RTX Spark wins on CUDA software compatibility and maximum model size.
Microsoft has confirmed a fall 2026 availability window but has not announced official pricing or a specific launch date. A Morgan Stanley analysis of Computex channel checks placed N1X-class flagship configurations at approximately $2,899 and entry N1 configurations near $1,799. The tester who benchmarked the prototype predicted the Surface Laptop Ultra itself will start no lower than $2,800, with the fully loaded 128GB configuration reaching approximately $4,500 — directly competing with Apple's MacBook Pro M5 Max lineup.
No — and this is the most important interpretive point for the Fouquin data. Engineering samples (labeled EV1.5 in this case) run pre-production firmware and early-stage drivers specifically not intended to represent final performance. GPU scores that place this prototype at RTX 3060 equivalent are almost certainly understating the retail chip's capability. The CPU benchmarks are more likely to be predictive, since CPU performance typically matures earlier in the development cycle than GPU driver stacks. The CUDA AI workload failures are the most significant data point, not because they predict retail behavior, but because they establish that the platform's core AI claim has not yet been independently verified on real hardware.
