
(Image credit: AMD)
Nvidia is expected to launch RTX Spark devices later this week, following a not-so-subtle tease at the end of last week (and the chip’s announced fall release). Ahead of Microsoft’s event on Wednesday, October 7 (where the launch is expected) AMD has shared some benchmarks for its new Gorgon Halo chips — the range that will directly compete with RTX Spark devices. The extra insight comes a matter of days after the first Gorgon Halo devices launched, some of which cost upwards of $7,099.
(Image credit: Tom's Hardware)
Short of a few questionable Geekbench leaks, we haven’t seen any performance results for the RTX Spark yet, so AMD isn’t using it as a comparison point. Rather, it’s comparing the Ryzen AI Max+ Pro 495 to Intel’s Core Ultra X9 388H. These chips aren’t in the same class of device, though AMD would argue that it’s comparing its top-of-stack part to Intel’s top-of-stack part.
AMD used ComfyUI to measure generative AI performance. AMD averaged multiple runs of various models, comparing total throughput to Intel’s competition. AMD used its top-spec 192GB configuration of the Ryzen AI Max+ Pro 495 and compared it to the Core Ultra X9 388H in a system with 64GB of memory (Panther Lake supports up to 128GB).
The performance advantage ranges from 1.1x up to 32.2x, though the end point is a clear outlier. We searched for Yuve on Hugging Face and didn’t find any results. It’s possible this delta comes down to an optimization issue, or that it’s simply too big to run on the Panther Lake machine.
(Image credit: AMD)
We don’t have gaming or general application performance, but we don’t expect a major swing compared to last-gen Strix Halo chips. Gorgon Halo is largely a refresh of that range.
Gorgon Halo isn’t getting into the ring with Panther Lake, however. It’s going mainly against the RTX Spark, and to a lesser extent, Apple’s larger M-series SoCs. This category of agentic PCs, as AMD calls it, is seemingly expanding, though it’s still far smaller than some of the hype around the RTX Spark would have you believe. AMD bragged in a prebriefing with the press about shipping “10s of millions” of AI PCs (read: laptops) before clarifying that it had slipped “over half a million” agentic PCs. Presumably, those are numbers for Strix/Gorgon Halo devices.
The matchup between Gorgon Halo and the RTX Spark has, up to this point, focused mainly on memory capacity. RTX Spark devices top out at 128GB of unified memory, same as the GB10 in the DGX Spark (the two chips are nearly identical). AMD, on the other hand, supports up to 192GB with Gorgon Halo. Higher capacity means running larger models locally, though at a lower performance level.
(Image credit: AMD)
In the GLM 5.3 Flash with 320 billion parameters, AMD saw peak throughput of 20 tokens per second, though using Unsloth's UD-IQ4_XS mixed-quantization format. For context, we clocked peak token throughput on the DGX Spark running GPT-OSS 120B with 4-bit quantization at 64 tokens per second, and the last-gen Ryzen AI MAx+ 395 at 56 tokens per second.
Note that although GLM 5.3 Flash has 320 billion total parameters, only 18 billion are activated for each token. Similarly, GPT-OSS 120B is another mixture-of-experts model with 120 billion total parameters, though only around 5 billion are active per token.
(Image credit: AMD)
AMD also shared Qwen 3.8 Flash Next performance, a multimodal MoE model with 125 billion main parameters, 51 billion embedding parameters, and about 4 billion parameters for multi-token prediction (MTP), with 5 billion parameters active per token. AMD says the Ryzen AI Max+ Pro 495 achieves up to 42 tokens per second with this model, once again using Unsloth’s dynamic 4-bit quantization and MTP.
That’s solid performance, but in both cases, “up to” carries a lot on its shoulders. The story of token throughput is told as the context length increases, showcasing what happens when someone actually runs these models locally, not just boots them up cold. Performance drops at higher context lengths, naturally, which could pose some issues for the larger models. If GLM 5.3 Flash provides up to 20 tokens per second, it could very easily decline into unusable territory as the context length increases.
We largely know what performance to expect out of the Ryzen AI Max+ Pro 495, and Gorgon Halo more broadly. It’s a refresh of Strix Halo, with notable spec changes being the bump up to 192GB of unified memory from 128GB, as well as a 100 MHz jump on boost clocks for the 495. Otherwise, the range is using identical core counts and microarchitectures as previous-gen Strix Halo chips.
With these proxies — Strix Halo for Gorgon Halo, and GB10 for RTX Spark — we can already get a good idea about how these parts will stack up. As you can see in our Ryzen AI Halo review (packing the Ryzen AI Max+ 395), AMD’s part universally underperformed compared to the DGX Spark in both time to first token and tokens per second across three models. More unified memory will allow you to run larger models, but that doesn’t mean those models will run faster.
The first Gorgon Halo devices are available for sale now, such as the Minisforum MS-S1 Max-P495. For the top-line configuration, prices sit around $7,000 right now, though we expect a broad range of prices once different devices are available, likely driving above that $7,000 mark.
(Image credit: AMD)
(Image credit: AMD)
(Image credit: AMD)
(Image credit: AMD)
(Image credit: AMD)
(Image credit: AMD)
(Image credit: AMD)
(Image credit: AMD)
(Image credit: AMD)
(Image credit: AMD)
(Image credit: AMD)
(Image credit: AMD)
(Image credit: AMD)
(Image credit: AMD)
(Image credit: AMD)
(Image credit: AMD)
(Image credit: AMD)
(Image credit: AMD)
(Image credit: AMD)
(Image credit: AMD)
(Image credit: AMD)
(Image credit: AMD)
(Image credit: AMD)
(Image credit: AMD)
View Original
View Original
View Original
View Original
View Original
View Original
View Original
View Original
View Original
View Original
View Original
View Original
View Original
View Original
View Original
View Original
View Original
View Original
View Original
View Original
View Original
View Original
View Original
View Original
Follow Tom's Hardware on Google News, or add us as a preferred source, to get our latest news, analysis, & reviews in your feeds.
