QUOPS: New Quantum Benchmark Compares Google, IBM, Quantinuum for First Time: All Trail by 100,000x
21 hour ago / Read about 45 minute
Source:TechTimes

Sandia.gov

Quantum computing finally has a report card that works across every hardware type — and the first results reveal a field that is simultaneously more stratified than most buyers realized and farther from practical use than any prior metric made measurable. A new benchmark called QUOPS (Quantum Universal Operations Performance System), developed by a team led by Sandia National Laboratories with contributions from Quantinuum and NVIDIA, posted the first direct cross-platform performance scores for quantum computers on September 10, 2026. The results: Quantinuum's Helios-1 trapped-ion processor achieves a Q score of 1,504 — roughly seven times larger than Google's superconducting Willow (Q = 216) and IBM's ibm_boston (Q = 204). But all three systems, including Helios-1, are roughly 100,000 times below the minimum Q score researchers estimate is needed to run useful scientific applications, according to the QUOPS benchmark paper.

The benchmark arrived at IEEE Quantum Week 2026, which opened Sunday in Toronto, alongside NVIDIA's announcement that QUOPS is now integrated into its open-source CUDA-Q platform. The timing puts a standardized measurement tool in front of government hardware buyers, HPC centers, and DARPA's Quantum Benchmarking Initiative at exactly the moment those stakeholders most need one — with new Stage A proposals for the QBI due September 30, 2026. NVIDIA's announcement introduced CUDA-Q Logical, an orchestration layer for developing and verifying fault-tolerant quantum applications, on the same day QUOPS was integrated into the platform.

Why Qubit Count Has Never Been Enough

Quantum hardware benchmarking has long suffered from what critics call "spec-sheet theater." A machine might boast thousands of qubits, but if those qubits are too noisy to sustain a large circuit, the count means little for practical computation. Conversely, a system with fewer but higher-quality qubits — and sophisticated error correction — may far outperform in real workloads. Traditional metrics like raw qubit count, two-qubit gate fidelity, and gate speed remain vital for engineering diagnosis, but they do not predict system-level performance, and they fail entirely once fault-tolerant error correction enters the picture. The QUOPS paper addresses all of these limitations of prior benchmarks directly.

IBM introduced Quantum Volume in 2019 as the first attempt at a single holistic score, and it became the field's flagship system-level benchmark for several years. But Quantum Volume has documented limitations: its classical verification cost grows exponentially with qubit count, its metric is only loosely connected to real computational utility, and it only measures square circuits (equal width and depth), when real workloads are almost never square. The field needed a successor, and QUOPS is the first design to address all three gaps simultaneously.

Read more: Quantum Computing Roadmap: Coalition Sets Hard Milestones for Neutral-Atom Advantage

How QUOPS Works: Two Numbers That Tell the Full Story

QUOPS evaluates full system performance by running randomized circuit layers across varying shapes — specifically, arbitrary-angle single-qubit rotations and CNOT gates across circuits of varying widths (number of qubits) and depths (number of gate layers). Crucially, the measurement captures the integrated effect of compilation, error correction, syndrome decoding, and error mitigation — not just the performance of isolated components. A vendor cannot optimize for the benchmark on one dimension while hiding deficits elsewhere, because the framework includes explicit anti-gaming provisions.

From these experiments, the benchmark reports two summary metrics:

Q — the largest benchmark circuit size (defined as 2 × width × depth) that a system can execute successfully above a fidelity threshold of 1/√e ≈ 61% (the mean process polarization), within a utility-motivated region of circuit shapes.

Ω (Omega) — the effective quantum operations per second at that peak circuit size, accounting for all system overhead including error mitigation sampling cost.

Together, Q and Ω give a two-dimensional view of a machine's capability: how large a computation it can handle, and how quickly it can deliver results. Unlike Quantum Volume, which scores only square circuits, QUOPS evaluates rectangular circuits in both dimensions — producing a capability map that reflects the actual shapes of workloads a quantum system would need to run.

Leaderboard: What the Scores Actually Reveal

The paper provides the first direct, cross-platform comparison of leading quantum processors under a common methodology, confirmed by the Quantum Computing Report benchmark table:

SystemArchitectureQ ScoreΩ (QUOPS/s)

Quantinuum Helios-1

Trapped-ion (QCCD, 98 qubits)

1,504 (1,824 with leakage postselection)

303

Quantinuum H2-1

Trapped-ion (QCCD, 56 qubits)

1,320 (1,392 with postselection)

353

Google Willow

Superconducting (105 qubits, 2D grid)

216

~20,000,000

IBM ibm_boston

Superconducting (156 qubits, heavy-hex)

204

~310,000

Helios-1 (Logical FTQC)

Steane [[7,1,3]] code, 8 logical qubits

40

4.9

Helios-1's Q score of 1,504 is roughly seven times larger than either superconducting system — a striking lead in circuit-size performance data. But the comparison is not one-dimensional.

Google's Willow registers an Ω rate of approximately 20 million operations per second, reflecting superconducting architecture's inherent speed advantage. Helios-1's trapped-ion design achieves far larger computations but at roughly 303 QUOPS/s — about 66,000 times slower than Willow. The QUOPS framework treats this as a feature rather than a flaw: by plotting Q against Ω, it quantifies a speed-versus-depth tradeoff that has long been discussed qualitatively but never measured on a common scale.

The architectural reason for the gap is measurable in the data. Quantinuum's QCCD (Quantum Charge-Coupled Device) design physically shuttles ions between zones, enabling any qubit to interact with any other regardless of position. As a result, Helios-1's Q score declines relatively little as circuit width increases — there is no connectivity penalty for deep, wide circuits. Google's 2D-grid superconducting design, by contrast, can only entangle physically adjacent qubits; longer-range interactions require added routing operations that increase circuit depth and error accumulation. The Q-score dependence on circuit width is the benchmark's way of making that structural difference visible and numerically comparable.

Extending QUOPS to Fault-Tolerant Logical Qubits

The benchmark is designed to grow with the field. The paper also demonstrates QUOPS applied to a fully fault-tolerant logical architecture, running circuits on up to eight Steane-encoded logical qubits on Helios-1 — with active magic-state injection and syndrome extraction. The logical-qubit results were Q = 40 and Ω = 4.9 QUOPS/s, confirmed in the benchmark cross-platform data table.

Those numbers are far below Helios-1's physical-qubit scores, reflecting the overhead that current error-correction codes impose. But the framework's value is precisely that it makes this overhead visible and measurable — providing a baseline from which progress toward practical fault-tolerant computing can be tracked across successive hardware generations.

The paper includes historical QUOPS scores inferred from quantum volume data going back to 2018, showing that both trapped-ion and transmon processors have improved steadily along both the Q and Ω dimensions. Those trend lines provide a roadmap-compatible view of whether current improvement rates are sufficient to close the utility gap within any given planning horizon.

Read more: Quantum Error Correction Validated in Nature: Microsoft and Quantinuum Log 800-Fold Improvement

How Far Are We Really? The 100,000x Gap to Useful Computation

The benchmark's most consequential contribution may be quantifying exactly how far current machines are from where they need to be. By mapping the resource requirements of recognized scientific challenge problems into effective QUOPS circuit sizes, the paper identifies specific utility targets that give the field something concrete to aim for:

  • RSA-2048 factoring (breaking the encryption standard protecting most internet banking): Q ≈ 250,000,000 (2.5 × 10⁸)
  • FeMoco energy eigenvalue calculation (a key quantum chemistry milestone): Q ≈ 340,000,000 (3.4 × 10⁸)

Helios-1's Q score of 1,504 means today's leading machine is approximately 100,000 times below the RSA-2048 target — five full orders of magnitude. That specific, measurable deficit is something no prior benchmark made directly visible.

For security architects and enterprise IT leaders, this quantification matters because it connects quantum hardware progress to the regulatory clock already running. The National Security Agency's Commercial National Security Algorithm Suite 2.0 mandates that all new national security system acquisitions be quantum-safe starting January 2027. The National Institute of Standards and Technology finalized its first post-quantum cryptography standards in August 2024 and has called for RSA-2048 to be deprecated after 2030. The "harvest now, decrypt later" threat — in which adversaries capture encrypted communications today with the intent to decrypt them once a sufficient quantum machine exists — means the relevant deadline is not when a cryptographically relevant quantum computer is built, but when the data being protected today must remain confidential, as detailed in NIST's IR 8547 guidance.

NVIDIA Integration and Open Availability

QUOPS does not exist in isolation. NVIDIA integrated a reference implementation into its open-source CUDA-Q platform on September 14, announced as part of a broader expansion that also introduced CUDA-Q Logical — an orchestration layer for developing and verifying fault-tolerant quantum applications. The CUDA-Q Logical integration means researchers and procurement teams can run QUOPS on any CUDA-Q-connected hardware without reimplementing the benchmark from scratch.

The practical impact is already measurable. Fermilab used CUDA-Q Logical to cut fault-tolerant architecture development from roughly five months to three weeks — a sevenfold speedup. Iceberg Quantum used the platform to model 1,000 logical qubits from just 150,000 physical qubits for Diraq's architecture — approximately ten times fewer than prior estimates had suggested would be required. The benchmark code is additionally available as an open-source implementation on GitHub, and the full preprint is freely accessible on arXiv.

What QUOPS Means for Government Procurement and DARPA

The timing of the QUOPS release is not incidental. DARPA's Quantum Benchmarking Initiative (QBI) — launched to determine whether utility-scale fault-tolerant quantum computing can be achieved by 2033 — has advanced 11 companies to Stage B research-and-development phase. Proposals for a new Stage A solicitation are due September 30, 2026. DARPA program leadership has said it now considers it "likely that someone will build a utility-scale quantum computer by 2033" while emphasizing that which approach will succeed remains uncertain.

QUOPS arrives as a natural complement to that program. With a single Q score and Ω rate, government procurement officers, HPC centers, and commercial buyers gain a simpler system-level reference point that does not require them to parse competing vendor claims about code distance, magic-state factories, or qubit modality. Quantinuum has explicitly framed the benchmark in those terms, calling on buyers and agencies to require QUOPS scores in RFPs rather than relying on vendor-selected metrics.

What Should Buyers Watch: Caveats on an Early Standard

Industry observers have noted that the initial dataset covers only four platforms — Helios-1, H2-1, Willow, and ibm_boston — and that broader adoption across the full hardware landscape will be needed before QUOPS can serve as a definitive industry standard. The anti-gaming provisions are an important design choice, but independent verification across labs will ultimately determine whether the benchmark earns the trust of the broader research community.

There is also a structural observation worth noting: Quantinuum's researchers contributed to the benchmark's development, and Quantinuum's hardware scores highest on the Q dimension. This does not invalidate the methodology — the benchmark was led by Sandia National Laboratories, a government research institution with no commercial stake in the outcome — but buyers using QUOPS to compare vendors should be aware that the benchmark's current champion is also one of its architects, as confirmed in the paper's author affiliations.

There is also the question of completeness. QUOPS measures a specific class of random circuits as a proxy for general computational capability; application-specific benchmarks, hybrid quantum-classical workflows, and domain-specific suites remain necessary alongside it. The benchmark's authors acknowledge this, positioning QUOPS as a system-level yardstick rather than a replacement for the full benchmarking ecosystem.

What the field now has, for the first time, is a number — Q — that translates the dizzyingly technical landscape of quantum hardware into a single comparable figure: the largest circuit a machine can actually run. In a field long accused of overpromising and underdelivering, that clarity is itself a milestone — even if the gap it reveals is far larger than most timelines have acknowledged.


Frequently Asked Questions

What is QUOPS, and how is it different from Quantum Volume?

QUOPS (Quantum Universal Operations Performance System) is an architecture-agnostic benchmark that measures two things simultaneously: Q, the largest quantum circuit a machine can execute above a fidelity threshold, and Ω, how many operations per second it can perform at that scale. Details are available in the full QUOPS benchmark paper. Quantum Volume — developed by IBM in 2019 — only scores square circuits, requires classical simulation to verify (which becomes impossible at large qubit counts), and its score is only loosely connected to real scientific utility, as documented in the Quantum Volume overview. QUOPS evaluates rectangular circuits of any shape, works without exponentially costly classical simulation, and is calibrated directly to the resource requirements of recognized scientific challenge problems like RSA-2048 factoring and molecular energy calculations. Both are system-level benchmarks that capture compilation, connectivity, and error effects — QUOPS simply goes further and scales further.

Why do trapped-ion machines score seven times higher on Q but far lower on speed?

The Q-score gap reflects a fundamental architectural difference in qubit connectivity. Trapped-ion computers like Quantinuum's Helios-1 physically move ions between zones in a microfabricated trap, allowing any qubit to interact with any other without added routing. Superconducting chips like Google's Willow fix qubits to positions on a 2D grid; interactions between non-adjacent qubits require a chain of intermediate operations that lengthen circuits and accumulate errors. QUOPS measures this connectivity cost directly — superconducting systems show a steeper Q-score drop as circuit width grows, while trapped-ion systems maintain depth-capability more consistently. The speed gap exists because moving ions between zones takes hundreds of nanoseconds to a few microseconds, while superconducting gates complete in tens of nanoseconds. Both dimensions matter; which one dominates depends on the specific workload.

How far are today's best quantum computers from breaking RSA-2048 encryption?

By the QUOPS framework's own estimates, factoring a 2,048-bit RSA key requires a Q score of approximately 250 million — more than 100,000 times Helios-1's current score, as shown in the QUOPS cross-platform benchmark analysis. That gap is the reason the paper recommends fault-tolerant approaches: no amount of improvement in physical-qubit quality alone can close five orders of magnitude. Reaching RSA-2048 capability requires the error-correction overhead to shrink dramatically, the logical qubit count to scale by orders of magnitude, and operational throughput to improve alongside circuit size. NIST has called for RSA-2048 to be deprecated after 2030 and formally disallowed after 2035, and the NSA's CNSA 2.0 framework requires all new national security system acquisitions to be quantum-safe starting January 2027. No credible timeline puts a cryptographically relevant quantum computer within a few years.

Can government procurement teams actually use QUOPS scores right now in hardware decisions?

The benchmark is usable today, with important caveats. A QUOPS reference implementation is already integrated into NVIDIA's open-source CUDA-Q platform and available on GitHub. Quantinuum has explicitly called on government agencies and HPC centers to require QUOPS Q and Ω scores in hardware requests for proposals, arguing that a common score avoids vendors cherry-picking favorable metrics, as detailed in the Quantinuum QUOPS procurement guidance. DARPA's Quantum Benchmarking Initiative — with Stage A proposals due September 30, 2026 — could adopt QUOPS as a standardized evaluation input without requiring new methodology development. The main limitation: only four systems have been publicly benchmarked so far, and Google and IBM have not yet independently verified or published their own QUOPS assessments of their hardware. Independent replication across labs is what will ultimately determine whether QUOPS earns the trust its institutional provenance suggests it deserves.