Willow’s Five-Minute Quantum Benchmark Explained
Google’s Willow chip completed a random-circuit-sampling benchmark in under five minutes, while its separate peer-reviewed error-correction result showed logical errors falling as encoded qubit grids grew.
Timeline
- December 9, 2024: Google introduced Willow and reported the benchmark and error-correction results.
- February 27, 2025: Nature published the peer-reviewed paper on below-threshold quantum error correction.
Google introduced Willow as a 105-qubit superconducting quantum chip and reported two different research results. One was a random circuit sampling benchmark completed in under five minutes. The other was a quantum error-correction experiment in which larger encoded systems had lower logical error rates. Combining them into one claim can obscure what each experiment actually tested. [1][2]
Random circuit sampling, or RCS, instructs a quantum processor to run circuits built from sequences of randomly chosen gates and then sample the resulting bit strings. Researchers compare the output distribution with what the circuit should produce. The task is designed to stress the whole processor and becomes extremely difficult to simulate classically as circuit size and depth increase. [1]
Google said Willow performed its chosen RCS computation in less than five minutes and estimated that one of the fastest classical supercomputers would require 10 to the 25th years under the modeled assumptions. That enormous comparison is Google’s estimate, not an independently timed classical run lasting that long. Classical algorithms and hardware can improve, so such estimates depend on the best known simulation methods and assumptions at the time. [1]
The benchmark does not mean Willow solved a five-minute commercial problem. Google explicitly wrote that RCS has no known real-world application and described the next challenge as a useful computation beyond classical reach. The experiment demonstrates processor performance on a deliberately difficult sampling task; it does not show that current quantum computers can replace ordinary computers for banking, web browsing, drug design or general artificial intelligence. [1]
The error-correction result addresses a different obstacle. Physical qubits are fragile, so researchers combine many of them into an encoded logical qubit and repeatedly detect errors. In the Nature paper, Willow-based surface-code systems scaled from distance three to distance five and distance seven. The reported logical error rate fell as code distance increased, which is the below-threshold behavior needed for scalable error correction. [1][2]
Below threshold is a milestone rather than a complete fault-tolerant computer. Useful large calculations would require many reliable logical qubits, long computations, control systems and further reductions in error. The paper also reports remaining error mechanisms and emphasizes the engineering needed to scale. Willow demonstrates a direction in which adding physical qubits can improve encoded reliability instead of merely adding more opportunities for failure. [2]
The most accurate summary separates performance from usefulness: Willow produced a classically difficult RCS sample quickly according to Google’s benchmark and showed peer-reviewed below-threshold error correction in encoded memories. The first supports a quantum-advantage claim for a specialized test; the second supports progress toward reliable logical qubits. Neither result by itself establishes a commercially useful, general-purpose quantum computer. [1][2]
Sources
- Google Quantum AI — Meet the Willow quantum chip
- Nature — Quantum error correction below the surface code threshold