Run the numbers on Wolfgang Mauerer’s new quantum computing reproducibility study, and you land somewhere around nine percent. That’s the share of quantum computing papers whose results a stranger could actually reproduce, according to the team he leads at the Technical University of Applied Sciences Regensburg, and I’ve been turning that number over for a day now because it isn’t coming from an outside critic with an axe to grind. It’s coming from inside the field, from researchers who went looking for exactly this problem and found it worse than they expected.
Mauerer’s group split the work into two passes. First, they hand-checked 127 papers from the last five years against five criteria: whether the paper included code, how much documentation came with it, whether there was enough hardware detail to understand what ran it, and, critically, whether the code actually executed when someone tried it. Only 24.4 percent of those 127 papers included code you could even attempt to run. Of that quarter, 64.5 percent failed outright. Multiply those two numbers and you get the nine percent I opened with. The second pass ran an automated version of the same check across nearly 5,000 papers, skipping the execution test for scale, and found only 26.8 percent provided enough information to attempt a replication at all.
None of this surprises Mauerer. His team ran a smaller version of the same study four or five years ago and found the field already struggling. “We thought the numbers would be better by now,” he told New Scientist, which is a fairly quiet way of saying the community had years to fix this and mostly didn’t.
He’s also careful not to make it a laziness story. Quantum hardware itself is the harder problem: a cloud accessible quantum processor can behave differently from one week to the next in ways a classical CPU simply doesn’t, so a result that held on Tuesday’s calibration run might not hold on Thursday’s. “I don’t just want to attribute the lack of reproducibility to, say, some laziness of authors or non-requirements by the community,” Mauerer said. “Quantum machines in themselves are unusually variable.” That’s a real and separate problem from the culture one, and conflating the two would let the field off easy.
The reactions split, as you’d expect. William Zeng at the Unitary Foundation isn’t surprised and hopes AI coding tools will quietly fix a chunk of this by generating reproduction code straight from a paper’s stated results. Fred Chong at the University of Chicago isn’t alarmed at all, arguing reproducibility becomes a priority only once a field matures, and that conventional computer science took decades to get its own house in order. I don’t fully buy the second argument. Software engineering didn’t issue press releases about world-changing breakthroughs every quarter as it matured. Quantum computing companies want it both ways: cite the field’s youth when someone tries to check the work, skip the caveat entirely when the funding round or the keynote slide needs a bold number.
This is the exact hole I kept circling a couple of weeks ago writing about IBM’s quantum advantage claims getting a second opinion from two rival verification outfits, Algorithmiq and Qedma, in a piece about who actually gets to check a quantum result. The problem there was that a quantum computer can produce an answer no classical machine can check, so trust has to come from somewhere else, ideally an independent replication. Mauerer’s numbers say that independent replication mostly isn’t happening, not because nobody’s trying but because most papers don’t leave enough of a trail to try.
It reframes a few things I’ve covered this month without saying any of them are wrong exactly. D-Wave’s 99.9 percent two-qubit gate fidelity number is a real, specific, well-reported claim, and I still think the fidelity number matters less than the roadmap behind it, but I’d feel better about that number if I knew how many outside groups had actually tried to hit it on their own hardware. Oracle folding a Quantinuum trapped ion machine directly into its own cloud is a genuine structural bet, and it still rests, like almost everything in this field, on benchmark numbers the vendor itself is the one reporting.
It also raises the stakes on something closer to home for this blog. Illinois’s quantum park is deliberately betting on four incompatible qubit architectures at once rather than picking a winner, which is a defensible hedge if you can’t yet tell which approach is actually ahead. A reproducibility crisis makes that judgment harder, not easier, because the papers you’d use to rank photonic against superconducting against neutral atom against spin qubits are exactly the papers Mauerer’s team says mostly can’t be checked.
I don’t know whether agentic coding tools actually close this gap or just produce more code that also fails to run under different circumstances, and I’m not going to pretend I do. What I do know is that the next quantum press release I read, mine or anyone else’s, gets read a little differently now. Nine percent is not a number you build billion-dollar infrastructure decisions on faith around, and yet that’s roughly what the entire industry has been doing.
Sources
- New Scientist, Quantum computing may be facing a replication crisis, August 17, 2026
- arXiv preprint, Mauerer et al., reproducibility analysis of quantum computing papers, 2026