Skip to content
Open access

Reinforcement learning control of quantum error correction

V. Sivak A. Morvan M. Broughton R. Cortiñas J. Bausch A. W. Senior M. Neeley A. Eickbusch N. Shutty L. Beni J. S. Spencer Francisco J. H. Heras Thomas Edlich D. Abanin Amira Abbas R. Acharya G. Aigeldinger R. Alcaraz S. Alcaraz T. Andersen M. Ansmann F. Arute K. Arya W. Askew N. Astrakhantsev J. Atalaya B. Ballard J. Bardin Hector Bates A. Bengtsson M. Karimi A. Bilmes S. Bilodeau F. Borjans A. Bourassa J. Bovaird D. Bowers L. Brill Peter Brooks D. Browne B. Buchea B. Buckley T. Burger B. Burkett N. Bushnell J. Busnaina A. Cabrera J. Campero Hung-Shen Chang Silas Chen B. Chiaro Liang-Ying Chih A. Cleland B. Cochrane M. Cockrell J. Cogan R. Collins P. Conner Harold Cook W. Courtney A. Crook B. Curtin M. Damyanov Sayan Das D. Debroy S. Demura P. Donohoe I. Drozdov A. Dunsworth V. Ehimhen A. Elbag Lior Ella M. Elzouka David Enriquez C. Erickson V. Ferreira Marcos Flores L. Burgos E. Forati J. Ford A. Fowler B. Foxen M. Fukami A. Fung L. Fuste S. Ganjam G. García C. Garrick R. Gasca H. Gehring R. Geiger É. Genois W. Giang D. Gilboa J. Goeders E. Gonzales R. Gosula Stijn J. de Graaf A. Dau D. Graumann J. Grebel A. Greene J. Gross Jose Guerrero L. Le Guevel T. Ha S. Habegger T. Hadick A. Hadjikhani Michael C. Hamilton M. Harrigan S. Harrington J. Hartshorn S. Heslin P. Heu O. Higgott R. Hiltermann Hsin-Yuan Huang M. Hucka C. Hudspeth A. Huff W. Huggins E. Jeffrey Shaun Jevons Zhang Jiang Xiaoxuan Jin C. Joshi P. Juhás A. Kabel D. Kafri Hui Kang K. Kang A. Karamlou R. Kaufman K. Kechedzhi T. Khattar M. Khezri Seon Kim C. Knaut B. Kobrin F. Kostritsa J. Kreikebaum Ryuho Kudo B. Kueffler Arun Kumar V. Kurilovich V. Kutsko N. Lacroix D. Landhuis T. Lange-Dei B. Langley P. Laptev K. Lau J. Ledford Joy Lee Kenny Lee B. Lester W. Leung Lily Li W. Li M. Li A. Lill W. Livingston M. Lloyd A. Locharla Laura de Lorenzo D. Lundahl A. Lunt S. Madhuk Aniket Maiti A. Maloney S. Mandrá L. Martin O. Martin Eric Mascot P. Das D. Maslov M. Mathews C. Maxfield J. McClean M. McEwen S. Meeks K. Miao Z. Minev R. Molavi S. Molina S. Montazeri C. Neill Michael Newman A. Nguyen M. Nguyen Chia-Hung Ni M. Niu L. Oas Raymond Orosco K. Ottosson A. Pagano A. D. Paolo S. Peek David Peterson A. Pizzuto E. Portolés R. Potter O. Pritchard Michael Qian C. Quintana A. Ranadive M. Reagor R. Resnick D. Rhodes Daniel Riley G. Roberts Roberto Rodriguez E. Ropes Lucia B De Rose Eliot Rosenberg E. Rosenfeld D. Rosenstock E. Rossi P. Roushan D. Rower R. Salazar K. Sankaragomathi M. Sarihan K. Satzinger Max Schäfer Sebastian Schroeder H. Schurkus A. Shahingohar M. Shearn A. Shorter V. Shvarts S. Small W. C. Smith D. Sobel B. Spells S. Springer G. Sterling J. Suchard A. Szasz A. Sztein M. Taylor J. Thiruraman D. Thor D. Timucin E. Tomita A. Torres M. Torunbalci H. Tran A. Vaishnav J. Vargas S. Vdovichev G. Vidal C. Heidweiller M. Voorhees S. Waltman J. Waltz Shannon Wang B. Ware J. Watson Yonghua Wei T. Weidel Theodore White K. Wong B. Woo Christopher J. Wood M. Woodson C. Xing Z. Yao P. Yeh B. Ying Juhwan Yoo N. Yosri Elliot Young G. Young A. Zalcman Ran Zhang Yaxing Zhang N. Zhu N. Zobrist Zhenjie Zou R. Babbush Dave Bacon S. Boixo Yu Chen Zijun Chen M. Devoret M. Hansen J. Hilton Cody Jones J. Kelly A. Korotkov E. Lucero A. Megrant H. Neven W. D. Oliver G. Ramachandran V. Smelyanskiy P. Klimov
Jul 2026 · Nature · Vol 655, pp. 879 - 884 · 0 citations · 57 references
Medicine

Abstract

Quantum error correction (QEC) is the primary strategy for protecting a quantum computer from the environment1,2. The prerequisite of QEC is that errors must remain sufficiently rare, which requires perpetually adapting the control parameters of the computer to the drifting environmental conditions. The current solution to this problem is to terminate the entire quantum computation for recalibration, but it is incompatible with the long runtimes of future quantum algorithms3,4. Here we address this challenge by unifying calibration with computation. We grant the QEC process5, 6, 7, 8, 9, 10–11 a dual role: its error-detection events are not only used to correct the logical quantum state but are also repurposed as a learning signal, teaching a reinforcement learning agent12, 13, 14, 15–16 to continuously steer the control parameters and stabilize the quantum system during computation. We experimentally demonstrate this framework on a Willow superconducting processor, improving the logical stability of the surface code 3.5-fold against injected drift. By synthesizing our full suite of technological advances, we achieve record performance of the surface and colour codes, with average logical error per cycle of 7.72(9) × 10−4 and 8.19(14) × 10−3, respectively. Numerical simulations of large codes with tens of thousands of control parameters confirm the scalability of our RL framework, revealing an optimization speed that is independent of system size. This work thus enables a new paradigm: a quantum computer that learns from its errors and never stops computing. By integrating reinforcement learning with quantum error correction, a quantum computer continuously self-calibrates during computation, achieving record logical error rates and enhanced resilience to drift.

Read PDF