MEMORY INDUSTRY INTELLIGENCE
오류로부터 학습하는 양자 컴퓨터를 향하여
한국어 번역·요약·분석
원문 제목: Towards a quantum computer that learns from its errors
핵심 요약
구글 리서치는 양자 컴퓨터의 제어 파라미터 드리프트 문제를 해결하기 위해 양자 오류 검출 데이터를 강화학습(RL)의 능동적 학습 신호로 활용하는 프레임워크를 Nature에 발표했다. 이들은 플래그십 초전도 프로세서 Willow에서 인위적 드리프트를 주입한 실험을 통해 RL 제어가 논리적 안정성을 3.5배 향상시키고, 전문가 캘리브레이션 후에도 논리 오류율을 추가로 20% 억제했으며, 표면 코드에서 1,000회 오류 정정 주기당 1회 미만, 컬러 코드에서 100회당 1회의 기록적 낮은 논리 오류를 달성했다고 밝혔다. 수백 큐비트와 수만 개 제어 파라미터의 수치 시뮬레이션에서 필요한 RL 훈련 반복 횟수가 시스템 크기와 무관함을 확인했으나, 완전한 잠재력 실현을 위해서는 에이전트와 양자 프로세서 간 통신 주기 단축 및 고도화된 머신러닝 기법 적용이 필요하다고 언급했다. 이 연구는 오류로부터 학습하며 계산을 멈추지 않는 양자 컴퓨터라는 새 패러다임을 제시한다. 다만 이는 양자 컴퓨팅 제어 소프트웨어·알고리즘 연구로, 메모리 산업과의 직접적 연결 근거는 원문에서 확인되지 않는다.
메모리 산업 영향 분석
이 문서는 양자 컴퓨터의 제어 보정과 오류 정정을 강화학습으로 개선하는 구글의 연구로, 메모리 산업과의 직접적 연결 근거는 원문에서 확인되지 않는다. 양자 컴퓨터는 고전적 메모리 계층(HBM·DRAM·NAND·eSSD 등)을 사용하지 않는 별도 컴퓨팅 패러다임이며, 원문에는 메모리 제품·고객·거래 관계에 대한 언급이 전혀 없다. 따라서 관련 제품·고객·직접 효과·계층 이동·사용량 효과·적용 범위를 원문 근거로 도출할 수 없다. 분석가 가설로는 장기적으로 양자 컴퓨팅이 상용화될 경우 고전 컴퓨팅 수요 일부를 대체하거나 보완할 수 있다는 정도의 간접적 시나리오가 가능하나, 이는 원문이 뒷받침하지 않는 미확인 가정이다. 확인할 지표로는 양자 프로세서의 논리 큐비트 수·오류율·상용화 시점 등이 있으나, 이는 메모리 산업 지표가 아니다.
한국어 번역 읽기
수집된 원문 v2의 전체 본문 기준 · 8269자
복잡한 명곡을 연주하는 심포니 오케스트라를 상상해 보자. 만약 바이올린이 몇 마디마다 음정이 어긋난다면, 앙상블은 계속 멈추고 악기를 조율해야 할 것이다. 다행히 오케스트라에서는 악기가 안정적으로 음정을 유지하기 때문에 이런 일은 일어나지 않는다. 그러나 이것이 바로 양자 컴퓨터 운영의 현재 현실이다.
양자 컴퓨터는 근본적으로 드리프트에 민감한 아날로그 기계이므로, 신뢰할 수 있는 운영을 유지하려면 제어 파라미터, 즉 큐비트를 안무하는 아날로그 신호의 주파수, 진폭, 위상을 끊임없이 재보정해야 한다. 오늘날 이는 전체 양자 계산을 완전히 종료해야 한다. 계산과 보정의 이러한 완전한 분리는 미래의 근본적 병목으로, 유용한 양자 알고리즘은 며칠 또는 몇 달 동안 연속적으로 실행되어야 하기 때문이다.
이를 해결하기 위해, [_Nature_](https://www.nature.com/)에 게재된 “[Reinforcement learning control of quantum error correction](https://www.nature.com/articles/s41586-026-10759-2)”에서 우리는 자율 에이전트가 양자 오류 검출로부터 학습하여 수천 개의 제어 파라미터를 지속적으로 조정하고 계산 중 드리프트에 대해 양자 시스템을 안정화하는 [강화학습](https://en.wikipedia.org/wiki/Reinforcement_learning)(RL) 프레임워크를 시연했다. 요컨대, 우리는 음악이 연주되는 동안 악기를 조율하는 방법을 찾았다.
## 양자 오류 다루기
콘서트홀에서는 음정이 어긋난 악기가 즉시 들린다. 양자 영역은 그런 사치를 제공하지 않는다. 듣는 행위 자체가 공연을 망치는 것처럼, 큐비트를 측정하면 양자 중첩 상태가 붕괴된다. 양자 정보를 보존하기 위해 우리는 대신 [양자 오류 정정](https://research.google/blog/making-quantum-error-correction-work/)(QEC)을 사용하는데, 이는 중복성을 활용하여 다수의 물리적 큐비트로부터 “논리적 큐비트”를 만들고, 물리적 큐비트에 대한 특수한 패리티 검사를 사용하여 아날로그 잡음을 이진 오류 검출 이벤트로 디지털화하는 기술이다.
불행히도 이 비트들은 양자 회로의 제한된 시공간 영역 어딘가에서 오류가 발생했다는 것만 알려줄 뿐, 정확한 위치는 알려주지 않는다. 이는 어느 연주자가 연주했는지 정확히 모른 채 삑사리 소리를 듣는 것과 같다. 가능한 오류 위치를 특정하고 필요한 정정을 계산하기 위해 우리는 신경망 디코더 [AlphaQubit](https://blog.google/innovation-and-ai/models-and-research/google-deepmind/alphaqubit-quantum-error-correction/)(실제 데이터로 훈련됨)과 알고리즘 디코더 [Tesseract](https://arxiv.org/html/2503.10988v1)와 같은 QEC 디코더에 의존한다. 오류가 충분히 드물면 이 디코더들은 오류 검출 데이터를 분석하여 논리적 양자 정보를 성공적으로 복원할 수 있다. 그러나 디코더는 결정적인 질문 하나에 답하지 않는다. 왜 그 오류들이 애초에 발생했는가?
일부 오류는 양자 시스템과 주변 환경의 불가피한 상호작용에서 비롯되며, 이는 [결어긋남](https://physicstoday.aip.org/features/decoherence-and-the-transition-from-quantum-to-classical)으로 이어진다. 이 무자비한 과정은 거시적 양자 중첩을 파괴하여 양자 컴퓨터를 사실상 고전 컴퓨터로 만든다. 이 근본적 현상은 너무나 만연하여 우리에게 익숙한 고전적 실재가 자연의 근저 양자 법칙으로부터 출현하게 한다. 이러한 환경적 오류는 완전히 방지할 수 없지만, 다른 많은 오류는 부정확한 제어 보정과 하드웨어 드리프트의 발현이며, 이는 우리가 완화할 수 있는 결함이다.
## 전통적 물리 모델을 넘어서
전통적으로 양자 보정은 물리 모델에 의존했다. 그 기술은 수십 년간의 양자 제어 연구를 통해 정교해졌다. 그러나 기술 영역 전반에서 인간이 만든 모델은 필연적으로 성능 한계에 부딪힌다. 초기 컴퓨터 비전은 엄격한 기하학적 규칙에 의존하다가 [정체](https://en.wikipedia.org/wiki/ImageNet#History_of_the_ImageNet_challenge)되었다. [전통적 로봇공학](https://arxiv.org/abs/2506.13498)은 접촉 역학과 마찰의 지저분한 현실을 포착하지 못하는 운동학 방정식으로 여전히 어려움을 겪고 있다. 마찬가지로 단백질 접힘 예측이라는 수십 년 된 난제는 [AlphaFold](https://deepmind.google/blog/alphafold-five-years-of-impact/)와 같은 딥러닝 시스템이 전례 없는 정확도를 달성할 때까지 전통적 물리 모델로는 대부분 해결 불가능했다. 이 분야들에서 새로운 돌파구는 데이터로부터 직접 학습하는 접근으로 전환할 때 발생했다.
최근 [AlphaQubit](https://blog.google/innovation-and-ai/models-and-research/google-deepmind/alphaqubit-quantum-error-correction/)은 가장 강력한 알고리즘 QEC 디코더의 정확도를 능가했다. 이제 양자 제어도 같은 한계에 직면했다. 양자 프로세서가 제조 및 하드웨어의 발전을 통해 개선됨에 따라, 그 오류는 전통적 모델링과 보정에 어려운 복잡한 현상에 의해 지배되게 된다. 머신러닝이 새로운 진전을 가져올 수 있을까?
## 오류를 정정만 하지 말고, 오류로부터 학습하라!
구글 리서치는 전통적 프로그래밍으로는 너무 복잡한 문제를 해결하기 위해 RL을 개척한 풍부한 역사를 가지고 있다. 명시적 지시에 의존하는 알고리즘과 달리, RL은 경험을 통해 작동한다. 자율 에이전트는 다양한 행동을 시험하고 그 결과 오류로부터 직접 학습하여 전략을 개선한다. 정확하고 연속적인 양자 보정을 달성하기 위해 RL을 적용하는 것은 거의 불가피하게 느껴졌다. QEC가 이미 꾸준한 검출 이벤트 흐름을 생성하므로, 우리는 이 데이터에 보완적 역할을 부여했다. 이를 디코딩하고 오류를 정정하는 것 외에도, 우리는 검출 이벤트를 능동적 학습 신호로 사용한다. 계산이 진행됨에 따라 RL 에이전트는 이 데이터를 모니터링하고 제어 파라미터를 동적으로 조정하여 드리프트에 대응하고 새로운 오류를 방지하는 법을 학습한다.
## 우리의 양자 제어 실험
우리는 플래그십 [Willow](https://blog.google/innovation-and-ai/technology/research/google-willow-quantum-chip/) 초전도 프로세서에서 이러한 RL 양자 제어를 검증했다. 제어 파라미터의 인위적 드리프트를 의도적으로 주입함으로써, RL 조정이 오류 정정 코드의 논리적 안정성을 3.5배 향상시켜 프로세서가 신뢰할 수 있는 “양자 메모리” 장치로 작동하는 시간을 연장함을 보였다.
일반적으로 프로세서를 최고 성능으로 조율하는 것은 “인간 개입” 접근에 크게 의존했으며, 과학자들이 자동화하기 어려운 예외적 경우를 해결하기 위해 물리적 직관을 적용했다. 그러나 이 철저한 전문가 보정 후에도, 후속 RL 미세 조정은 논리 오류율을 추가로 20% 체계적으로 억제했다.
이 실험에서 우리 기술의 종합은 양자 메모리의 논리 오류를 기록적 최저 수준으로 감소시켰다. [표면 코드](https://research.google/blog/making-quantum-error-correction-work/)에서 오류 정정 주기 1,000회당 1회 미만, [컬러 코드](https://research.google/blog/a-colorful-quantum-future/)에서 100회당 1회이다.
## 확장되는가?
머신러닝 및 양자 연구자들의 중요한 질문은 이 RL 접근이 미래의 대형 양자 컴퓨터로 확장될 수 있는지이다. 이를 테스트하기 위해 우리는 수백 큐비트와 수만 개의 제어 파라미터로 수치 시뮬레이션을 수행했다.
시뮬레이션은 우리의 예상을 확인했다. 필요한 RL 훈련 반복 횟수(에포크)는 오류에 대한 QEC 검출 이벤트의 국소적 민감성 덕분에 시스템 크기와 무관하다. 그러나 RL 프레임워크의 완전한 잠재력을 실현하려면 더 긴밀한 통합이 필요하다. 에이전트와 양자 프로세서 간의 통신 주기를 가속화하고 더 발전된 머신러닝 방법을 사용함으로써, 우리는 상당한 추가 개선을 열기를 희망한다.
따라서 우리의 연구는 새로운 패러다임을 가능하게 한다. 오류로부터 학습하고 계산을 멈추지 않는 양자 컴퓨터이다.
## 감사의 글
_하드웨어, 소프트웨어, 극저온 및 전자 인프라를 구축하고 유지하는 것을 포함한 공동 저자들의 기여에 감사드립니다. 이 작업은 구글 리서치의 Google Quantum AI 팀이 Google DeepMind 팀과 협력하여 가능하게 했습니다._
양자 컴퓨터는 근본적으로 드리프트에 민감한 아날로그 기계이므로, 신뢰할 수 있는 운영을 유지하려면 제어 파라미터, 즉 큐비트를 안무하는 아날로그 신호의 주파수, 진폭, 위상을 끊임없이 재보정해야 한다. 오늘날 이는 전체 양자 계산을 완전히 종료해야 한다. 계산과 보정의 이러한 완전한 분리는 미래의 근본적 병목으로, 유용한 양자 알고리즘은 며칠 또는 몇 달 동안 연속적으로 실행되어야 하기 때문이다.
이를 해결하기 위해, [_Nature_](https://www.nature.com/)에 게재된 “[Reinforcement learning control of quantum error correction](https://www.nature.com/articles/s41586-026-10759-2)”에서 우리는 자율 에이전트가 양자 오류 검출로부터 학습하여 수천 개의 제어 파라미터를 지속적으로 조정하고 계산 중 드리프트에 대해 양자 시스템을 안정화하는 [강화학습](https://en.wikipedia.org/wiki/Reinforcement_learning)(RL) 프레임워크를 시연했다. 요컨대, 우리는 음악이 연주되는 동안 악기를 조율하는 방법을 찾았다.
## 양자 오류 다루기
콘서트홀에서는 음정이 어긋난 악기가 즉시 들린다. 양자 영역은 그런 사치를 제공하지 않는다. 듣는 행위 자체가 공연을 망치는 것처럼, 큐비트를 측정하면 양자 중첩 상태가 붕괴된다. 양자 정보를 보존하기 위해 우리는 대신 [양자 오류 정정](https://research.google/blog/making-quantum-error-correction-work/)(QEC)을 사용하는데, 이는 중복성을 활용하여 다수의 물리적 큐비트로부터 “논리적 큐비트”를 만들고, 물리적 큐비트에 대한 특수한 패리티 검사를 사용하여 아날로그 잡음을 이진 오류 검출 이벤트로 디지털화하는 기술이다.
불행히도 이 비트들은 양자 회로의 제한된 시공간 영역 어딘가에서 오류가 발생했다는 것만 알려줄 뿐, 정확한 위치는 알려주지 않는다. 이는 어느 연주자가 연주했는지 정확히 모른 채 삑사리 소리를 듣는 것과 같다. 가능한 오류 위치를 특정하고 필요한 정정을 계산하기 위해 우리는 신경망 디코더 [AlphaQubit](https://blog.google/innovation-and-ai/models-and-research/google-deepmind/alphaqubit-quantum-error-correction/)(실제 데이터로 훈련됨)과 알고리즘 디코더 [Tesseract](https://arxiv.org/html/2503.10988v1)와 같은 QEC 디코더에 의존한다. 오류가 충분히 드물면 이 디코더들은 오류 검출 데이터를 분석하여 논리적 양자 정보를 성공적으로 복원할 수 있다. 그러나 디코더는 결정적인 질문 하나에 답하지 않는다. 왜 그 오류들이 애초에 발생했는가?
일부 오류는 양자 시스템과 주변 환경의 불가피한 상호작용에서 비롯되며, 이는 [결어긋남](https://physicstoday.aip.org/features/decoherence-and-the-transition-from-quantum-to-classical)으로 이어진다. 이 무자비한 과정은 거시적 양자 중첩을 파괴하여 양자 컴퓨터를 사실상 고전 컴퓨터로 만든다. 이 근본적 현상은 너무나 만연하여 우리에게 익숙한 고전적 실재가 자연의 근저 양자 법칙으로부터 출현하게 한다. 이러한 환경적 오류는 완전히 방지할 수 없지만, 다른 많은 오류는 부정확한 제어 보정과 하드웨어 드리프트의 발현이며, 이는 우리가 완화할 수 있는 결함이다.
## 전통적 물리 모델을 넘어서
전통적으로 양자 보정은 물리 모델에 의존했다. 그 기술은 수십 년간의 양자 제어 연구를 통해 정교해졌다. 그러나 기술 영역 전반에서 인간이 만든 모델은 필연적으로 성능 한계에 부딪힌다. 초기 컴퓨터 비전은 엄격한 기하학적 규칙에 의존하다가 [정체](https://en.wikipedia.org/wiki/ImageNet#History_of_the_ImageNet_challenge)되었다. [전통적 로봇공학](https://arxiv.org/abs/2506.13498)은 접촉 역학과 마찰의 지저분한 현실을 포착하지 못하는 운동학 방정식으로 여전히 어려움을 겪고 있다. 마찬가지로 단백질 접힘 예측이라는 수십 년 된 난제는 [AlphaFold](https://deepmind.google/blog/alphafold-five-years-of-impact/)와 같은 딥러닝 시스템이 전례 없는 정확도를 달성할 때까지 전통적 물리 모델로는 대부분 해결 불가능했다. 이 분야들에서 새로운 돌파구는 데이터로부터 직접 학습하는 접근으로 전환할 때 발생했다.
최근 [AlphaQubit](https://blog.google/innovation-and-ai/models-and-research/google-deepmind/alphaqubit-quantum-error-correction/)은 가장 강력한 알고리즘 QEC 디코더의 정확도를 능가했다. 이제 양자 제어도 같은 한계에 직면했다. 양자 프로세서가 제조 및 하드웨어의 발전을 통해 개선됨에 따라, 그 오류는 전통적 모델링과 보정에 어려운 복잡한 현상에 의해 지배되게 된다. 머신러닝이 새로운 진전을 가져올 수 있을까?
## 오류를 정정만 하지 말고, 오류로부터 학습하라!
구글 리서치는 전통적 프로그래밍으로는 너무 복잡한 문제를 해결하기 위해 RL을 개척한 풍부한 역사를 가지고 있다. 명시적 지시에 의존하는 알고리즘과 달리, RL은 경험을 통해 작동한다. 자율 에이전트는 다양한 행동을 시험하고 그 결과 오류로부터 직접 학습하여 전략을 개선한다. 정확하고 연속적인 양자 보정을 달성하기 위해 RL을 적용하는 것은 거의 불가피하게 느껴졌다. QEC가 이미 꾸준한 검출 이벤트 흐름을 생성하므로, 우리는 이 데이터에 보완적 역할을 부여했다. 이를 디코딩하고 오류를 정정하는 것 외에도, 우리는 검출 이벤트를 능동적 학습 신호로 사용한다. 계산이 진행됨에 따라 RL 에이전트는 이 데이터를 모니터링하고 제어 파라미터를 동적으로 조정하여 드리프트에 대응하고 새로운 오류를 방지하는 법을 학습한다.
## 우리의 양자 제어 실험
우리는 플래그십 [Willow](https://blog.google/innovation-and-ai/technology/research/google-willow-quantum-chip/) 초전도 프로세서에서 이러한 RL 양자 제어를 검증했다. 제어 파라미터의 인위적 드리프트를 의도적으로 주입함으로써, RL 조정이 오류 정정 코드의 논리적 안정성을 3.5배 향상시켜 프로세서가 신뢰할 수 있는 “양자 메모리” 장치로 작동하는 시간을 연장함을 보였다.
일반적으로 프로세서를 최고 성능으로 조율하는 것은 “인간 개입” 접근에 크게 의존했으며, 과학자들이 자동화하기 어려운 예외적 경우를 해결하기 위해 물리적 직관을 적용했다. 그러나 이 철저한 전문가 보정 후에도, 후속 RL 미세 조정은 논리 오류율을 추가로 20% 체계적으로 억제했다.
이 실험에서 우리 기술의 종합은 양자 메모리의 논리 오류를 기록적 최저 수준으로 감소시켰다. [표면 코드](https://research.google/blog/making-quantum-error-correction-work/)에서 오류 정정 주기 1,000회당 1회 미만, [컬러 코드](https://research.google/blog/a-colorful-quantum-future/)에서 100회당 1회이다.
## 확장되는가?
머신러닝 및 양자 연구자들의 중요한 질문은 이 RL 접근이 미래의 대형 양자 컴퓨터로 확장될 수 있는지이다. 이를 테스트하기 위해 우리는 수백 큐비트와 수만 개의 제어 파라미터로 수치 시뮬레이션을 수행했다.
시뮬레이션은 우리의 예상을 확인했다. 필요한 RL 훈련 반복 횟수(에포크)는 오류에 대한 QEC 검출 이벤트의 국소적 민감성 덕분에 시스템 크기와 무관하다. 그러나 RL 프레임워크의 완전한 잠재력을 실현하려면 더 긴밀한 통합이 필요하다. 에이전트와 양자 프로세서 간의 통신 주기를 가속화하고 더 발전된 머신러닝 방법을 사용함으로써, 우리는 상당한 추가 개선을 열기를 희망한다.
따라서 우리의 연구는 새로운 패러다임을 가능하게 한다. 오류로부터 학습하고 계산을 멈추지 않는 양자 컴퓨터이다.
## 감사의 글
_하드웨어, 소프트웨어, 극저온 및 전자 인프라를 구축하고 유지하는 것을 포함한 공동 저자들의 기여에 감사드립니다. 이 작업은 구글 리서치의 Google Quantum AI 팀이 Google DeepMind 팀과 협력하여 가능하게 했습니다._
브리프용 요약 초안
구글 리서치가 Nature에 양자 오류 검출 데이터를 강화학습 신호로 활용해 계산 중 제어 파라미터 드리프트를 보정하는 프레임워크를 발표했다. Willow 프로세서 실험에서 논리적 안정성 3.5배 향상, 논리 오류율 추가 20% 억제, 표면 코드 1,000주기당 1회 미만의 기록적 낮은 오류를 달성했다. 다만 이는 양자 컴퓨팅 제어 연구로, 메모리 산업과의 직접적 연결 근거는 원문에서 확인되지 않는다.
원문 텍스트
원문 열기 ↗Imagine a symphony orchestra performing a complex masterpiece. If the violins drifted out of tune every few measures, the ensemble would constantly have to stop and retune their instruments. Thankfully, this doesn't happen in an orchestra because the instruments reliably stay in tune. However, it is the current reality of operating a quantum computer.
Since quantum computers are fundamentally analog machines that are sensitive to drift, maintaining reliable operation requires perpetually recalibrating their control parameters, i.e., the frequencies, amplitudes, and phases of the analog signals choreographing the qubits. Today, this requires fully terminating the entire quantum computation. This complete decoupling of computation and calibration represents a fundamental bottleneck for the future, as useful quantum algorithms must run continuously for days or even months.
To address this, in “[Reinforcement learning control of quantum error correction](https://www.nature.com/articles/s41586-026-10759-2)”, published in [_Nature_](https://www.nature.com/), we demonstrated a [reinforcement learning](https://en.wikipedia.org/wiki/Reinforcement_learning) (RL) framework in which an autonomous agent learns from quantum error detections to continuously steer thousands of control parameters, stabilizing the quantum system against drift during the computation. In short: we found a way to tune the instruments while the music plays.
## Dealing with quantum errors
In a concert hall, a detuned instrument is immediately heard. The quantum realm offers no such luxury. As if the very act of listening ruined the performance, measuring the qubits collapses their quantum superposition states. To preserve the quantum information, we instead employ [Quantum Error Correction](https://research.google/blog/making-quantum-error-correction-work/) (QEC), a technique that exploits redundancy to create “logical qubits” out of many physical qubits, and uses specialized parity checks on the physical qubits to digitize the analog noise into binary error detection events.
Unfortunately, these bits only tell us that an error occurred somewhere within a bounded spacetime region of the quantum circuit, not its exact location. It is like hearing a sour note without knowing exactly which musician played it. To pinpoint the likely error locations and calculate the necessary corrections, we rely on QEC decoders, such as the neural network decoder [AlphaQubit](https://blog.google/innovation-and-ai/models-and-research/google-deepmind/alphaqubit-quantum-error-correction/) (trained on real data) and algorithmic decoder [Tesseract](https://arxiv.org/html/2503.10988v1). If errors are sufficiently rare, these decoders can successfully restore the logical quantum information by analyzing the error detection data. However, decoders leave a crucial question unanswered: why did those errors happen in the first place?
Some errors result from the unavoidable interaction of a quantum system with its surrounding environment, leading to [decoherence](https://physicstoday.aip.org/features/decoherence-and-the-transition-from-quantum-to-classical). This ruthless process destroys macroscopic quantum superpositions, effectively turning quantum computers into classical ones. This fundamental phenomenon is so pervasive that it causes our familiar classical reality to emerge from the underlying quantum laws of Nature. While these environmental errors can never be completely prevented, many others are manifestations of imprecise control calibration and hardware drift – flaws that remain within our power to mitigate.
## Beyond traditional physics models
Traditionally, quantum calibration relied on physics models. Its techniques were refined through decades of quantum control research. However, across technological domains, human-crafted models inevitably hit a performance ceiling. Early computer vision [stalled](https://en.wikipedia.org/wiki/ImageNet#History_of_the_ImageNet_challenge) when relying on strict geometric rules. [Traditional robotics](https://arxiv.org/abs/2506.13498) still struggles with kinematic equations that fail to capture the messy reality of contact dynamics and friction. Similarly, the decades-old challenge of predicting protein folding remained largely intractable for traditional physical models until deep learning systems like [AlphaFold](https://deepmind.google/blog/alphafold-five-years-of-impact/) achieved unprecedented accuracy. Across these fields, new breakthroughs occurred when the approach shifted toward learning directly from data.
Recently, [AlphaQubit](https://blog.google/innovation-and-ai/models-and-research/google-deepmind/alphaqubit-quantum-error-correction/) surpassed the accuracy of the most powerful algorithmic QEC decoders. Now quantum control faces the same ceiling. As quantum processors improve through progress in fabrication and hardware, their errors become dominated by complex phenomena that are challenging for traditional modeling and calibration. Can machine learning bring new advances?
## Don’t just correct errors, learn from them!
Google Research has a rich history of pioneering RL to solve problems too complex for traditional programming. Unlike algorithms relying on explicit instructions, RL operates through experience. An autonomous agent tests different behaviors and learns directly from resulting errors to refine its strategy. Applying RL to achieve accurate, continuous quantum calibration felt almost inevitable. Since QEC already generates a steady stream of detection events, we simply granted this data a complementary role. In addition to decoding it and correcting the errors, we employ the detection events as an active learning signal. As computation progresses, the RL agent monitors this data and learns to dynamically steer the control parameters, counteracting drift and preventing new errors.
## Our quantum control experiment
We validated such RL quantum control on our flagship [Willow](https://blog.google/innovation-and-ai/technology/research/google-willow-quantum-chip/) superconducting processor. By deliberately injecting artificial drift of control parameters, we showed that RL steering improved the logical stability of our error-correcting code 3.5-fold, prolonging the time during which the processor acts as a reliable “quantum memory” device.
Typically, tuning the processor to peak performance relied heavily on a “human-in-the-loop” approach, with scientists applying physical intuition to resolve edge cases that are difficult to automate. Yet, even after this exhaustive expert calibration, subsequent RL fine-tuning systematically suppressed the logical error rate by an additional 20%.
The synthesis of all our technologies in this experiment reduced the logical errors in quantum memories to a record low: fewer than one per thousand error correction cycles in the [surface code](https://research.google/blog/making-quantum-error-correction-work/), and one per hundred in the [color code](https://research.google/blog/a-colorful-quantum-future/).
## Does it scale?
A critical question from machine learning and quantum researchers is whether this RL approach can scale to large quantum computers of the future. To test this, we conducted numerical simulations with hundreds of qubits and tens of thousands of control parameters.
The simulations confirmed our expectation: the number of required RL training iterations (epochs) is independent of the system size, owing to the local sensitivity of the QEC detection events to errors. However, realizing the full potential of the RL framework requires tighter integration. By speeding up the communication cycle between the agent and the quantum processor, and employing more advanced machine learning methods, we hope to unlock significant additional improvements.
Our work thus enables a new paradigm: a quantum computer that learns from its errors and doesn’t stop computing.
## Acknowledgements
_We thank our co-authors for their contributions, including building and maintaining the hardware, software, cryogenics and electronics infrastructure. This work was made possible by the Google Quantum AI team at Google Research, in collaboration with teams from Google DeepMind._
Since quantum computers are fundamentally analog machines that are sensitive to drift, maintaining reliable operation requires perpetually recalibrating their control parameters, i.e., the frequencies, amplitudes, and phases of the analog signals choreographing the qubits. Today, this requires fully terminating the entire quantum computation. This complete decoupling of computation and calibration represents a fundamental bottleneck for the future, as useful quantum algorithms must run continuously for days or even months.
To address this, in “[Reinforcement learning control of quantum error correction](https://www.nature.com/articles/s41586-026-10759-2)”, published in [_Nature_](https://www.nature.com/), we demonstrated a [reinforcement learning](https://en.wikipedia.org/wiki/Reinforcement_learning) (RL) framework in which an autonomous agent learns from quantum error detections to continuously steer thousands of control parameters, stabilizing the quantum system against drift during the computation. In short: we found a way to tune the instruments while the music plays.
## Dealing with quantum errors
In a concert hall, a detuned instrument is immediately heard. The quantum realm offers no such luxury. As if the very act of listening ruined the performance, measuring the qubits collapses their quantum superposition states. To preserve the quantum information, we instead employ [Quantum Error Correction](https://research.google/blog/making-quantum-error-correction-work/) (QEC), a technique that exploits redundancy to create “logical qubits” out of many physical qubits, and uses specialized parity checks on the physical qubits to digitize the analog noise into binary error detection events.
Unfortunately, these bits only tell us that an error occurred somewhere within a bounded spacetime region of the quantum circuit, not its exact location. It is like hearing a sour note without knowing exactly which musician played it. To pinpoint the likely error locations and calculate the necessary corrections, we rely on QEC decoders, such as the neural network decoder [AlphaQubit](https://blog.google/innovation-and-ai/models-and-research/google-deepmind/alphaqubit-quantum-error-correction/) (trained on real data) and algorithmic decoder [Tesseract](https://arxiv.org/html/2503.10988v1). If errors are sufficiently rare, these decoders can successfully restore the logical quantum information by analyzing the error detection data. However, decoders leave a crucial question unanswered: why did those errors happen in the first place?
Some errors result from the unavoidable interaction of a quantum system with its surrounding environment, leading to [decoherence](https://physicstoday.aip.org/features/decoherence-and-the-transition-from-quantum-to-classical). This ruthless process destroys macroscopic quantum superpositions, effectively turning quantum computers into classical ones. This fundamental phenomenon is so pervasive that it causes our familiar classical reality to emerge from the underlying quantum laws of Nature. While these environmental errors can never be completely prevented, many others are manifestations of imprecise control calibration and hardware drift – flaws that remain within our power to mitigate.
## Beyond traditional physics models
Traditionally, quantum calibration relied on physics models. Its techniques were refined through decades of quantum control research. However, across technological domains, human-crafted models inevitably hit a performance ceiling. Early computer vision [stalled](https://en.wikipedia.org/wiki/ImageNet#History_of_the_ImageNet_challenge) when relying on strict geometric rules. [Traditional robotics](https://arxiv.org/abs/2506.13498) still struggles with kinematic equations that fail to capture the messy reality of contact dynamics and friction. Similarly, the decades-old challenge of predicting protein folding remained largely intractable for traditional physical models until deep learning systems like [AlphaFold](https://deepmind.google/blog/alphafold-five-years-of-impact/) achieved unprecedented accuracy. Across these fields, new breakthroughs occurred when the approach shifted toward learning directly from data.
Recently, [AlphaQubit](https://blog.google/innovation-and-ai/models-and-research/google-deepmind/alphaqubit-quantum-error-correction/) surpassed the accuracy of the most powerful algorithmic QEC decoders. Now quantum control faces the same ceiling. As quantum processors improve through progress in fabrication and hardware, their errors become dominated by complex phenomena that are challenging for traditional modeling and calibration. Can machine learning bring new advances?
## Don’t just correct errors, learn from them!
Google Research has a rich history of pioneering RL to solve problems too complex for traditional programming. Unlike algorithms relying on explicit instructions, RL operates through experience. An autonomous agent tests different behaviors and learns directly from resulting errors to refine its strategy. Applying RL to achieve accurate, continuous quantum calibration felt almost inevitable. Since QEC already generates a steady stream of detection events, we simply granted this data a complementary role. In addition to decoding it and correcting the errors, we employ the detection events as an active learning signal. As computation progresses, the RL agent monitors this data and learns to dynamically steer the control parameters, counteracting drift and preventing new errors.
## Our quantum control experiment
We validated such RL quantum control on our flagship [Willow](https://blog.google/innovation-and-ai/technology/research/google-willow-quantum-chip/) superconducting processor. By deliberately injecting artificial drift of control parameters, we showed that RL steering improved the logical stability of our error-correcting code 3.5-fold, prolonging the time during which the processor acts as a reliable “quantum memory” device.
Typically, tuning the processor to peak performance relied heavily on a “human-in-the-loop” approach, with scientists applying physical intuition to resolve edge cases that are difficult to automate. Yet, even after this exhaustive expert calibration, subsequent RL fine-tuning systematically suppressed the logical error rate by an additional 20%.
The synthesis of all our technologies in this experiment reduced the logical errors in quantum memories to a record low: fewer than one per thousand error correction cycles in the [surface code](https://research.google/blog/making-quantum-error-correction-work/), and one per hundred in the [color code](https://research.google/blog/a-colorful-quantum-future/).
## Does it scale?
A critical question from machine learning and quantum researchers is whether this RL approach can scale to large quantum computers of the future. To test this, we conducted numerical simulations with hundreds of qubits and tens of thousands of control parameters.
The simulations confirmed our expectation: the number of required RL training iterations (epochs) is independent of the system size, owing to the local sensitivity of the QEC detection events to errors. However, realizing the full potential of the RL framework requires tighter integration. By speeding up the communication cycle between the agent and the quantum processor, and employing more advanced machine learning methods, we hope to unlock significant additional improvements.
Our work thus enables a new paradigm: a quantum computer that learns from its errors and doesn’t stop computing.
## Acknowledgements
_We thank our co-authors for their contributions, including building and maintaining the hardware, software, cryogenics and electronics infrastructure. This work was made possible by the Google Quantum AI team at Google Research, in collaboration with teams from Google DeepMind._