TR2026-130
Reinforcement-Learning-Guided Multi-Branch Decoding of Quantum LDPC Codes
-
- , "Reinforcement-Learning-Guided Multi-Branch Decoding of Quantum LDPC Codes", IEEE International Conference on Quantum Computing and Engineering (QCE), September 2026.BibTeX TR2026-130 PDF
- @inproceedings{Nourozi2026sep2,
- author = {Nourozi, Vahid and Koike-Akino, Toshiaki and Mitchell, David},
- title = {{Reinforcement-Learning-Guided Multi-Branch Decoding of Quantum LDPC Codes}},
- booktitle = {IEEE International Conference on Quantum Computing and Engineering (QCE)},
- year = 2026,
- month = sep,
- url = {https://www.merl.com/publications/TR2026-130}
- }
- , "Reinforcement-Learning-Guided Multi-Branch Decoding of Quantum LDPC Codes", IEEE International Conference on Quantum Computing and Engineering (QCE), September 2026.
-
MERL Contact:
-
Research Area:
Abstract:
Belief propagation (BP) is attractive for quantum lowdensity parity-check (QLDPC) codes, yet short cycles, degeneracy, and trapping configurations can cause oscillation, nonconvergence, or convergence to an incorrect logical class. We propose RL-MBOSD, a reinforcement-learning-guided multi-branch decoder for QLDPC codes. For each code matrix, a separately trained 28- parameter linear action-value policy is shared across its variable nodes and ranks sequential message updates using degreenormalized local and global features. Deterministic score perturbations generate B complementary trajectories, each executed for at most T outer sweeps. Bounded component-wise OSD repairs selected stalled branches; the resulting X- and Z-component lists are paired and ranked using the joint negative log-likelihood under the Pauli channel. Syndrome-valid candidates are then grouped by logical equivalence class and selected using an aggregate logicalclass score. On the [[144, 12, 12]] bivariate-bicycle code and the A5 code, the proposed decoder gives lower error-rate point estimates than the re-plotted prior baselines.
