Skip to content

Reinforcement Learning for Compiled-CNOT-Efficient VQE Circuits

· 0 citations · 10 references

TL;DR

This work asks whether reinforcement learning (RL) can design circuits that reach a fixed accuracy target with fewer compiled CNOTs than strong hand-designed and greedy baselines than strong hand-designed and greedy baselines.

View source

Similar papers

Preprint Aug 2026

Automating Variational Quantum Sensing through Reinforcement-Learned Circuit Structures

Variational quantum sensing offers a promising route to high-precision parameter estimation, but its performance depends strongly on the circuit architectures used for probe preparation and measurement. Existing approaches typically optimize continuous parameters within predefined ans\"atze, restricting the accessible design space and limiting adaptation to sensing tasks and hardware constraints. Here, we introduce \textsc{AutoQSense}, a reinforcement-learning framework that searches circuit architectures using Fisher-information-based objectives. For few-qubit systems, a single agent sequentially constructs preparation and measurement circuits. For larger systems, a distributed formulation assigns local circuit design to subsystem agents and inter-block entanglement to a budgeted agent. Numerical results show that the learned architectures recover known benchmark strategies, adapt to dephasing noise, and outperform fixed hardware-efficient ans\"atze while using fewer entangling gates. These results establish \textsc{AutoQSense} as a resource-aware approach to adaptive and hardware-compatible quantum sensing.

Jie Liu, Xin Wang · 0 citations
Preprint Jul 2026

Variational Learning with Sparse Long-range Entangling Gates

This work examines when structured long-range connectivity provides a useful resource, focusing on sparse power-of-two (PWR2) coupling graphs, and identifies circuit geometry and qubit reconfigurability as task-dependent resources for variational algorithms.

Helene M. Losl, Aydin Deger, Andrew J. Daley · 0 citations
#artificial intelligence Preprint Aug 2026

AlphaClifford: Efficient Clifford Synthesis and Transpilation with Model-based RL

AlphaClifford is introduced, a model-based Reinforcement Learning framework designed to efficiently synthesize Clifford circuits from the fundamental gate set composed of H, S, and CNOT, demonstrating the broad applicability of the framework on two additional tasks: hardware-constrained Clifford transpilation, where it outperform existing RL-based compilers, and as a post-synthesis optimization component within a full Clifford+T logical synthesis pipeline.

Daniele Lizzio Bosco, Jacopo Cossio, Carla Piazza et al. · 0 citations
Preprint Jul 2026

Approximate Quantum State Preparation Through Proximal Policy Optimization

In this work, a quantum architecture search framework for approximate quantum state preparation (QSP) is proposed. QSP is a challenging task, since the search space grows exponentially with the number of qubits, making the identification of the optimal circuit non-trivial. To address this problem, deep reinforcement learning is employed through an agent based on proximal policy optimization. The objective of the agent is to identify the best possible approximation of the target state while simultaneously minimizing the number of gates used. At each step, the agent appends a new gate to the circuit and recomputes the fidelity between the approximated state and the target states. Various experiments have been performed from 2 to 5 qubits. Both predefined states, such as Bell, GHZ, W, and Dicke states, and completely random states are considered. The proposed framework is able to achieve approximation errors of $10^{-14}$.

Marco Mordacci, Michele Amoretti · 0 citations
Preprint Jul 2026

RubriQ: Rubric-Guided Group Relative Policy Optimization for Constraint-Aware Quantum Circuit Synthesis

RubriQ is introduced, a scalable framework that formulates circuit synthesis as a large language model (LLM) code-generation task, optimized via group relative policy optimization (GRPO), which establishes an automated, high-performance computing (HPC)-driven pipeline for generating hardware-ready, fault-tolerant quantum circuits at scale.

Ziqing Guo, Ziwen Pan · 0 citations
Preprint Aug 2026

Quantum circuit optimization using deep reinforcement learning: Applications across multiple gate sets

A reinforcement learning framework that embeds a deterministic Commutation-and-Reduction (CR) algorithm directly into the training environment, enabling the agent to focus its learning capacity on the non-trivial optimizations where reinforcement learning adds real value.

Khoa Dang Tao, Sumin Jin, M. Raza et al. · 0 citations