Skip to content
Review

Machine learning for sample-based quantum diagonalization: generative configuration recovery and the classical-simulability frontier

Aug 2026 · 2 citations · 88 references
Physics

TL;DR

This work confronts the field's central question -- whether the quantum sampler beats classical selected configuration interaction -- and reports a carefully scoped negative, and flags learning from quantum experiments, whose classical sample-complexity lower bound is an unconditional theorem, as the one adjacent frontier where a quantum advantage is provable but not yet bridged to chemistry.

Abstract

Sample-based quantum diagonalization (SQD), equivalently quantum-selected configuration interaction (QSCI), has in two years become a pragmatic centre of gravity of pre-fault-tolerant quantum chemistry: a quantum processor samples electronic configurations, and the many-electron Hamiltonian is diagonalized classically in the resulting determinant subspace. Its accuracy is set entirely by which configurations enter that subspace, a selection problem for machine learning made acute by a coupon-collector bottleneck. We critically review the ecosystem of generative and learned selectors, organizing it by the object each method generates and the importance signal it exploits, and expose one conspicuous gap: a reward-proportional generative-flow-network proposer built for tail discovery. We then confront the field's central question -- whether the quantum sampler beats classical selected configuration interaction -- and report a carefully scoped negative: across published same-active-space comparisons, strong classical selected CI matches or beats the quantum-sampled subspace, and the flagship single-layer circuits now admit polynomial-time classical energy estimation. We distil a benchmarking standard and turn the negative into a regime map, then test it with FCI-exact experiments that confirm one prediction and refute another: the cheap prior's rank correlation with the exact weights declines with multireference character (a usable coordinate), but a controlled single-molecule noise sweep shows the one generative advantage we find, robustness to valid-shot starvation, to be generic rather than the multireference-specific effect a confounded contrast first suggested. Finally, we flag learning from quantum experiments, whose classical sample-complexity lower bound is an unconditional theorem, as the one adjacent frontier where a quantum advantage is provable but not yet bridged to chemistry.

View source

Similar papers

Preprint Sep 2026

Quantum Wavefunction Augmentation via Variational Autoencoders

Sample-based quantum diagonalization (SQD) has emerged as a promising route for quantum-centric supercomputing, relying on classical diagonalization of the molecular Hamiltonian within a hardware-sampled determinant subspace. However, its accuracy degrades in strongly correlated regimes where the relevant determinant space exceeds what finite-shot sampling can capture. In this work, we introduce Quantum Wavefunction Augmentation via Variational Autoencoders (Q-WAVE), a hybrid method that combines determinants sampled via SqDRIFT Krylov circuits and configuration interaction singles and doubles (CISD) determinants with generative machine learning. Using a custom $\beta$-annealed variational autoencoder (VAE) model, Q-WAVE iteratively expands this basis toward the variational ground state. The VAE learns the wavefunction's primary support structure from the combined hardware and CISD seed in a continuous latent space, generating new dominant determinants beyond any fixed excitation hierarchy. The resulting compact wavefunction exceeds what can be extracted from raw hardware samples alone. We demonstrate sub-millihartree accuracy compared to full configuration interaction for $\text{H}_2\text{O}$ and $\text{N}_2$ dissociation. Finally, we establish Q-WAVE's scalability on a 52-qubit ethylene system (achieving sub-millihartree accuracy versus CCSD(T)) and a highly correlated 60-qubit $\text{Cr}_2$ stress test that attains chemical accuracy upon a final perturbative correction.

Sonaldeep Halder, Chayan Patra, Rahul Maitra · 0 citations
Open access Jul 2026

Accelerated quantum-centric supercomputing through perturbation-theoretic measures and generative machine learning

PIGen-SQD is introduced, an efficiently designed QCSC workflow that utilizes the capability of generative machine learning (ML) along with physics-informed configuration screening via implicit low-rank tensor decompositions for accurate fermionic state reconstruction.

Chayan Patra, D. Mondal, Sonaldeep Halder et al. · 1 citation · ⚡1
Preprint Sep 2026

DF-SQD: Deterministic Fields for Sampling-Based Quantum Diagonalization

Sampling-based quantum diagonalization method exploits Quantum-centric supercomputing platforms to sample bitstrings for Hamiltonian projection on a quantum computer, and then classically diagonalize the Hamiltonian to estimate the eigenvalues and eigenvectors. In current quantum devices an algorithm is useful when shallow quantum circuits with error mitigation support can discover better results while having either a proof of convergence or some method to explain trust in experiment. In this paper, we introduce DF-SQD, a hybrid algorithm that derives deterministic auxiliary-field circuits from selected double-factorization leaves of the two-electron tensor. The circuits propose occupation-number configurations, while selected configuration interaction evaluates the original active-space Hamiltonian and can recentre subsequent proposal rounds. On N2 (32 qubits; 6-31G basis) and a 40-qubit [Fe2S2(SCH3)4]2- active-space Hamiltonian, we show that DF-SQD improves the energy obtained from sampled determinant spaces while using shallow number-preserving circuits in both simulator and hardware runs. For N2, DF-SQD is 45x more accurate with a 11.23\% smaller subspace, and due to its ability to sample better bitstrings at lesser shots it is 2.93x faster than SQD in quantum devices. For the iron-sulfur cluster, DF-SQD generated a subspace dimension of 221M with 400K shots, while SQD needed 1.5M shots to generate a 238M subspace, thus we have better subspace recovery evident from the hardware at 3.75x reduced shots. At a matched 50M subspace dimension, DF-SQD is 1.32x more accurate (achieves a 24.5\% relative error reduction over standard SQD). So overall, our method is able to discover better results with shallower circuits, is sample efficient, uses configuration recovery (so has targeted error mitigation) and we have empirical convergence observation.

Kushagra Agarwal, Anupama Ray · 0 citations
Preprint Aug 2026

Dynamical spectral functions from bitstring-sampled quantum subspaces: entanglement, not one-body magic, tracks the sampling cost

Sample-based quantum diagonalization (SQD) and quantum-selected configuration interaction (QSCI) are the electronic-structure methods with most hardware traction, yet their canonical target -- the ground-state energy -- is where classical methods have caught up. We move the target to dynamics and the resource question. From one bitstring-sampling primitive -- computational-basis measurements of a shallow real-time circuit, with no Hadamard or controlled unitaries -- we reconstruct, from sampled subspaces, the single-particle spectral functions $A(\omega)$ and $A(k,\omega)$ and the neutral-sector dynamical structure factors $S(q,\omega)$ and $S^{zz}(q,\omega)$, each built classically in the Lehmann representation from its own subspace. The reconstruction matches exact diagonalization on Hubbard chains and, for $A(\omega)$, across nineteen molecules (FCI-verified to $<10^{-5}$ Ha), and runs on the IBM Heron processor. Second, we ask which resource controls the cost -- the determinant support $|\mathcal{S}|$ the sampler must populate. On number-conserving states the fermionic AntiFlatness collapses to one 1-RDM invariant, $\mathcal{F}_1 = 4\,\mathrm{tr}[\gamma(1-\gamma)] = 2N_u$. An orbital-rotation (Gaussian) invariant while $|\mathcal{S}|$ is basis dependent, $\mathcal{F}_1$ is provably decoupled from the cost; the cost is instead lower-bounded and tracked by the entanglement -- the minimal bond dimension $\chi$ (Spearman $\rho = 0.90$). One-body magic is thus a faithful multireference diagnostic but an unreliable cost predictor; any genuine advantage lives in the non-Gaussianity of the higher-body cumulants. We prove moment exactness and a sampling bound polynomial in $|\mathcal{S}|$, independent of Hilbert-space dimension. Self-consistent configuration recovery improves the subspace under device noise, while a learned generative model does not beat that classical baseline.

Nicolás Bonilla Vargas · 1 citation
Preprint Jul 2026

Generative AI Beyond Tokens: Quantum Resource Consumption of IQP Circuits

As IQP circuits produce remarkably low intermediate magic relative to phase-randomised states with the same sampling distributions, this renders IQP-based quantum generative models as promising candidates for resource-efficient demonstrations of quantum advantage on early fault-tolerant architectures.

Tom Krüger, Wolfgang Mauerer · 0 citations
Preprint Sep 2026

Quantum Representation Learning Beyond Pairwise Fidelity

Quantum contrastive, metric, and self-supervised learning often expose encoded quantum states to the learner through transition probabilities, especially fidelity. Quantum states are known to possess higher-order relational invariants, but their consequences for learned representations remain unclear. Here we show that a transition-probability-only learning interface can possess exact continuous blind directions in certain quantum-state families. We recover this missing information with a batch operator, built from coherent overlap amplitudes and negative masking, where its second moment $q_-$ retains four-state interference. Moreover, $q_-$ is directly measurable through two-copy interference and can enter variational learning via methods like parameter shift. In relational quartets derived from toric-code and double-semion states, this fidelity-blind signal encodes inequivalent modular data despite identical pairwise fidelities. Finally, in a four-photon benchmark with preparation drift, augmenting all six pairwise fidelities at two orthogonal probes with the corresponding normalized $q_-$ reduces the mean out-of-distribution phase error by $86\%$ at equal total shot budget. These results establish multistate relational observables as measurable, trainable, and physically consequential signals for quantum representation learning beyond pairwise fidelity.

Junpeng Hou, Chang-Bin Lu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.