Multi-Quantum OES: offline OES replay triage workbench with an isolated AI ∩ quantum toy lab
Abstract
Multi-Quantum OES is an offline, standard-library-only Python (3.11–3.13) workbench for explainable OES-512 telemetry replay triage, with a preregistered synthetic stress evaluation and a separate, isolated “AI ∩ quantum” toy lab. Everything runs locally on synthetic data; the quantum parts are exact classical simulations, not QPU runs. What is included OES-512 replay triage. Each frame has 512 finite values in sixteen 32-value blocks. A fixed block score (0.45·max|x| + 0.35·RMS + 0.20·mean|x|) flags a block when it is ≥ a preregistered threshold. It is reported next to a per-block max-abs baseline, with event recall, false-positive counts and episodes, block precision/recall/IoU, latency, ranked evidence, and SHA-256 hashes of the exact replay and protocol bytes. Thresholds come from a preregistration file and are never tuned by the CLI. Preregistered stress evaluation. Seeded generators produce a held-out test split and a separate calibration split (320 frames, 6 regimes, 11 labelled events each). Two protocols were locked before evaluation, a shared fixed threshold of 0.50 and thresholds calibrated on the calibration split to a 5% no-event false-positive rate, and a hash manifest records the data, preregistrations and tools. OES32 adaptive-tau stream model (float-level, models the earlier v1 triage branch logic; not bit-exact HLS), a 32-value residual check (fails only when R > tolerance), offline outage-snapshot classification, metadata-only failure-capsule validation, and a loopback-only local web interface. AI ∩ quantum toy lab. One exact statevector core (gates, Pauli-sum observables, parameter-shift gradients, full-unitary equivalence) powers four small applications: quantum-kernel classification versus a classical RBF kernel, QAOA MaxCut versus brute force, circuit compilation with unitary verification, and H2 VQE versus exact diagonalisation (H2 coefficients from the Qiskit Algorithms tutorial, Apache-2.0). It never enters the telemetry scoring path. Reproducibility. 72 unit tests; CI on Python 3.11, 3.12 and 3.13 checks that every committed report regenerates byte for byte. Results, including negative ones Stress evaluation: negative. At the shared 0.50 threshold the block score detected fewer events than the max-abs baseline (5/11 vs 8/11) with fewer false-positive frames (45/272 vs 81/272) but more false-positive episodes (26 vs 14). Calibrated to the same 5% false-positive rate, both detectors gave identical results (0/11 events, 17/272 false-positive frames). With only 11 invented events, none of these differences is statistically meaningful. Quantum kernel: at chance on the toy data (accuracy 0.531) versus 0.906 for a classical RBF kernel with an untuned γ. QAOA p=1 / p=2 reach approximation ratios 0.822 / 0.925 on a 5-node MaxCut; the compiler reduces a 15-gate circuit to 3 gates and catches a deliberately faulty compile; H2 VQE matches exact diagonalisation (|Δ| about 4e-16 Ha). Limits Synthetic data only; no real telemetry was used or inferred. Exact classical simulation of 2–5 qubit toy problems: not QPU execution, not evidence of quantum advantage, and not a chemistry workflow. Not an operational, alarm, safety or medical system. The OES32 stream model does not reproduce fixed-point rounding, saturation, AXI timing or synthesis. The stress results describe these generators and this calibration rule only. Version 0.1.1 changes metadata and documentation only; code and outputs are as in 0.1.0.