Edukaizen benchmark register
Pro Student Quantum Advantage List
8 student-scale project reports, with separate labels for local time-to-answer comparisons, runtime bounds, execution-metric ratios and diagnostic results. Timing scope, numerical accuracy and public access are stated for each entry.
The current list
| Project | Scale | Primary quantum timing | Classification |
|---|---|---|---|
| 1D Fermi-Hubbard | 120 qubits / 60 sites | 33.148928 s | Local time-to-answer separation |
| SU(2) hadron dynamics | 120 qubits / 60 sites | 1.425408 s | Paper-aligned local separation |
| Operator Loschmidt Echo Q80 | 80 qubits | 328 s | Local runtime lower bound |
| Random Graph Sampling | 70 qubits | 19 s | Diagnostic only |
| Floquet-Ising 51q | 51 qubits | 41 s | Local time-to-answer separation |
| 2D Hubbard Nighthawk | 72 qubits / 36 sites | 7 s | Local execution-metric ratio; accuracy unvalidated |
| XXZham | 80 spins | 76 s QPU usage | Local execution-metric lead; chi=512 passes one full refinement, convergence unfinished |
| Nighthawk RCS | 61 qubits | 19 s QPU usage | 65.7x local execution-metric lead over circuit-following MPS; process fidelity unvalidated |
Entry 1 · Local time-to-answer separation
Fermi-Hubbard dynamics on 120 qubits
A 60-site Fermi-Hubbard hardware workflow produced local charge, spin, and double-occupancy observables and was compared with local MPS and observable-specific Majorana calculations.
Measured comparison. The quantum execution proxy was 272.50x shorter than the local chi=256 MPS wall time for this declared instance.
Quantum result. The hardware produced a full 120-qubit observable profile; raw mean double occupancy was 0.22862549 and readout-corrected mean double occupancy was 0.23067722.
Classical baselines
| Method | Wall time | Status |
|---|---|---|
| Local quimb MPS with maximum bond dimension chi=256 | 9,033 s | not fully converged; maximum bond reached the requested cap |
| Local Majorana propagation with cutoff 2 | 20.96 s | faster than the quantum proxy but visibly inaccurate |
| Local Majorana propagation with cutoff 4 | 1,153.51 s | close to the chi=256 value for this selected observable |
Official sources
Complete implementation
Claim boundary
- This is a local time-to-answer result, not a reproduction of the paper's headline practical-advantage claim.
- The chi=256 MPS baseline did not establish full convergence.
- The fast Majorana route demonstrates that observable-specific classical methods can change the ranking.
Entry 2 · Paper-aligned local separation
Non-Abelian SU(2) hadron dynamics on 120 active qubits
A Loop-String-Hadron implementation follows a differential hadron signal on a 60-site lattice and compares the quantum route with local circuit-MPS checks and published tensor-network and Pauli-propagation baselines.
Measured comparison. Both the local circuit checks and the paper-native baselines show a substantial runtime separation under their declared timing definitions.
Quantum result. The local hardware route produced charge-sector and differential-observable data for the 120-qubit circuit family.
Classical baselines
| Method | Wall time | Status |
|---|---|---|
| Local Aer MPS on compiled QASM | 34.281305 s | completed local sanity baseline |
| Local ITensorMPS on compiled QASM | 174.611 s | completed local sanity baseline |
| Published Pauli propagation on CPU at step 5 | 477.4471 s | published baseline |
| Published Pauli propagation on GPU at step 5 | 547.581 s | published baseline |
| Published ITensor TDVP tensor network at step 5 | 584.092 s | published baseline |
Official sources
Complete implementation
Claim boundary
- Hardware-only time is not cloud wall time and excludes several service overheads.
- The local scalar normalization remains distinct from the tracker's published hadron scalar.
- The result supports runtime separation and circuit or sector validation, not an independent precision reproduction of every published observable.
Entry 3 · Local runtime lower bound
Operator Loschmidt Echo on 80 qubits
A tracker-compatible 80-qubit extension estimates an Operator Loschmidt Echo from finite computational-basis samples and compares the complete mitigated hardware action with a bounded tracker-linked BP-TN calculation.
Measured comparison. The incomplete bond-dimension-64 classical delta half alone exceeded the complete Fire Opal action by more than 2.75x on this machine.
Quantum result. The measured delta/delta0 OLE ratio was 0.74028847 +/- 0.01663657; all eight sample ratios were positive.
Classical baselines
| Method | Wall time | Status |
|---|---|---|
| Tracker-linked Heisenberg BP-TN at bond dimension 16 | 365.14 s | not converged; apparent ratio is not a valid physical estimate |
| Tracker-linked Heisenberg BP-TN delta half at bond dimension 32 | 342.42 s | not converged; value shifted by 86 percent from bond dimension 16 |
| Tracker-linked Heisenberg BP-TN delta half at bond dimension 64 | 901.01 s | timeout before producing a result |
Official sources
Complete implementation
Claim boundary
- The classical calculation did not converge and no matched-accuracy ratio was obtained.
- This is a tracker-compatible 80-qubit extension with N_init=8, not an official tracker instance or an N_init=500 reproduction.
- The observation is local and does not cover every classical implementation or optimized compute platform.
Entry 4 · Diagnostic only
Random Graph Sampling on 70 data qubits
A complete 70-data-qubit non-Clifford circuit was sampled on IBM hardware, alongside an independent 70+8-qubit stabilizer-verification workflow and local classical scaling studies.
Measured comparison. Hardware returned 256 samples in 19 quantum-seconds, while a local Aer fit projects about 6.89 million years for one 70-qubit sample; sample counts and output quality are not matched.
Quantum result. The complete circuit returned 256 samples. The separate checked dataset retained 4,519 of 184,320 shots and gave a graph-state-prefix point estimate of 0.01217; its predeclared one-sided 95 percent lower-bound test failed. The original restricted-access IBM Boston execution reported substantially stronger effective performance than this independently accessible Kingston reproduction.
Classical baselines
| Method | Wall time | Status |
|---|---|---|
| Local Qiskit Aer extended-stabilizer fit evaluated at 70 qubits | 217,512,854,796,362.625 s | extrapolated; not measured at 70 qubits and not quality matched |
| Local ITensorMPS at maximum bond dimension 64 | 205.36 s | completed but strongly truncated and not converged |
| Local exact MPS anchor at 14 induced qubits | 3.23 s | exact small-width validation; not a 70-qubit baseline |
Complete implementation
Claim boundary
- The Tracker result used restricted access to IBM Boston, whereas this independent reproduction used the available IBM Kingston route; backend access, physical mapping, and calibration window are therefore not matched.
- Boston produced a substantially stronger workload-level result, but its historical calibration and complete raw fidelity-analysis record are not public, so the result does not establish that Boston was universally better hardware than Kingston.
- The 70-qubit classical runtime is extrapolated from measurements ending at 12 qubits, not measured at full width.
- The quantum samples have no validated full-distribution fidelity, and the separate predeclared 95 percent stabilizer test failed.
- The post-hoc 75 percent lower bound is an exploratory sensitivity result, not 75 percent fidelity and not evidence of quantum advantage.
Entry 5 · Local time-to-answer separation
Floquet-Ising oscillation detection on 51 qubits
A complete 51-qubit IBM Fez PEA/ZNE trajectory detected a reproducible Floquet-Ising oscillation in less declared execution time than a local D=64 PEPS simple-update trajectory for the same coordination-two magnetization.
Measured comparison. The 41-QPU-second PEA/ZNE route detected the 51-qubit oscillation 8.93x sooner than the 366.17-second local D=64 PEPS-SU trajectory under the declared timing scopes.
Quantum result. The fitted PEA/ZNE period was 4.76 cycles with profile interval 4.62 to 4.92; the fitted oscillation amplitude was about 5.2 conditional fit standard errors, consistent with two independent raw plus M3 trajectories at periods 4.48 and 4.61.
Classical baselines
| Method | Wall time | Status |
|---|---|---|
| Local Quimb arbitrary-geometry PEPS simple update at D=64 | 366.169545 s | complete; D=40-to-D=64 convergence accepted through cycle 8 only, with cycles 9-16 diagnostic rather than converged |
| Local Quimb arbitrary-geometry PEPS simple update at D=40 | 118.33 s | complete; used with D=64 to define the low-D convergence window |
| Published PEPS-BP production calculations at D=512 and D=700 | not available | paper baseline; 35.1 hours at D=512 and 599.8 hours at D=700, with late-cycle convergence limitations |
Official sources
Complete implementation
Claim boundary
- This is a partial, task-specific practical time-to-signal advantage, not a general or complexity-theoretic quantum-advantage claim.
- The 41-second timing is QPU execution only and excludes queueing, orchestration, retrieval, analysis, and classical mitigation processing.
- The classical D=64 PEPS-SU trajectory is converged by the declared low-D diagnostic only through cycle 8; cycles 9-16 are not a claim-ready classical reference.
- Across the eight trusted cycles, PEA/ZNE has diagnostic SRMSE 3.154 and maximum absolute z-score 6.932, so the preregistered matched-accuracy thresholds are not met.
- The local PEPS-SU implementation is not the paper's PEPS-BP production method at D=512 and D=700, and a stronger observable-specific classical result may change the ranking.
Entry 6 · Local execution-metric ratio; accuracy unvalidated
2D Local Quantum Advantage: 6×6 Hubbard on Nighthawk
A student/hobby project ran full-fermion 6×6 Hubbard circuits on 72 modes and measured charge, spin and doublons. Its local milestone is circa 20x less registered QPU usage than the current laptop-side chi64 MPS kernel time. The timing ratio is measured; numerical accuracy and end-to-end advantage remain unvalidated.
Measured comparison. The local chi64 kernel took 150.819180 s versus 7 registered QPU s for the complete paired job: 21.55x, or circa 20x. The separate 8-QPU-second job gives 18.85x. This is an execution-metric ratio, not an end-to-end or matched-accuracy speedup.
Quantum result. Readout+TFLO original gave N=31.39633 and D/site=0.090130; compact gave N=31.23064 and D/site=0.077194. Sitewise charge/spin/doublon RMS differences from the uncertified chi64 reference were 0.083889/0.091851/0.050805 (original) and 0.109304/0.125864/0.064116 (compact). Both TFLO holdout checks failed; reconstructed local probabilities reached -0.077303 and -0.092726. These target estimates are unvalidated.
Classical baselines
| Method | Wall time | Status |
|---|---|---|
| Local Quimb finite-circuit MPS, chi=64, one thread | 150.81918 s | completed but not converged or certified; reference N=31.9999992 and D/site=0.087896; chi128 was not run |
| Same chi64 calculation, cold worker | 156.905497 s | completed; broader timing of the same calculation, not a second independent reference |
Complete implementation
Access. Nine detailed Edukaizen articles and this register entry are public. The complete source, raw runs and theory archive remain in a private GitHub repository; access requires permission. Admission uses the public project-report route, not a claim that the complete implementation is publicly downloadable.
Claim boundary
- The approximately 20x milestone compares local classical kernel time with registered QPU usage. It is not a 20x shorter end-to-end run or a matched-accuracy quantum advantage.
- The 7 seconds cover the entire paired job, not each arm. The earlier 8-second DD/TFLO result is a separate job; mitigation settings and budgets must not be mixed.
- The latest provider running-to-finished interval was about 181 seconds, already longer than the 150.819-second classical kernel; queueing, preparation and local analysis add other overheads.
- The chi64 reference has no certified 6×6 error bound and is not converged. Chi128 was not performed, and its estimated cost cannot be counted as extra measured quantum speedup. Concurrent local work also limits timing comparability.
- Both TFLO holdout checks failed and some reconstructed probabilities are negative. The reported charge, spin and doublon estimates remain unvalidated; apparent agreement of individual observables does not validate the estimator.
- Exact N=32 is a model constraint, not a replacement for measured N. Original and compact results retain their own measured values and reference differences.
- This local resource comparison does not establish superiority to all classical algorithms, Google/Bonsai results, or a superconducting-material simulation. The 72-mode model and two-step finite circuit define the tested task.
- The public reports describe the evidence, but the raw research archive remains private. Full independent reproduction therefore requires access permission; listing the project does not remove this limitation.
Entry 7 · Local execution-metric ratio; convergence unvalidated
XXZham: imbalance dynamics of 80 spins
An IBM Kingston Heron R2 run returned the full 31-point imbalance curve for an 80-spin XXZ chain in 76 registered QPU seconds. Local chi=64, chi=128, chi=256 and chi=512 TEBD runs at dt=1/24 took 130.69 s, 1,503.99 s, 4,571.54 s and 35,432.31 s. These yield descriptive QPU execution-time ratios of 1.72x, 19.79x, 60.15x and 466.21x. The chi=256 to 512 curve change passes both selected refinement limits; the full convergence protocol is unfinished.
Measured comparison. QPU usage of 76 s is shorter than local chi=512 TEBD execution (35,432.31 s) by 466.21x. The source-TN-mean RMSE improves from 0.00514 at chi=256 to 0.00454 at chi=512. The chi=256 to 512 curve change is 0.00267 RMS and 0.00676 at its largest point, within the chosen 0.005 and 0.01 limits. The preceding chi=128 to 256 maximum change failed its limit, and time-step convergence at chi=512 was not tested. This is a measured execution-time contrast, not a matched-quality quantum-advantage result.
Quantum result. The completed IBM job produced all 31 points with RMSE 0.036943 against the stated tensor-network mean. A later 40-second-capped quantum attempt failed and yielded no complete curve.
Local classical trials
| Method | Execution | Convergence status |
|---|---|---|
| Local one-thread TEBD, chi=16, dt=1/6 | 8.012 s median | Excluded: adjacent time-step refinement changes 0.02461 RMS and 0.06443 maximum |
| Local TEBD, chi=64, dt=1/24 | 130.69 s | Not converged: source-mean RMSE 0.00951; chi=32 to 64 change 0.01570 RMS / 0.04393 maximum |
| Local TEBD, chi=128, dt=1/24 | 1,503.99 s | Improved source-mean RMSE 0.00638; chi=64 to 128 change 0.00769 RMS / 0.02158 maximum still exceeds the stability limits |
| Local TEBD, chi=256, dt=1/24 | 4,571.54 s | Source-mean RMSE 0.00514; chi=128 to 256 change 0.00448 RMS passes, 0.01348 maximum fails |
| Local TEBD, chi=512, dt=1/24 | 35,432.31 s (9 h 50 min 32 s) | 31 points to t=5; source-mean RMSE 0.00454; chi=256 to 512 change 0.00267 RMS and 0.00676 maximum, both pass |
| Published H200 Pauli propagation | 1,673.49 s | Separate upstream Heron R3 benchmark; not a local laptop measurement |
Official sources
Complete implementation
Claim boundary
- The fast chi=16 timing is real: five complete runs reached all 31 times through t=5. It is excluded under the selected convergence gate because its coarse cap and time step fail refinement stability.
- Chi=128 improved the observable error and took 25 minutes 4 seconds on one local CPU thread. It still changed too much from chi=64, and time-step convergence at chi=128 was not measured. The chi=256 run took 76 minutes 12 seconds and had source-mean RMSE 0.00514; its chi=128 to 256 maximum change of 0.01348 failed the 0.01 limit. The completed chi=512 run took 9 hours 50 minutes 32 seconds, produced 31 finite points through t=5 and had source-mean RMSE 0.00454. Its chi=256 to 512 changes of 0.00267 RMS and 0.00676 maximum pass both limits. A second consecutive passing chi refinement and time-step refinement are still missing.
- The 76-second QPU usage and local CPU execution have different clock scopes. Queue, service overhead and local setup are not matched in this ratio.
- The reported 5.63x QPU/H200 ratio belongs to a different upstream Heron R3 benchmark, not the laptop comparison.
- The later frozen quantum trial had no successful complete result. These local measurements show an execution-metric lead over the named chi=64, chi=128, chi=256 and chi=512 TEBD settings, but do not yet establish matched-quality or end-to-end quantum advantage.
Entry 8 · Local execution-metric ratio; process fidelity unvalidated
Nighthawk random-circuit sampling on 61 qubits
The released 61-qubit, 36-cycle circuit returned one million bitstrings on IBM Phoenix in 19 registered QPU seconds. A local chi=128 MPS simulation of the same logical circuit took 1,248.95 seconds for 1,000 bitstrings. This is a 65.7x local execution-metric lead over that circuit-following simulator; process-level fidelity has not been independently matched.
Measured comparison. The 19-second QPU usage was 65.7x shorter than the 1,248.95-second local chi=128 MPS run of the released circuit. Its RMS difference from the IBM data across 61 one-bit probabilities was 0.09047, within an exploratory hobby tolerance of 0.10. That low-order check does not establish equal process fidelity. The shot counts and timing scopes also differ.
Quantum result. The IBM job returned one million 61-bit strings. It repeated the released circuit’s scale and sampling count, but this project has not independently determined its full-distribution XEB or fidelity. The original paper used mirror and patched-XEB estimators for that purpose.
Classical process simulations
| Method | Wall time | Status |
|---|---|---|
| Local Qiskit Aer MPS, chi=128 | 1,248.95 s | Same logical circuit, 1,000 samples; one-bit RMS 0.09047, inside the exploratory 0.10 band |
| Local Qiskit Aer MPS, chi=64 | 173.81 s | Same logical circuit, 1,000 samples; one-bit RMS 0.10923, outside the band |
| Local Qiskit Aer MPS, chi=8 | 1.45 s | Same logical circuit, 1,000 samples; one-bit RMS 0.19505, outside the band |
Complete implementation
Claim boundary
- A circuit-independent random-bit generator is not a competitor for simulating this quantum process. It serves only as a negative control showing that one-bit RMS alone cannot verify the process.
- The 0.10 one-bit RMS band is an exploratory, post-measurement hobby criterion; it is much looser than the approximately 0.018 noise-only 95% threshold for 1,000 shots.
- The 19 seconds are provider QPU usage for one million shots; the local MPS wall times cover 1,000 shots. Queue, preparation, service overhead and analysis are excluded from the quantum timing.
- Passing the marginal check does not validate the full 61-bit distribution, ideal-circuit fidelity or full-width XEB. Our run did not reproduce the paper’s mirror or patched-XEB fidelity estimates.
- The 65.7x figure is a timing lead over one named circuit-following capped-MPS implementation, not a demonstrated matched-fidelity, best-classical or end-to-end quantum advantage.
What this list does not claim
A stronger classical implementation is a successful challenge, not a problem. Every result is conditional on its stated observable or task, accuracy or convergence status, timing scope, and available resources. The list does not certify formal complexity-theoretic advantage.


