The latest paired run shows two things at once. Readout plus TFLO produces a much clearer spin profile than the raw data. However, in this test the compact circuit version does not agree more closely with our classical chi64 reference on local observables than the original version.
This is job daf6foe42tqs73avhoeg on ibm_phoenix, with registered QPU usage of 7 seconds for the entire job. Both circuit arms and their training, holdouts and calibrations are included. The model is 6×6, N=32, U=8, t’=-0.25 and T=0.4, with two symmetric steps.
First the values, then the differences
N is the total particle number. D is the average double occupancy per site. Spin RMS measures the magnitude of the local spin profile, not its error.
| Method | N | D per site | Spin RMS |
|---|---|---|---|
| Original, raw | 35.463 | 0.24065 | 0.01632 |
| Original, readout + TFLO | 31.396 | 0.09013 | 0.65683 |
| Compact, raw | 35.868 | 0.24854 | 0.01935 |
| Compact, readout + TFLO | 31.231 | 0.07719 | 0.67128 |
| Classical chi64 reference | 32.000 | 0.08790 | 0.70346 |
The contrast between raw and corrected spin is particularly striking. However, the correction uses information from U=0 training. We therefore need to test whether the final U=8 estimator is reliable.
The local RMS differences use all 36 sites and compare each observable separately against the same chi64 reference:
| Method | Charge | Spin | Doublons |
|---|---|---|---|
| Original, readout + TFLO | 0.08389 | 0.09185 | 0.05080 |
| Compact, readout + TFLO | 0.10930 | 0.12586 | 0.06412 |
These numbers are not known absolute errors: chi64 is not yet converged. They do show what the current comparison actually yields. Looking only at average D would hide local differences that can cancel in the sum.
Why we do not simply replace N with 32
The ideal Hamiltonian conserves particle number. We therefore know that N=32 is the model value, and may state it as a constraint. But replacing the measured N cell with 32 and then claiming that the hardware reproduced charge almost perfectly would be misleading.
If we genuinely want to incorporate a constraint into an estimator, the method must be specified in advance, and the local values and their uncertainties must change consistently with it. Repairing one number afterwards is not a validated mitigation method.
How far can we trust the table?
The independent U=0 holdouts did not meet our screening requirements. Some reconstructed local probabilities are negative. We preserve those outcomes without clipping. The bootstrap intervals describe only the chosen statistical assumptions; they do not include every source of drift, transfer bias or classical truncation error.
This is therefore a concrete, reproducible hardware milestone and an informative mitigation result. The next scientific step is not a more attractive table, but an estimator that also passes independent checks.
Source: analysis/summary.json and CONCLUSIE.md under results_cuprate_2d/hardware_compact_tflo_2026-09-07_v1, together with the saved classical_full_6x6_mps_T04_chi64_v1 reference.


