The next step is not automatically a larger lattice. Our project already has an interesting local timing ratio and a complete 6×6 implementation. The main open question is how much reliable U=8 information the quantum measurement genuinely adds after mitigation.
Test the estimator first
We choose in advance which observables to estimate and what differences we will accept. Charge, spin and doublons each receive their own tolerance. Otherwise, averaging across all types of observable could hide a poor spin estimate behind an easier charge check.
We then test the mitigation on settings whose answers we know, but which were not used for training. Our latest U=0 holdouts show that more work is needed. If the method does not generalise well even there, an attractive U=8 target value is insufficient evidence.
The control without target hardware input remains important. When an estimator gives approximately the same answer without using the interacting hardware data, that answer does not convincingly measure the intended quantum dynamics. We specifically want to demonstrate that those data contribute useful, testable information.
Two refinement ladders, not one
On the classical side, we refine the bond dimension. On the quantum and circuit side, we must distinguish noise, repetitions and time-step error. More chi does not repair the wrong circuit; more shots do not repair an incorrect fermion mapping; better readout does not fix every gate error.
A useful design holds T fixed, compares several time-step sizes and always uses the correct reference for the selected finite circuit. Only then can the deviation from continuous Hamiltonian dynamics be assessed. Under a cost limit, it is better to test one clear question than to change many factors at once.
A stronger timing test
The approximately 20x ratio in this series uses registered QPU time and the existing local classical kernel time. A stronger test also records preparation, training, calibration, waiting and analysis. We explicitly state whether a calibration is performed once or reused across several calculations.
The finish line is then not ‘the job is complete’, but ‘the requested observables meet the agreed error tolerance’. This allows a genuine time-to-acceptable-answer comparison against our local computer. That remains a limited, achievable goal; it does not require beating every classical computer in the world.
A record that others can inspect
The separate Nighthawk repository preserves code, theory, raw measurements, calibrations, references, checksums and the results of less successful approaches. Large files are compressed losslessly and have a restoration manifest. The original project directory remains in place.
This series is therefore not a final declaration, but a well-documented intermediate result. We can be proud of what has been achieved within a limited budget while remaining precise about what has not yet been proved. The next hardware job will be submitted only after fresh, explicit budget approval.
Sources: theory/RESULTS_AND_CLAIMS.md, HARDWARE_COMPACT_TFLO_PROTOCOL.md, the holdout tables and the saved chi32/chi64 checks in the project repository.


