Our goal does not start with beating every supercomputer. It starts with a much more personal question: can we, as a student/hobby project, have a quantum computer perform an interesting calculation using less recorded computation time than our own classical computer needs?
That limited goal is the common thread of this third Hubbard series. In the first series we explored the one-dimensional model. In the second we moved to two dimensions. We now focus on IBM Nighthawk, more physically faithful fermion circuits and a measurable local timing comparison.
The first milestone: approximately twenty times
For our archived 6×6 instance, the classical chi64 calculation took 150.819 seconds in its kernel. A Nighthawk job recorded 8 seconds of QPU usage; a later paired original/compact job recorded 7 seconds. The ratios are approximately 18.85 and 21.55. We summarise this as approximately 20x.
That ratio genuinely follows from the saved timing records. But it answers a specific question. QPU usage is not the same as the time between clicking ‘start’ and reading a validated answer. Queueing, preparation, network communication and local analysis are not all included on that QPU clock. Moreover, our classical reference has not been shown to be converged: chi64 is a calculation setting, not a guarantee of correctness.
Our claim is therefore that the quantum job used approximately twenty times less registered QPU time than the kernel of our current local classical baseline. We do not yet claim a twentyfold reduction in total elapsed time at demonstrably matched accuracy.
What does the quantum computer actually calculate?
We simulate a simplified model of moving, mutually repelling electrons: the two-dimensional Fermi-Hubbard model. The lattice has 36 sites. With two spin modes per site, that gives 72 fermion modes represented on 72 qubits. The state contains 32 particles, not 72.
We do not read out the complete quantum state. We ask for local charge, spin and double occupancy. These observables describe where the particles are and how the chosen initial state evolves. A useful quantum run must be judged on those requested quantities.
Working within a limited budget
The hardware experiments grew from small pilots to 6×6. We worked with 512 shots per setting and explicit QPU time limits, eventually a maximum of 15 seconds for an authorised job. More settings, extra mitigation circuits and additional shots all consume budget. Maximum depth, many time points and perfect statistics cannot be combined for free.
That makes this project educational. We had to decide which check would provide the most information, when a circuit becomes too deep, and when an attractive corrected value mainly tells us something about the correction method.
The new repository therefore contains more than appealing final tables. Source code, raw data, unsuccessful pilots and theory are preserved as well. While the repository is private, access is restricted to authorised readers.
Sources: our job records under hardware_tflo_dd6_2026-09-06_v1 and hardware_compact_tflo_2026-09-07_v1, the chi64 baseline, and IBM’s explanation of workload usage.


