Part 7 of the series From quarks to quantum advantage
After six parts we can finally ask the question that is often asked in the first paragraph: has the quantum processor won over the classical computer? The honest answer is not one word. The hardware time is much shorter than the published TN and Pauli propagation times for the chosen observable-estimation instance. But shorter only becomes quantum advantage when task, accuracy and time accounting are sufficiently equal.
The paper provides strong evidence for a practically interesting hardware route and a coherent non-Abelian hadron signal in the early time window. It does not yet provide a general complexity separation, a complete end-to-end benchmark, and no proof that each late time point accurately corresponds to one known classical truth.
Four questions, not one
A benefit claim must answer four separate questions:
- Physics: Does the approximate dynamics follow the full LSH Hamiltonian?
- Circuit: Does the ideal circuit calculate the intended observable?
- Hardware: Does the signal remain recognizable after gates, decoherence and measurement?
- Performance: Is the quantum route for the same task faster or more scalable?
flowchart LR
H["full LSH theory"] -->|"TN versus PP"| C["ideal circuit"]
C -->|"PP versus QPU"| Q["hardware data"]
Q -->|"decoder and normalization"| O["physical observable"]
O -->|"equal error and accounting"| A["advantage claim"]
A quick hardware measurement without the first three steps can efficiently produce the wrong answer. In contrast, a perfect small validation without performance comparison shows no benefit. The strength of this project lies in the chain, but each link has its own reach.
What does the paper compare physically?
The tensor network route evolves the full LSH Hamiltonian with a flux cutoff and \(D_{\max}=200\). Pauli propagation evolves the measured operators through the ideal, approximate Trotter circuit. The QPU runs that circuit with hardware errors.
This results in two diagnostic differences:
TN - PP -> model approximation, cutoff difference, and Trotterization
PP - QPU -> hardware, compilation, and measurement errors
The similarity in the early spacetime heatmaps and in the derived breathing frequency supports the interpretation of the hardware pattern as hadron dynamics. That's stronger than just a conserved global charge.
The comparison becomes more difficult later in evolution. The TN bond dimension becomes limiting, PP has to handle more Pauli terms, and the deep QPU circuits accumulate more hardware error. The methods may then all deteriorate, but not necessarily in the same direction.
The same observable is a hard condition
The global charge
\[Q=\sum_r[n_i(r)+n_o(r)]
\]
has an exactly intended value of 60. It is excellent for sector control. At step 20, the tracker reports \(Q_{\mathrm{QPU}}=59,99402699\), a deviation of approximately 0.00597. The included charge-based error bar is 0.00571065; which is close, but not identical, to the raw absolute deviation.
However, the hadron question is in \(n_f(r,t)\), the difference heatmap and the breathing frequency. A correct (Q) does not rule out large local profile deviations. This is also evident from the tracker values for the global \(n_f(20)\) scalar:
| method | n_f(20) |
|---|---|
| TN | 0.277835 |
| PP CPU | 0.408178 |
| PP GPU | 0.245628 |
| QPU | -0.077234 |
The method standard deviation there is approximately 0.2062 and the QPU is 0.32 to 0.49 below the three classical values. So step 20 is not a point where all methods coincide numerically. The paper claim relies more broadly on the early spacetime pattern, symmetry conservation and the spectroscopic frequency.
This also sets the limit for our reproduction. Our local SCV versus meson scalar is derived from the same QASM counts, but does not yet demonstrably use the same summation and normalization as tracker-\(n_f(t)\). The two columns should therefore not be subtracted from each other as error. First, the paper definition must be recorded exactly in the analyzer.
The paper-native timing
The tracker reports times per evaluated Trotter depth for the main instance. For steps 5 and 20, the published values are:
| step | QPU | TN | PP CPU | PP GPU | fastest classic / QPU |
|---|---|---|---|---|---|
| 5 | 20.0 s | 584.092s | 477.447s | 547.581 s | 23.9x |
| 20 | 20.0 s | 2,484.369 s | 2,930.440 s | 8,949.110s | 124.2x |
At step 5, PP-CPU is the fastest of the three classic routes; at step 20, TN is the fastest. Within the accounting of the paper, the QPU remains at 20 seconds due to the fixed shot budget, while the classical methods become more expensive with the evolution depth and the required representation size.
This table supports a clear practical timing gain within the reported benchmark protocol. It does not prove that every classical algorithm family must be at least as slow on every machine. They are measured or reported implementations with concrete cutoffs and resource budgets.
Our Fire Opal timing
Our local hardware-only time uses estimated_duration * shots for the two main circuits, SCV and meson. With 512 shots we found 0.987136 seconds for step 5 and 2.768896 seconds for step 20.
To just equalize the shot budgets, we can scale linearly to 10,000 shots:
| step | measured locally, 512 shots | estimated at 10,000 | paper-QPU | fastest paper-classic/local estimate |
|---|---|---|---|---|
| 5 | 0.987136 s | 19.28s | 20.0 s | 24.8x |
| 20 | 2.768896 s | 54.08s | 20.0 s | 45.9x |
At step 5, the linear estimate is very similar to the paper QPU time. At step 20, our estimated ibm_fez hardware time is about 2.7 times the paper value of 20 seconds. The quantum route remains well faster than the published classical baselines in hardware time: approximately 45.9 times compared to TN, 54.2 times compared to PP-CPU and 165.5 times compared to PP-GPU.
The 54.08 seconds have not been measured. They assume linear shot scaling and the same optimized circuit duration. A real 10,000-shot run can be different due to service behavior, calibration or batching.
Three clocks give three answers
At least three time definitions are relevant for our local runs:
| clock | contains | excludes |
|---|---|---|
| hardware only | estimated QPU execution of the shots | queue, API, analysis |
| action wall | Fire Opal action from submit to result | local preparation and post-processing |
| end to end | preparation, service, hardware, retrieval and decoding | nothing within the chosen workflow |
The main action for step 5 had a wall time of 133 seconds. The combined main action for steps 10, 15 and 20 lasted 304 seconds. The hardware-only parts within are much shorter. That's not contradictory: a cloud service can wait, compile, schedule, and process multiple circuits as a group.
The paper QPU time and our local hardware time are appropriate for how the physical shot execution scales. They do not answer how long a researcher waits for a final result from a source file. An end-to-end benefit has therefore not been established with the current tables.
There is a second accounting question. The QPU uses a separate circuit for each time point. A classic TDVP run can save multiple time points from the same evolution along the way. The paper reports its own compute up to step protocol, but an independent whole-path benchmark should explicitly capture whether all time points or just the last point are queried.
Accuracy and time go together
A method that becomes faster by truncation more aggressively only has an advantage if the error is still within the agreed budget. Three types of errors are visible for this comparison:
- TN: flux cutoff, finite bond dimension and loss with growing entanglement;
- PP: truncation or budgeting of a growing Pauli expansion;
- QPU: gate, decoherence, readout and shot noise.
The paper uses global charge drift and method spread as practical diagnostics because the exact \(N=60\) real-time solution is not available. However, method spread is not a statistical confidence interval around a known truth. All methods may have a shared bias or may diverge due to different approaches.
A final benchmark would choose one error measure in advance, for example RMSE of the entire \(\Delta n_f(r,t)\) heatmap in a time window, plus a tolerance for (Q) and the breathing frequency. Then all routes must be converged to that same tolerance and only then timed.
Four possible meanings of quantum advantage
We can now formulate the conclusion precisely:
| Claim | Current status |
|---|---|
| Asymptotic complexity advantage | not demonstrated |
| Universal classical intractability | not demonstrated |
| End-to-end runtime advantage | not yet measured |
| Hardware time savings versus reported TN/PP baselines | clearly present |
| Coherent physical signal on a large non-Abelian instance | supported in the early time window |
The strongest defensible formulation is therefore a practical hardware time saving for the defined observable-estimation task, supported by a layered but approximate physical validation. That is scientifically interesting without making it bigger than the data allows.
What else is needed for our own strong claim?
Local reproduction has four concrete next steps:
- Implement exactly the paper normalization of \(n_f(t)\) and test it on synthetic and ideal QASM data.
- Preferably perform step 20 with 10,000 shots instead of just scaling linearly.
- Report bootstrap or binomial uncertainties for SCV, meson and their differential, in addition to the (Q) sector check.
- Measure hardware-only, action wall and full end-to-end time with the same start and end points for quantum and classical.
For an entire route claim, one more choice must be made: does the task only count \(t=20\), or all twenty intermediate points? That workload must be identical for all methods.
Final conclusion of the series
The path from quark physics to a timing table involves much more than a large QASM file. The non-Abelian Gauss laws are first solved in the LSH basis. A weak-coupling approximation reduces the local dynamics to two qubits per site. Trotterization turns this into a shallow local circuit. Differential SCV versus meson measurement extracts a coherent signal from noisy hardware. TN and PP then control various transitions in that chain.
Within that carefully defined protocol, the QPU hardware is much faster than the reported classical baselines and physically recognizable early hadron dynamics remain visible. The current data do not justify a blanket statement that classical simulation has been defeated. They do show that a physics-native encoding takes a non-Abelian real-time computation to a scale where the available classical references quickly become expensive and difficult to verify.
That is not a slogan, but a concrete research result with a clear next test: the same paper-normalized observable, the same error limit and the same clock on all platforms.


