Two experiments, kept explicitly apart. One tests whether the tyre generalizes to circuits it never saw; the other calibrates the offset the live predictor uses. The same circuit can score two opposite numbers, and that is the most honest thing on this page.
Fit two tyre numbers to one real lap, freeze them, then score the full-lap sim on the real driven lines of circuits the model never saw. Anything inside the shaded band counts as generalizing. The two that fall outside it are on the chart as well.
This experiment records the v1 regime, superseded at the summer re-freeze. At the time there was no real line before a session, so the live predictor ran on generated minimum-curvature lines, and the offset was calibrated in that same regime. Calls from Zandvoort onward use the venue’s own 2025 pole lap instead. But these four circuits are the calibration set, which makes this an in-sample fit: μ is their mean and σ their sample scatter, not an out-of-fold estimate. With n that small the number needs an asterisk, and μ leans hard on a single circuit.
μ corrects the lap; σ sets the 1σ/2σ bands. But the four circuits are the calibration set, so σ is an in-sample scatter. It understates the true predictive spread, and μ leans on one circuit: drop Shanghai and it flips to +0.14%. With the real band is wider than 1σ implies. The two scored calls showed this in different ways. Spa at 2.04σ was still consistent with an understated σ. Budapest at 4.19σ was not, which makes it a bias rather than scatter, and it is what triggered the documented re-freeze at the summer break. The offset math →
Budapest's debrief pre-declared a re-freeze at the summer break. It happened, under four declarations hashed before the runs they govern. The short version: v1's tight σ was two errors cancelling (an idealized racing line hiding an optimistic grip model), the frozen car was pulling a physically impossible 5.97 g, the mid-season rule change had never entered the physics, and air density was being treated as a constant when the real sessions ran anywhere from 1.07 to 1.22 kg/m³. Then a failed experiment paid for itself: trying to drop the survey-map dependency exposed a curvature bias in the line reconstruction that the tyre fit had been absorbing all along. The v2.2 line is now the real pole lap itself, and two circuits the model had never fit, Miami and Barcelona, were called blind against their 2026 poles: −0.29% and −0.24%. And one omission was hiding in plain sight: the model deployed a fixed 4 MJ and harvested nothing, in a formula whose whole story is the MGU-K recovering 350 kW in every braking zone. With the energy ledger in, deployment lands at the 7–9 MJ the superclipping era actually runs, and the tyre fit relaxed for the third time. That re-freeze closed at σ 1.32%, seven-fold out-of-fold, against a teammate noise floor of 0.37%. It did not stay there. An audit then found five defects in the physics, one of them cancelling another, and moved it to 1.46%; a leak in the calibration itself, an input production cannot have, brought it to the current 1.414%. Each move is documented in the repo: predictions/refreeze-2026-summer.md, refreeze-2026-v3.md and refreeze-2026-v4.md.
Backtests are one thing. The real test is the pre-registered calls, scored against public FastF1 timing. This table grows with the season, and the misses stay on it.