Stat Explainer

Three Fixes, One Knob

In four days this site measured three defects in its own forecasting and shipped none of them: a playoff simulation too sure of itself, a quarterback change the engine cannot see, and division rematches swept more often than two independent prices imply. Applied together to 768 opening-day team-seasons and scored out of sample on 608, they turn out to be very nearly one correction. Rating uncertainty alone takes the Brier score from .22388 to .21348 and the calibration slope from 0.607 to 1.007, better in 16 of 19 seasons; the rematch link adds 0.000013 on top of that, because the first correction already drives the simulated sweep rate past what history shows; and the quarterback fix cannot be applied in advance at all, because nobody knows in September who will change quarterback in November.

By C. B. Zakarian · Published September 19, 2026

Three Defects, Measured Separately

In the last four days this site has measured three things wrong with its own forecasting and shipped none of them. The playoff simulation is too sure of itself, because it treats every rating as exactly right; the repair is to assume each rating is wrong by a normal draw of some number of Elo points. The engine cannot see a quarterback change, which is worth 78 Elo points in hindsight. And division rematches are swept more often than two independent prices imply, which a latent correlation of 0.0945 reproduces.

Each was fitted on its own, and each page was careful to say it was not proposing a change to the shipped numbers. The obvious next question is what happens when all three are applied at once: how much of the calibration gap do they close together, and do they add up?

They do not add up, and my position is that this is the useful finding rather than a disappointment. Priced together on 768 team-seasons and scored out of sample on 608, the rating uncertainty does essentially the whole job — Brier .22388 to .21348, calibration slope 0.607 to 1.007, better in 16 of 19 test seasons. The rematch correlation, applied on top of it, is worth 0.000013 of Brier score with a season-level standard error of 0.000037. And the quarterback correction cannot be applied going forward at all. Three separately-fitted repairs turn out to be very nearly one repair, and the two that survive are the same repair twice.

The Test Bed, and Why It Is the Opening Day

Every figure here is scored on the opening-day board of each season from 2002 to 2025: the ratings as they stood before a single game was played, carried through the site's own 20,000-season simulation, and graded against who actually made the playoffs. That is 24 seasons and 768 team-seasons, of which the 19 seasons from 2007 are scored out of sample, with each season's parameter chosen only on the seasons before it.

Two reasons for that choice. The first is that the repair page fitted its parameter on the board after week 1 and left opening day as a footnote, so this is the harder and less-examined version of the same question. The second is that an opening-day board is finished history: it is computed from the previous season's results and nothing else, so no game played this year can move a number on this page. The 2026 column below is the ledger's frozen August board, locked on 2026-08-12, with Seattle at 1674.6. Nothing here previews or grades a 2026 result.

The harness reproduces the earlier pages before it does anything new. The after-week-1 board comes back at a Brier score of .2082 and a slope of 0.684, and its repair at .1997, which are the three numbers the repair page published. The opening-day board comes back at .2247 and 0.596, which is that page's footnote. Nothing below is a new simulation; it is the same one, asked a longer question.

The Exhibit

Left panel: a calibration plot of 608 opening-day team-seasons scored out of sample from 2007 to 2025, with the band-average simulated playoff probability on the horizontal axis and the share that made the playoffs on the vertical axis, against a dotted diagonal. Grey circles for the simulation as it ships run far flatter than the diagonal, from 24 per cent where it promised 6 to 89 per cent where it promised 94. Dark blue diamonds for the rating-uncertainty version and green triangles for rating uncertainty plus the rematch link lie on top of each other along the diagonal, the two sets of markers overlapping at every band. A box gives slope 0.61 then 1.01 then 1.01, Brier 0.2239 then 0.2135 then 0.2135, log loss 0.6598 then 0.6163 then 0.6163. Right panel: four bars showing how often the same team sweeps a division home-and-home in the simulated 2026 season - 54.41 per cent as it ships, 56.89 with the rematch link alone, 62.10 with rating uncertainty alone and 64.05 with both - against a dashed red line at the historical 57.31 per cent, which the last two bars stand well above.
Left: the second and third sets of markers are indistinguishable, which is the finding. Right: why — rating uncertainty already produces more correlation between a division pair's two games than history shows, so the correction built to supply that correlation has nothing left to add. Data: nflverse game log 1999–2025, replayed through nfl_elo.py, and the site's own season simulation.

How Wrong the Ratings Should Assume They Are

The first correction has one parameter: the standard deviation of the rating error the simulation assumes before each simulated season. Here is the sweep across the 768 team-seasons, with and without the rematch link applied on top:

Rating uncertainty (Elo points)BrierCalibration slopeBrier, with the rematch linkWhat the link adds
00.224750.5960.22437+0.000383
400.221560.6350.22132+0.000245
800.216480.7420.21636+0.000121
1000.214920.8110.21488+0.000034
1200.213960.8890.21399-0.000030
1400.213590.9730.21358+0.000013
1600.213741.0570.21372+0.000016
2000.214721.2310.21474-0.000014

Two things to read here. The opening-day board wants 140 points of assumed rating error, not the 120 the after-week-1 board wanted, and the minimum is interior: 200 is worse than 140. That is the right direction for the wrong-looking reason. A board built in February has had a whole offseason to become wrong about rosters, and by late September it has at least seen everyone play once, so the earlier board should trust itself less. The repair page said as much in a sentence; this is the number.

The last column is the second finding arriving early. On the raw board the rematch link is worth +0.000383, a small positive. By 80 points of rating uncertainty it is a third of that, and from 120 onward it is bouncing either side of zero at the fifth decimal place. The correction is not being outvoted; it is being made redundant.

All Three, Scored Out of Sample

Fitted walk-forward — each test season's parameter chosen only on seasons before it, nineteen test seasons, 608 team-seasons — the arms come out like this:

Out of sample, 2007–2025BrierLog lossCalibration slopePlayoff seats a season
The simulation as it ships0.223880.659760.60712.632
Rating uncertainty (v68)0.213480.616301.00712.632
Rating uncertainty + rematch link (v68 + v73)0.213460.616271.01012.632

Rating uncertainty gains 0.01040 of Brier score with a season-level standard error of 0.00309, which is 3.4 standard errors, and it helps in 16 of the 19 seasons. On the opening-day board it is worth a third more than the repair page measured after week 1, which follows from the board being worse to begin with. The fitted parameter is steady: across nineteen expanding windows it never leaves 140 to 160, and it ends at 140.

Adding the rematch link to that moves the Brier score by 0.000013 against a standard error of 0.000037, and the log loss by three hundred-thousandths. The first correction is more than a hundred times the second. All three arms still hand out 12.632 playoff seats a season, which is the real average field over this window — the test the shrinkage version of the repair failed, and the reason the repair page preferred this one.

Why the Second Correction Has Nothing Left to Do

A null is more convincing when you can say what happened to the effect, so here is the mechanism in one measurement. Take the 48 division home-and-home pairs on the 2026 schedule and count how often the simulation has the same team winning both:

simulated sweep rate, 48 division pairs, 20,000 seasons
  as it ships (independent draws)                 .5441
  with the rematch link alone                     .5689
  with rating uncertainty alone                   .6210
  with both                                       .6405

what 1,340 historical pairs actually did          .5731

The shipped simulation draws every game independently and comes in below the historical sweep rate, which is exactly the defect the rematch page identified. Applied on its own, the copula moves it to .5689 and lands close to the real number. But rating uncertainty, which was fitted for something else entirely, takes it to .6210 — already past history in the other direction. Adding the link on top pushes it to .6405.

That is not a coincidence. Perturbing a team's rating once per simulated season and then pricing all seventeen of its games off the perturbed number induces correlation between every pair of that team's games, including the two legs of a division series. The rematch page argued that the two results are dependent because both are driven by the same unobserved facts about how good the teams are; rating uncertainty is a model of precisely that, so it produces the dependence as a by-product. The copula then adds a second helping of the same thing.

Which means the honest reading of the pair is not “the rematch finding was wrong”. It was measured on the shipped simulation, and on the shipped simulation it is real. It is that two corrections aimed at different symptoms turn out to be aimed at the same cause, and the one fitted on 608 team-seasons of playoff outcomes dominates the one fitted on 1,340 rematches.

The Third Fix Cannot Be Applied at All

The quarterback correction is different in kind, and the difference is not subtle: nobody knows in September who will change quarterback in November. The 78-point dock is a number fitted on games where the change had already happened. There is no version of it that can be applied to an opening-day board, because the input does not exist yet.

What can be done is to ask how much of the rating uncertainty it accounts for. Split the 768 team-seasons by whether that team started more than one quarterback during the season, which is knowable only afterwards:

Team-seasons, split by what happened at quarterbackTeam-seasonsSeptember saidMade playoffsGapBest rating uncertainty
Kept one starting quarterback all season30042.6%52.3%+9.8120
Changed at least once46836.8%30.6%-6.3160

The 300 team-seasons that kept one starter made the playoffs 9.75 points more often than the opening-day board said they would; the 468 that did not came in 6.25 points under. Sixteen points of playoff probability separate two groups that are indistinguishable on the day the board is published. And the group that changed quarterback wants 160 Elo points of assumed rating error against 120 for the group that did not.

So the third finding is not a fourth column in the table above. It is an explanation of the second one. A large part of what the 140-point knob is paying for is the fact that a fifth of the league will be starting a different quarterback by December and the September board has no way to know which fifth. Priced as a separate correction it would be double-counting; priced as a reason the knob has to be that large, it is the most useful thing on this page.

Worked Through: Why Noise Flattens

The arithmetic that makes the whole page work is one line of convexity, and it is worth doing by hand. Take the two ends of the frozen August board — Seattle at 1674.6 and Tennessee at 1349.8 — and suppose they met at Seattle:

gap        = 1674.6 + 48 (home) - 1349.8              = 372.8
p(Seattle) = 1 / (1 + 10 ** (-372.8 / 400))            = .8953

now assume each rating is wrong by 140 points, and take the two equally likely cases
  Seattle 140 low, Tennessee 140 high:  gap = 92.8     = .6305
  Seattle 140 high, Tennessee 140 low:  gap = 652.8    = .9772
  average of the two branches                          = .8038

The average of the two perturbed probabilities is 9.15 points lower than the unperturbed one, because the logistic curve is concave in that region: the branch where Seattle is even better cannot gain as much as the branch where it is worse loses. Run that over 272 games and 20,000 seasons and every confident team's odds come down while every hopeless team's come up, which is the whole of the correction. It is also why the correction can never reorder a board, and why nothing in the 2026 table below changes who is favoured.

The Frozen 2026 Board, All Three Ways

For scale, the ledger's frozen August board carried through the same three arms. This is not the site's playoff-odds tablethe August page stands as written, and this is an argument about it. It is also deliberately the opening-day board rather than a current one, so that every number here is fixed rather than something that moves each weekend.

TeamAs it shipsRating uncertaintyAll threeChange
Seattle93.777.277.2−16.5
Houston85.469.069.0−16.4
Denver84.368.268.2−16.2
Philadelphia81.964.864.8−17.0
Buffalo80.665.865.8−14.8
LA Rams76.561.961.8−14.7
New England72.861.061.0−11.8
Baltimore71.660.260.1−11.5
Detroit70.759.058.8−11.8
Jacksonville67.858.158.0−9.8
Minnesota61.853.453.2−8.6
San Francisco56.750.850.7−6.0
Pittsburgh47.246.646.6−0.6
Green Bay44.944.945.0+0.1
Tampa Bay44.843.943.9−0.8
LA Chargers42.344.644.6+2.4
Cincinnati39.642.442.4+2.8
Kansas City37.642.342.3+4.7
Chicago37.240.840.8+3.6
Atlanta36.639.139.2+2.6
New Orleans28.534.935.0+6.5
Indianapolis27.936.836.8+8.9
Washington19.631.731.6+12.0
Dallas19.331.631.6+12.3
Cleveland19.031.631.6+12.6
Miami16.230.030.1+13.8
Carolina13.625.725.8+12.2
NY Giants11.124.624.7+13.6
Arizona3.315.715.8+12.5
Tennessee2.915.615.6+12.7
NY Jets2.814.414.4+11.7
Las Vegas1.913.413.5+11.5

The third column and the fourth are the same column. Across all 32 teams the three fixes together move 313.2 points of playoff probability, of which the rating uncertainty supplies 311.8 and the rematch link adds 1.7. The largest single contribution from the second correction anywhere on the board is two tenths of a point, on Detroit.

The shape is the familiar one: seven teams lose more than ten points and ten teams gain more than ten, nobody crosses fifty per cent, and the bottom of the board stops being written off — Las Vegas from 1.9% to 13.5%. Whether that is an improvement is not a matter of taste; over nineteen out-of-sample seasons the flattened version scored better in 16 of them.

What This Page Does Not Show

Three corrections fitted separately, combined afterwards. That is the overfitting risk in this kind of work and it is worth naming plainly. Two of the three were fitted on the same 24 seasons they are then scored on, in the sense that the grid of candidate values was chosen by looking at the data; only the walk-forward column protects against it, and it protects the rating-uncertainty parameter rather than the choice of which repairs to try. The order in which I applied them also matters: the rematch link looks worthless because rating uncertainty went first. Applied the other way around, the link is worth +0.000383 and the uncertainty still fixes almost everything that is left.

This is not the site's playoff-odds table and must not be read as one. The published odds are the August page's. Adopting any of this means changing the simulation in the build, which is a separate decision, and on this evidence I would change one thing rather than three.

Nineteen test seasons is nineteen observations. The gain of 0.01040 carries a season-level standard error of 0.00309, and six of those seasons have a different playoff format from the other thirteen.

The noise model is the simplest possible one. Independent, normal, identical for every team, drawn once a season and held. The quarterback split above is direct evidence that it should not be identical for every team — some teams are more forecastable than others — but a correction that needs to know which teams in advance is not available.

The sweep-rate comparison is not like-for-like. The .5731 is 1,340 historical pairs with both meetings decided; the simulated rates are 48 pairs of one particular schedule drawn 20,000 times. They are close enough in construction to make the point about direction and size, and not close enough to read the third decimal.

Ties are still broken at random in the simulation, with no head-to-head and no division records, exactly as the August page disclosed.

The random stream differs from the shipped one. Applying a copula needs a normal draw where the shipped simulation compares a uniform draw directly — the same coin, thrown differently. The raw column above reproduces the published August table to within 1.25 points on all 32 teams, and 0.37 on average, which is the Monte-Carlo noise of 20,000 seasons rather than a difference of method.

Method and Sources

Two files and one module. The June 2026 nflverse bundle (/data/games.csv, served at /data/games.csv): 7,276 games played from 1999 to 2025, with its 272 2026 rows schedule-only, so no 2026 result can reach a figure here. The frozen ledger at /data/predictions.json, from which this page takes only the August rating fields, never the scoreboard. And explainer_src/nfl_elo.py, imported rather than copied, whose replay reproduces all 32 frozen ratings. The simulation is the one in make_playoff_odds_chart.py, with the field size as a parameter. The harness is explainer_src/make_three_fixes_chart.py. The three corrections are these lines:

# v68: one rating draw per team per simulated season
e   = rng.normal(0.0, sigma, size=(20000, 32))
gap = (rating[home] + e[:, home]) - (rating[away] + e[:, away]) + home_field

# v73: the two legs of a division pair, linked (the venue flips, so the sign does)
Z[:, leg2] = -rho * Z[:, leg1] + sqrt(1 - rho ** 2) * Z[:, leg2]

# v71: not applicable - the input does not exist until the change has happened

The script asserts the engine constants, the bundle's row counts and the absence of any 2026 score, the 24-season census and its playoff-field sizes, the reproduction of the earlier pages on both boards, the full sigma sweep with its interior minimum and monotone slope, both walk-forward arms with their season-level standard errors and win counts, the seat counts in all three states, the fitted parameter path, the four simulated sweep rates against the historical one, the quarterback split with both gaps and both fitted parameters, the 2026 board cell by cell against the published August table, the worked example to four places, and this page's own figures: 59 assertions, all green as of September 19, 2026.

Sources: the nflverse public game log (games.csv). The Brier score is Glenn Brier's, from Verification of Forecasts Expressed in Terms of Probability (1950); the dependence structure is a Gaussian copula, after Abe Sklar's theorem (1959); the rating method is Arpad Elo's.

Further reading

About the author

C. B. Zakarian

C. B. Zakarian is an independent analyst who writes about what he can measure. He builds every model, chart, and calculator on this site himself from the public nflverse play-by-play and game-log releases, shows the working, and never invents a number. The dataset behind the exhibits is served openly at /data/, and the method behind every figure is spelled out so you can check it against the same file. When the data can't answer a question, he says so.

More Explainers
Do Favorites Cover? Scoring by Week Most Common Scores Favorite Win Rates Playoff Football Division Games & Home Field DVOA EPA vs. DVOA CPOE Passer Rating vs. QBR Pythagorean Wins Air Yards & YAC Fourth-Down Analytics Strength of Schedule ANY/A RYOE Pass Protection Coverage Metrics Special Teams PROE & Game Script Red Zone Efficiency Explosive Plays Third Down Time of Possession Turnovers & Luck Win Probability YAC Over Expected Snaps & Usage Points Per Drive Success Rate Pressure Rate Play-Action Yards After Contact RPO Two-Point Conversions Yards per Route Run Block Win Rates Target Share & WOPR Home-Field Advantage Expected Points Point Spread Accuracy Weather & Scoring Rest & Scheduling Scoring Trend Overtime Over/Under Accuracy Key Numbers (3 & 7) Thursday & Primetime Grass vs. Turf One-Score Games Stadium Scoring Referee Effects QB Continuity Week 1 Signal Shutouts 2026 Schedule Strength 2026 Schedule Quirks Best Record vs. Super Bowl Win & Loss Streaks Division Repeats Close-Game Luck The Prediction Model The Week 1 Slate AFC East 2026 AFC North 2026 AFC South 2026 AFC West 2026 NFC East 2026 NFC North 2026 NFC South 2026 NFC West 2026 Preseason Signal 2026 Preseason 2026 Win Totals 2026 Playoff Odds 2026 International Games 2026 Miss Budget AFC vs NFC The 17-Game Era The Coach Ledger Week 1 Predictions Opening Night 2026 SB Rematch Effect Road Favorites The Chiefs' Rating Division Leverage The Learning Curve The Board, Sorted September, Priced The Offseason Haircut Fair Prices What One Game Moves The Shortest Lines The Week 10 Problem The New-Coach Bounce Same Record, Different Rating The Opener, Graded Rankings After the Opener The First Miss, Graded What Sunday Can Do SF and LA, Re-Priced No Good Years, Only Lucky Ones Week 2 Overreaction Sunday, Graded Week 1 Scoring, 2026 What 0-2 Costs Monday Night, Graded Playoff Odds After Week 1 Calibrated Playoff Odds Thursday: DET at BUF Game-Pick Calibration The QB Blind Spot DET at BUF, Graded The Second Meeting Three Fixes, One Knob The Engine's Constants The Margin Multiplier All explainers

Go deeper

Want the code behind these metrics? Work through the 45-chapter NFL analytics tutorial.

Browse tutorials Free tools