Stat Explainer

The Model Cannot See a Quarterback Change

This site's engine reads final scores and says plainly that it does not know who the quarterback is. Across 656 games since 2010 in which exactly one side started a different quarterback than the week before, that team won 36.5% of the time; the model gave it 45.8% and the closing market 38.9%. The model's 9.28-point shortfall is 5.2 standard errors from zero and the market's 2.39 is 1.4, and no draw of 2,000 random samples came close. Pricing it would mean docking the changing team 78 Elo points, more than the engine's entire home-field edge. It survives removing the weeks when contenders rest starters, and the ratings close most of it on their own within three games.

By C. B. Zakarian · Published September 17, 2026

The One Thing It Admits It Cannot See

The engine behind this site's weekly numbers reads final scores and nothing else. The page that states it in full says so in plain terms: it does not know who the quarterback is. That is usually filed as a charming limitation. This page puts a number on it.

Across the sixteen seasons the model has graded, there are 656 regular-season games in which exactly one of the two teams started a different quarterback from the one it started the week before. The model gave those teams a 45.8% chance of winning. The closing market gave them 38.9%. They won 36.5%. The model was 9.28 points too high, the market 2.39, and the model's shortfall is 5.2 standard errors from zero where the market's is 1.4.

Priced as a rating rather than a probability, the model would have needed to dock the changing team 78 Elo points to get those games right — more than the 48-point home-field edge it applies for playing at home. That is the size of the blind spot, and my position is that it is the largest single correction available to this model that it can never make, because the input does not exist in the file it reads.

How a Change Is Counted

The game log records both starting quarterbacks for every played game, by player id. The harness walks the file in date order, and flags a team as having changed if it has already played that season and the id differs from its previous game. A team's first game of a season is never a change, so a new starter installed in the offseason does not count — this is strictly about changes the model could not have absorbed from a full week's scores.

That gives 766 team-games starting a new quarterback from 2010 through 2025, spread across 286 team-seasons, between 27 and 64 of them a season: fewest in 2012, most in 2021. Of the games those teams played, 55 had both sides changing at once, which cancels the comparison and is set aside. The 656 that remain are the clean cases, and all of them carry a two-way closing moneyline, so both forecasters can be graded on the same games.

The comparison is deliberately one-sided in the model's favour. The model's rating already knows everything the team's scores have shown, including any games the departing quarterback played badly. The only thing it cannot know is who is taking the snaps on Sunday. The market knows that, and the gap between the two is what the information is worth.

The Exhibit

Left panel: seven groups of three bars showing the chance that a team starting a new quarterback wins, as the model expected, as the closing market expected, and as it actually happened. For all 656 games the bars read 45.8, 38.9 and 36.5 per cent, labelled minus 9.3. For weeks 1 to 16, 45.0, 39.1 and 37.9, labelled minus 7.1. For weeks 17 and 18, 50.9, 37.6 and 27.6, labelled minus 23.3. For weeks 1 to 8, 46.7, 39.7 and 35.8. For weeks 9 to 16, 43.9, 38.7 and 39.2. For 2010 to 2017, 46.1, 40.1 and 38.2. For 2018 to 2025, 45.5, 37.8 and 35.1. The model bar is higher than the actual bar in every group but one. Right panel: three lines across six points, by how many games have passed since that team last changed quarterback. The team actually won 38.5, 39.4, 42.0, 40.7, 44.7 and 43.9 per cent; the model expected 43.8, 44.2, 43.3, 43.0, 44.2 and 45.3; the market expected 41.1, 42.6, 42.2, 42.4, 43.3 and 45.4. The actual line starts far below both forecasts and joins them by the third game.
Left: the same three numbers under seven cuts of the 656 games. Right: the shortfall by how long the new starter has been in the job, for every team-game after a change. Data: nflverse game log (June 2026 bundle, 1999–2025) with its starting-quarterback ids and closing moneylines, replayed through nfl_elo.py.

Seven Cuts, One Direction

The changing team's chance of winningGamesModelMarketActualModel gapMarket gap
All 656 games65645.8%38.9%36.5%-9.3-2.4
Weeks 1–1656945.0%39.1%37.9%-7.1-1.2
Weeks 17–188750.9%37.6%27.6%-23.3-10.1
Weeks 1–822646.7%39.7%35.8%-10.8-3.9
Weeks 9–1634343.9%38.7%39.2%-4.7+0.5
2010–201730446.1%40.1%38.2%-8.0-2.0
2018–202535245.5%37.8%35.1%-10.4-2.7

The third row is the obvious objection, and it is a good one. Weeks 17 and 18 are when teams with a playoff seat settled rest their starters, and those 87 games are the worst the model has ever looked: it expected the changing team to win half the time and the team won 27.6%, a gap of 23.3 points. If that were the whole story, this page would be about the late-season rest pattern and not about quarterbacks.

It is not the whole story. Strip those two weeks out and the remaining 569 games still run 7.1 points under the model's number, which is four standard errors of the season-level estimate. The effect is worst in the first half of the season, 10.8 points across weeks 1 to 8, and smallest from week 9 to week 16, where it is 4.7 points and the market is actually a shade low. It has grown a little across eras, from 8.0 points in 2010–2017 to 10.4 since, which is what you would expect in a league that now changes quarterbacks 50-odd times a year rather than 30-odd.

In all seven cuts the market sits between the model and the truth, and much closer to the truth. That is the pattern that makes this a story about missing information rather than about a broken rating.

Could 656 Games Do This by Chance?

Two checks. Averaged season by season, the model's shortfall is 9.30 points with a standard error of 1.79, and the market's is 2.37 with a standard error of 1.72 — 5.2 standard errors against 1.4. The model overrated the changing team in 14 of the 16 seasons; the market in nine, which is what a fair coin does. The two seasons that went the other way for the model, 2014 and 2015, were also the two best seasons for the market, so they are a shared quiet spell rather than evidence against the effect.

The second check does not assume anything about standard errors. Draw the same number of games from each season at random, assign each one a random side, and measure how far that side's results fell below the model's expectation. In 2,000 such draws, the shortfall was never as large as the real one: the random draws centre on zero with a spread of about a point, and the observed value is 9.3 points below it. Whatever is happening in these games, it is not the accident of having chosen 656 of them.

Graded as Forecasts

Probabilities are cheap; scores are not. On the 656 change games the model's Brier score is .2329 and the market's is .2025. On the 3,463 control games where neither side changed, the model scores .2181 and the market .2119. So the market's usual edge of about six thousandths widens to about thirty on the change games — a factor of roughly five — and the market's score on them is better than its score on ordinary games, because a game with a backup quarterback in it is an easy game to price if you know that.

Straight up, on the 655 decided change games, the model goes .6305 and the market .6855. For scale, the model's overall accuracy is .6469 and the market's is a little under .68, so these games cost the model about a point and a half of hit rate while the market gains. The changing team is also outscored by 4.30 points a game on average, which is close to a field goal and a half of margin that no rating in this system can anticipate.

Worked Through: The Worst One

The single worst game in the file for the model is week 17 of 2020, the Chargers at Kansas City. The Chiefs had clinched and started Chad Henne; the engine, which cannot know that, had the best rating in the league next to a team it had beaten all year:

gap      = (1782.7 Kansas City + 48 home) - 1430.6 Chargers        = 400.1
p(KC)    = 1 / (1 + 10 ** (-400.1 / 400))                          = .9091

the market, same game
  Chargers -282 -> .7382     Kansas City +244 -> .2907     hold 2.89%
  p(KC)    = .2907 / (.7382 + .2907)                               = .2825

result: Chargers 38, Kansas City 21
  Brier, model  = (.9091 - 0) ** 2 = .8264        Brier, market = (.2825 - 0) ** 2 = .0798

Two forecasters 62.7 points apart on the same game, and the gap is one roster decision. The model's .8264 on that game is worth about a fifth of a point of its Brier score for the entire 2020 season on its own. In the five worst such games the market was lower on the changing team than the model every time, which is the cleanest way to see that this is information the model lacks rather than variance it was unlucky with.

The Engine Catches Up in About Three Games

A rating system that reads scores will learn a new quarterback the hard way: by watching the team play worse and marking it down. Here is every team-game after a change, grouped by how many games the new starter has had:

Games since that team changed quarterbackTeam-gamesModelMarketActualModel gap
The first game with the new starter66343.8%41.1%38.5%-5.3
1 game later44744.2%42.6%39.4%-4.8
2 games later33143.3%42.2%42.0%-1.3
3 games later24643.0%42.4%40.7%-2.3
4 games later19044.2%43.3%44.7%+0.6
5 or more games later47345.3%45.4%43.9%-1.4

The bill is paid in the first two games, 5.3 points and then 4.8, and by the third it is inside a point and a half, where it stays. Two poor results at K = 20 are worth enough rating points to close most of a 78-point gap, which is the update rule doing its job slowly and after the fact. The market's line is flatter throughout, as it should be: it knew on day one, so it has nothing to learn. Note that these rows pool the whole population of post-change team-games rather than the 656 clean matchups, which is why the first row holds 663 and not 656.

What This Page Does Not Show

The file does not say why. An injury, a benching and a rested starter in week 17 all look identical here — a different id in the quarterback column. Those three have different meanings, and the week 17–18 row is the only one of them this page can isolate. The continuity page has the season-level version of the same ambiguity.

Selection is real and only partly controlled. Teams change quarterbacks because something has gone wrong, and a team in that state may be worse for reasons beyond the position. The model's rating absorbs the team's results to date, and the market's price is the benchmark for everything else, but neither is a randomised experiment. Nothing here says the change caused the shortfall.

Seven cuts and a decay table. That is a lot of slicing on one dataset, and the 87-game week 17–18 cell is small. The claims I would defend are the pooled 9.28 points, its season-clustered standard error, the permutation test and the market comparison, all of which are computed on the full 656 before any cut.

78 Elo points is a description, not a patch. It is the constant that would have priced these games correctly in hindsight, fitted on the same games. It cannot be applied going forward, because the 2026 rows in the game file name no quarterback at all until a game is played, and the ledger's frozen rows have no quarterback field in any week.

The market numbers are the file's closing prices, de-vigged proportionally. A different de-vig moves them by a fraction of a point. I have no independent record of when each was captured, and a closing price already contains the news this page treats as the market's advantage.

This is history, not a 2026 claim. Every figure above comes from games played between 2010 and 2025.

Method and Sources

One file and one module. The June 2026 nflverse bundle (/data/games.csv, served at /data/games.csv): 7,276 games from 1999 to 2025, every played one carrying both starting quarterback ids and, from 2010, a two-way closing moneyline; its 272 2026 rows are schedule only and name nobody. And explainer_src/nfl_elo.py, imported rather than copied, whose replay reproduces the published backtest. The harness is explainer_src/make_qb_change_chart.py. The core of it:

changed = prev_starter.get((season, team)) not in (None, qb_id)   # a first game is never a change
# 656 games where exactly one side changed, 2010-2025
model  = mean(p assigned to the changing team)      = .4579
market = mean(vig-free closing price, same side)    = .3890
actual = mean(it won)                               = .3651
# what rating shift would price it:  solve for d in  mean(P(win | rating - d)) == actual  ->  78

The script asserts the engine constants, the bundle's row counts and the absence of any 2026 score or quarterback name, the replay's 4,175 graded regular-season games, the change counts and the team-seasons they span, all three headline rates and both gaps, the season-clustered standard errors and the count of seasons in each direction, all seven cuts cell by cell, the permutation test, both Brier scores and both control Brier scores, the straight-up records, the average margin, the Elo solution, the worked example down to the ratings, the hold and both Brier scores on it, the six decay rows, and the pre-kickoff snapshot's frozen scoreboard: 71 assertions, all green as of September 17, 2026.

Sources: the nflverse public game log (games.csv), including its starting-quarterback ids and closing moneylines. The Brier score is Glenn Brier's, from Verification of Forecasts Expressed in Terms of Probability (1950); the permutation argument is R. A. Fisher's, from The Design of Experiments (1935). The rating method is Arpad Elo's.

Further reading

About the author

C. B. Zakarian

C. B. Zakarian is an independent analyst who writes about what he can measure. He builds every model, chart, and calculator on this site himself from the public nflverse play-by-play and game-log releases, shows the working, and never invents a number. The dataset behind the exhibits is served openly at /data/, and the method behind every figure is spelled out so you can check it against the same file. When the data can't answer a question, he says so.

More Explainers
Do Favorites Cover? Scoring by Week Most Common Scores Favorite Win Rates Playoff Football Division Games & Home Field DVOA EPA vs. DVOA CPOE Passer Rating vs. QBR Pythagorean Wins Air Yards & YAC Fourth-Down Analytics Strength of Schedule ANY/A RYOE Pass Protection Coverage Metrics Special Teams PROE & Game Script Red Zone Efficiency Explosive Plays Third Down Time of Possession Turnovers & Luck Win Probability YAC Over Expected Snaps & Usage Points Per Drive Success Rate Pressure Rate Play-Action Yards After Contact RPO Two-Point Conversions Yards per Route Run Block Win Rates Target Share & WOPR Home-Field Advantage Expected Points Point Spread Accuracy Weather & Scoring Rest & Scheduling Scoring Trend Overtime Over/Under Accuracy Key Numbers (3 & 7) Thursday & Primetime Grass vs. Turf One-Score Games Stadium Scoring Referee Effects QB Continuity Week 1 Signal Shutouts 2026 Schedule Strength 2026 Schedule Quirks Best Record vs. Super Bowl Win & Loss Streaks Division Repeats Close-Game Luck The Prediction Model The Week 1 Slate AFC East 2026 AFC North 2026 AFC South 2026 AFC West 2026 NFC East 2026 NFC North 2026 NFC South 2026 NFC West 2026 Preseason Signal 2026 Preseason 2026 Win Totals 2026 Playoff Odds 2026 International Games 2026 Miss Budget AFC vs NFC The 17-Game Era The Coach Ledger Week 1 Predictions Opening Night 2026 SB Rematch Effect Road Favorites The Chiefs' Rating Division Leverage The Learning Curve The Board, Sorted September, Priced The Offseason Haircut Fair Prices What One Game Moves The Shortest Lines The Week 10 Problem The New-Coach Bounce Same Record, Different Rating The Opener, Graded Rankings After the Opener The First Miss, Graded What Sunday Can Do SF and LA, Re-Priced No Good Years, Only Lucky Ones Week 2 Overreaction Sunday, Graded Week 1 Scoring, 2026 What 0-2 Costs Monday Night, Graded Playoff Odds After Week 1 Calibrated Playoff Odds Thursday: DET at BUF Game-Pick Calibration The QB Blind Spot All explainers

Go deeper

Want the code behind these metrics? Work through the 45-chapter NFL analytics tutorial.

Browse tutorials Free tools