Stat Explainer

The Pattern That Only Works Backwards

Sunday's six misses averaged 33.0 combined points and its eight hits 42.9, which looks like a mechanism: fewer possessions, fewer chances for the better team to prove it. Across 4,350 graded games since 2010 the pattern is real — the lowest-scoring fifth grades 62.2% against a claimed 65.1%. Then it dies twice. The gap is 1.67 standard errors, on a hypothesis that came from the slate rather than from theory. And sorting the same games by the PREGAME total line, the only version available before kickoff, flattens it completely: a correlation of +.0085 against +.0287. The market is unbiased about scoring and still tells you nothing about when this model is wrong.

By C. B. Zakarian · Published September 21, 2026

The Pattern Was Real. It Was Also Useless.

Sunday's six misses averaged 33.0 combined points. Its eight correct picks averaged 42.9. The lowest-scoring game on the board, Minnesota 9 at Chicago 3, was a miss; the two highest, Dallas 37–20 and Kansas City 33–30, were both hits. Yesterday's grading note declined to read anything into a fourteen-game slate, which was the right call for the things it was refusing. This is the one pattern in that slate worth following, because it comes with a mechanism: a low-scoring game has fewer possessions, fewer possessions means the better team gets fewer chances to be the better team, and a forecast built on ratings should therefore land less often.

It does. Across the 4,350 games this engine has graded since 2010, sort them into equal fifths by final combined points and the lowest fifth — games at 27 points on average — grades at 62.2% against a claimed 65.1%. The best bucket, games landing between 48 and 56, grades at 67.3%. That is a five-point spread in accuracy, and the Brier score moves with it: .2299 in the lowest bucket against .2087 in the best.

Then it falls apart, in two stages, and my position is that the second stage is the one worth learning. Sort the same 4,350 games by the pregame total line instead — the one version of "this will be a low-scoring game" you can have before kickoff — and the gradient disappears. The correlation between the final total and whether the pick landed is +.0287. Between the pregame line and whether the pick landed it is +.0085, on a standard error of .0152. The market's best estimate of how much scoring is coming tells you nothing about whether this model is about to be wrong.

The Exhibit

Two bar charts side by side, both showing the share of the model's picks that won, on a vertical axis from 57 to 70 percent, with one-standard-error bars and a dashed red line marking what the model claimed in each bucket. Left panel, by final combined points in equal fifths: 62.2 percent at a bucket mean of 27 points over 852 games, 64.0 at 37 points over 762, 64.6 at 44 points over 914, 67.3 at 52 points over 903, and 65.1 at 65 points over 919. The claimed line is nearly flat near 65 percent across all five. The spread is 5.1 points and the correlation is plus 0.0287. Right panel, by pregame total line in equal fifths: 63.4 percent at a line mean of 39.1 over 841 games, 66.4 at 42.6 over 892, 64.2 at 44.7 over 796, 64.4 at 47.0 over 851, and 64.9 at 51.1 over 970. There is no visible trend. The spread is 3.0 points and the correlation is plus 0.0085.
Left: sorted by what the game turned out to be. Right: sorted by what the market said it would be, in advance. Same games, same model, same buckets. Data: nflverse game log, 4,350 decided games 2010–2025, replayed through the engine.

Both Sorts, Side by Side

Sorted by final combined pointsGamesClaimedWonGapBrier
Under 34 — mean 2785265.05%62.21%−2.84.2299
34–40 — mean 3776265.22%64.04%−1.18.2254
41–47 — mean 4491465.49%64.55%−0.94.2230
48–56 — mean 5290365.84%67.33%+1.49.2087
57 and up — mean 6591965.55%65.07%−0.48.2192
Sorted by pregame total lineGamesClaimedWonGapBrier
Under 41.5 — mean 39.184164.75%63.38%−1.37.2283
41.5–44 — mean 42.689265.88%66.37%+0.49.2156
44–46 — mean 44.779665.50%64.20%−1.30.2243
46–48.5 — mean 47.085165.80%64.39%−1.40.2212
48.5 and up — mean 51.197065.29%64.95%−0.34.2167

The top table has a shape. The bottom table has five numbers between 63.4 and 66.4 in no order at all, and its largest single gap belongs to the second bucket, not the lowest. Same games, same model, same bucket sizes. The only difference is whether the sorting variable was available before kickoff.

How Big Is the Real Effect? Smaller Than It Looks

Before drawing a lesson from the top table, the top table deserves a harder look, because I do not think it is as solid as its shape suggests.

Compare the lowest-scoring fifth against the other four together. The low bucket grades 62.21% over 852 games; everything else grades 65.29% over 3,498. The difference is 3.09 percentage points:

p1 = 0.6221  n1 =  852        p2 = 0.6529  n2 = 3498
diff = 0.0309
se   = sqrt( p1(1-p1)/n1 + p2(1-p2)/n2 ) = 0.0185
z    = 0.0309 / 0.0185 = 1.67

z = 1.67. That is short of the conventional threshold, and it should be read as short of it by more than the arithmetic suggests, because this hypothesis did not arrive from theory. It arrived from looking at one Sunday, noticing that the misses were the low-scoring games, and then going to find out whether that held. That is the wrong order, and it is exactly the order that turns noise into findings. A test built to confirm something you have already seen needs a higher bar than 1.67, not a lower one.

The whole-sample correlation says the same thing more bluntly. At n = 4,350 the standard error on a correlation is .0152. The final-total correlation of +.0287 is 1.9 standard errors; the pregame-line correlation of +.0085 is barely half of one. Neither is the kind of number anyone should build a model change on. The first one is at least pointing somewhere.

Why the Forecastable Version Is Flat

The obvious objection is that the pregame total is simply a bad forecast of the final total, so of course the gradient washes out. It is not a bad forecast. Across the same 4,350 games the mean line is 45.05 and the mean final total is 45.56 — the market is off by half a point on average, which is about as unbiased as an NFL number gets, a result this site has measured directly.

The market is unbiased and still useless here, and the two facts sit together comfortably once you separate the average from the scatter. A total line is a good estimate of the mean of a distribution whose spread is enormous: individual games miss their line by a 13.4-point standard deviation. A game listed at 39.0 is genuinely more likely to land under 34 than a game listed at 51.0, but it is nowhere near certain to, and the low-scoring games that actually hurt the model are scattered across every bucket of the line.

Put it the other way round, which is the version that matters for a forecaster: the thing that makes a game hard for this model is not that it was expected to be low-scoring. It is that it came out low-scoring — the defensive slog, the two turnovers inside the twenty, the rain nobody priced. Those are realisations, not conditions. You cannot bet on a realisation and you cannot condition a pregame probability on one.

That distinction is worth more than the effect it kills. A variable that predicts your errors is only worth having if you hold it before you commit. This site has already retired one repair for exactly this reason: the quarterback-change correction is worth 78 Elo points in hindsight and cannot be applied forward, because nobody knows in September who will change quarterback in November. Final scoring is the same defect wearing different clothes.

Worked Through: Sunday Under Both Sorts

The slate was genuinely low-scoring. Its fourteen games averaged 38.6 combined points against the sixteen-year mean of 45.6, which puts it in the bottom quarter of the 298 slates of thirteen or more games in the dataset. Three of the fourteen had a team score three points or fewer.

So the hindsight sort explains Sunday well. The forecastable sort does not get the chance to: at the moment the picks were locked, the model knew the ratings and the schedule, and nothing in what it knew said "this will be the fortnight's ugly slate". Take the two extremes:

Minnesota at Chicago   pick CHI at 53.4%   final  9- 3  =  12 pts   MISS
Indianapolis at KC     pick KC  at 70.1%   final 33-30  =  63 pts   HIT

after the fact: only 14 of the 4,350 games finished at 12 points or fewer
before kickoff: nothing separated these two games except the ratings themselves,
                which is the information the 53.4% and the 70.1% already contain

The Chicago pick was a 53.4% coin flip that lost, and the model said so at the time. Reclassifying it afterwards as "one of the hard ones" adds nothing, because the 53.4% had already said it was hard. That is the trap in the top table: it looks like a correction waiting to be applied, and what it actually describes is information the probability was already carrying.

What This Page Does Not Show

The hypothesis was generated by the data it is tested on, once removed. Sunday suggested it; history tested it. That is better than testing it on Sunday, and it is worse than having predicted it in advance. I have reported the test as weak rather than rounding 1.67 up into a finding.

Five buckets is a choice, and so is the variable. Equal-count fifths on combined points is one of many defensible cuts. I did not sweep bucket counts looking for a version with a bigger gap, and if I had, the first thing to say about any gap that appeared would be that I went looking for it.

This is not a claim that scoring environment is irrelevant to football. It is a narrow claim about one rating model's pick accuracy: the pregame total line does not improve it. Other models, and other questions about the same games, may well find the line informative.

No closing lines are used. The bundle's total_line is the game's listed total as nflverse records it, not a verified close, and this page makes no claim about any market's efficiency or about what anything would have returned.

Sunday's fourteen games are not part of the 4,350. The historical replay ends in 2025. The slate appears here only as the thing that raised the question.

Method and Sources

One file: the June 2026 nflverse bundle (/data/games.csv, served at /data/games.csv), replayed game by game through explainer_src/nfl_elo.py with 1999–2009 as burn-in and 2010–2025 graded — 4,350 decided games, every one of them carrying a pregame total line. The bundle's 2026 rows are schedule-only and scoreless, and the harness asserts that before it starts, so no live result reaches a historical figure. Sunday's own fourteen games are read from the pin yesterday's page captured (_honest_not_sharp_2026-09-21.json), because a site build republishes the June bundle over the served 2026 scores.

for each graded game:
    p    = engine probability on the side it picked
    hit  = 1 if that side won
    tot  = home_score + away_score      # known only afterwards
    tl   = total_line                   # known before kickoff

for key in (tot, tl):
    five equal-count buckets by key
    report claimed = mean p, actual = mean hit, and their gap
    corr(key, hit) against a standard error of 1/sqrt(n-3)

The harness is explainer_src/make_hindsight_variable_chart.py. It asserts the bundle carries no 2026 score; the replay's game count and its reproduction of the published backtest at .6469 accuracy and .2210 Brier; that every graded game has a total line, and the two mean totals; both sets of quintile cuts, that all ten buckets hold between 700 and 1,000 games and that each set partitions the sample exactly; every claimed rate, actual rate, gap and Brier in both tables; that the lowest-scoring bucket is the worst of its five on both measures; that the final-total spread exceeds five points while the total-line spread stays under three; both correlations, their shared standard error and the ratio between them; the two-proportion test down to its z; and from the pin, Sunday's split of six misses and eight hits with their mean totals, the slate mean, its rank among the 298 comparable slates in the bundle, the twelve-point Minnesota game and the three teams held to three points or fewer: 68 assertions, all green as of September 21, 2026.

Sources: the nflverse public game log (games.csv), whose total_line field supplies every pregame total. The rating method is Arpad Elo's, from The Rating of Chessplayers, Past and Present (1978); the Brier score is Glenn Brier's, from Verification of Forecasts Expressed in Terms of Probability (1950).

Further reading

About the author

C. B. Zakarian

C. B. Zakarian is an independent analyst who writes about what he can measure. He builds every model, chart, and calculator on this site himself from the public nflverse play-by-play and game-log releases, shows the working, and never invents a number. The dataset behind the exhibits is served openly at /data/, and the method behind every figure is spelled out so you can check it against the same file. When the data can't answer a question, he says so.

More Explainers
Do Favorites Cover? Scoring by Week Most Common Scores Favorite Win Rates Playoff Football Division Games & Home Field DVOA EPA vs. DVOA CPOE Passer Rating vs. QBR Pythagorean Wins Air Yards & YAC Fourth-Down Analytics Strength of Schedule ANY/A RYOE Pass Protection Coverage Metrics Special Teams PROE & Game Script Red Zone Efficiency Explosive Plays Third Down Time of Possession Turnovers & Luck Win Probability YAC Over Expected Snaps & Usage Points Per Drive Success Rate Pressure Rate Play-Action Yards After Contact RPO Two-Point Conversions Yards per Route Run Block Win Rates Target Share & WOPR Home-Field Advantage Expected Points Point Spread Accuracy Weather & Scoring Rest & Scheduling Scoring Trend Overtime Over/Under Accuracy Key Numbers (3 & 7) Thursday & Primetime Grass vs. Turf One-Score Games Stadium Scoring Referee Effects QB Continuity Week 1 Signal Shutouts 2026 Schedule Strength 2026 Schedule Quirks Best Record vs. Super Bowl Win & Loss Streaks Division Repeats Close-Game Luck The Prediction Model The Week 1 Slate AFC East 2026 AFC North 2026 AFC South 2026 AFC West 2026 NFC East 2026 NFC North 2026 NFC South 2026 NFC West 2026 Preseason Signal 2026 Preseason 2026 Win Totals 2026 Playoff Odds 2026 International Games 2026 Miss Budget AFC vs NFC The 17-Game Era The Coach Ledger Week 1 Predictions Opening Night 2026 SB Rematch Effect Road Favorites The Chiefs' Rating Division Leverage The Learning Curve The Board, Sorted September, Priced The Offseason Haircut Fair Prices What One Game Moves The Shortest Lines The Week 10 Problem The New-Coach Bounce Same Record, Different Rating The Opener, Graded Rankings After the Opener The First Miss, Graded What Sunday Can Do SF and LA, Re-Priced No Good Years, Only Lucky Ones Week 2 Overreaction Sunday, Graded Week 1 Scoring, 2026 What 0-2 Costs Monday Night, Graded Playoff Odds After Week 1 Calibrated Playoff Odds Thursday: DET at BUF Game-Pick Calibration The QB Blind Spot DET at BUF, Graded The Second Meeting Three Fixes, One Knob The Engine's Constants The Margin Multiplier Honest, Not Sharp A Hindsight Variable All explainers

Go deeper

Want the code behind these metrics? Work through the 45-chapter NFL analytics tutorial.

Browse tutorials Free tools