Stat Explainer

The Opener, Graded: Seattle by Three, and What That Proves

Seattle 13, New England 10, and the model's .6785 is one for one. This page grades the site's own pre-registrations: the move landed on Tuesday's table to the hundredth — +8.42, a 13th-percentile night against 7,276 historical moves — the 24-point crossover never came into play, and no week-2 pick flipped. Then the harder part. Four pre-kickoff numbers all picked Seattle, so the winner separated none of them; the model's .1033 Brier beats the June market's .1262 and a records-only .1861 only because it was the most confident, and the order reverses in the counterfactual. The spread pushed at the close. The miss budget is .32 lighter. One game grades a pick; it cannot grade a probability.

By C. B. Zakarian · Published September 10, 2026

Seattle 13, New England 10

The first number of the season went up on August 12 and came down on Wednesday night. Seattle beat New England 13–10 at Lumen Field, the model's pick was Seattle at .6785, and the ledger now reads one for one with a Brier score of .1034. That is the grade. The rest of this page is about what the grade is worth, which is a different and smaller thing than it looks.

Two pages on this site made specific, dated claims about this game before it was played, and both are now checkable. What One Game Moves, published on September 8, priced every possible margin in both directions and said where each would land on a scale of 7,276 historical rating moves. Two 14–3 Teams, 82 Points Apart, published the morning of the game, put three competing probabilities on the record: a records-only view at about 57%, the model at 67.9%, and the vig-free market at 64.5%, plus a model spread of 4.61 points against the market's 3.5. I am going to grade all of it, in public, using a harness that recomputes every figure from the game file rather than from my memory of what I wrote.

The short version: the mechanics landed to the hundredth, because they are arithmetic. The probabilities all "won", because all of them favoured Seattle, and one game cannot tell you which of them was right in the sense that matters for a probability. The spread pushed at the close. Nothing on the board changed rank. And the miss budget is 0.32 lighter than it was on Tuesday, which is exactly what a 68% favourite winning is supposed to cost.

Where the Night Landed on Tuesday's Table

The September 8 page published a table of eleven margins, both branches, with the resulting ratings. A three-point Seattle win was the second row: +8.42 rating points to Seattle, Seattle to 1683.0, New England to 1584.4. The pipeline's published board on Thursday morning reads Seattle 1683.0 and New England 1584.4. That is not a prediction coming true; it is a formula being a formula. The update rule is

multiplier = log(3 + 1) * (2.2 / (0.001 * 129.79 + 2.2))  = 1.3863 * 0.944292 = 1.3091
shift      = 20 * 1.3091 * (1 - 0.678549)                  = 8.416
Seattle     1674.589 + 8.416 = 1683.005    New England  1592.802 - 8.416 = 1584.386

and the harness for this page replays the engine over all 7,276 played games in the file, reproduces the frozen 1674.6 and 1592.8 to the tenth, feeds it the real row, and lands on the pipeline's board for all 32 teams. The other thirty ratings did not move by a tenth of a point, because nothing else has been played.

Where that move sits historically is the part worth keeping. Against the 7,276 rating moves the engine has made since 1999, an 8.42-point move is a 13.5th-percentile night: 980 games moved a rating less, 6,296 moved one more. The median move in the file is 17.05, so Wednesday was very nearly half a normal game. Among the 428 week-1 games in the file only 57 moved a rating less. The reason is the one Tuesday's page gave in advance: the favourite won, and won narrowly, and the engine pays for surprise at the odds it quoted. Had New England won by the same three points the move would have been 17.77 the other way, a 53rd-percentile night, because 2.111 — the prior odds on Seattle — was the exchange rate between the two branches and it did not change.

A model's pick winning by three is the least informative decided result there is, and the update rule agrees: across the 938 games in the file where the model's pick won by three or fewer, the average move was 8.24 points. Wednesday was an ordinary member of that group.

The Exhibit

Left panel: two curves showing Seattle's rating change against the final margin of the opener, the Seattle-wins branch rising from about plus 4 at one point to plus 23 at 45 points and the New England-wins branch falling from minus 9 to minus 49, with a green dot on the Seattle branch at three points marking the realised move of plus 8.42 and a hollow red circle at minus 17.77 on the other branch at the same margin; a dotted vertical line at 24 points is labelled as the crossover that never came into play. Right panel: grouped bars for four pre-kickoff probabilities, the model at .679, the June market at .645, the closing market at .600 and records-only at .569, each with a green bar for its Brier score on the actual result, 0.1033, 0.1262, 0.1603 and 0.1861, and a grey bar for what it would have scored had New England won, 0.4604, 0.4157, 0.3596 and 0.3234, with a dashed line at the coin-flip score of .25.
Left: the two branches priced on September 8, with Wednesday's result placed on them. Right: the four probabilities on the record before kickoff, scored on what happened and on what did not. Data: nflverse game log and this site's frozen ledger, replayed through nfl_elo.py.

What Did Not Happen, and What Did

Tuesday's headline number was 24: the margin by which New England needed to win to pass Seattle at the top of the board. It was never in play. Seattle won, so the only question was how much further ahead it would get, and the answer was 16.8 points on the gap between the two teams: 81.8 before kickoff, 98.6 after, which is twice the move because both ratings travel. What that did to the rest of the board is small but real, and the companion page today works through it team by team:

  • Seattle's lead over second-place Buffalo grew from 58.7 to 67.1. The tier line the rankings page draws below #1 was already there; it is wider.
  • New England's hold on sixth shrank from 11.6 points over Philadelphia to 3.2. Sixth is still sixth. No team on the board changed rank.
  • The gap between fifth-place Houston and New England widened from 12.8 to 21.2.

Tuesday's page also said no week-2 pick would flip in any branch, and none did. New England's price at home to Pittsburgh on September 20 fell from .6723 to .6615; Seattle's price at Arizona rose from .7919 to .7998. The other fourteen week-2 games are untouched because none of their teams played. Those are the largest consequences a three-point win by the favourite can have under this rule, and they are about a point of probability each.

Grading Three Numbers, Honestly

The morning-of page put three probabilities side by side and said the game was well designed as a test because they disagreed. Here is how each scored, with one number added that was not available at the time of writing: the closing market, which the game file now carries at Seattle −3 with moneylines of +140 and −166. Those imply .4167 and .6241, a 4.07% hold, and a vig-free 60.0% for Seattle. Between the June snapshot and kickoff the market moved four and a half points toward New England. The model, which cannot move between games, did not.

Pre-kickoff numberP(Seattle)Brier, actualLog lossBrier had NE won
The model, frozen August 12.6785.1033.3878.4604
The market, June snapshot (+170 / −205).6447.1262.4389.4157
The market at the close (+140 / −166).5996.1603.5114.3596
Records only (two 14–3 teams plus home field).5686.1861.5645.3234

Read the third column and the model wins the night, by .0828 of Brier over the records-only view and by .0229 over the June market. Then read the last column and notice that the order is exactly reversed: had New England won, the model would have posted the worst score of the four, .4604, and the records-only view the best. That is not a coincidence and it is not a finding. All four numbers picked Seattle. The winner did not separate them at all; the only thing that separated them was how confident each was, and a proper scoring rule rewards confidence on a hit and punishes it on a miss by construction. One game grades the pick. It cannot grade the probability.

So what does "the model was right" mean here? Precisely this: it said Seattle was more likely to win, and Seattle won. The same sentence is true of the market and of the records-only view. The stronger claim — that .6785 was the correct amount of confidence — is the one that actually distinguishes a rating system from a coin with a lean, and it is only answerable in bulk. Across the sixteen graded seasons, games the model priced between .65 and .70 have gone 434 of 652, a realised rate of .6656 with a standard error of .0183. A .6785 call landing is exactly what that band has always done about two times in three. It is consistent with the model being calibrated. It is also consistent with the model being three points too confident, or three points too timid. One game has a standard error of .467 on that question; separating .679 from .629 at two standard errors takes about 349 games, which is roughly a season and a third of this model's whole ledger.

The one genuinely interesting fact in the table is the market's move. In June it sat between the model and the records-only view, closer to the model. By kickoff it had moved to 60.0%, roughly halfway back toward the records-only number, and the closing spread of 3 was a half point shorter than the 3.5 the morning page quoted. The market spent the summer learning things the model cannot see — rosters, camps, a quarterback's August — and everything it learned pointed toward New England. Then Seattle won by three. I do not know what to make of that yet, and neither does anyone else after one game; I note it because it will matter if it keeps happening.

The Spread Column

The morning page translated the model's edge into points using the exchange rate fitted for the new-coach study, 0.03552 points per Elo point, and got 129.8 × 0.03552 = 4.61, against the market's 3.5. Seattle won by 3. On the closing number of 3 the game pushed; on the 3.5 the morning page quoted, New England covered; and the model's 4.61 missed the margin by 1.61 points. The total of 44.5 was missed by 21.5 in a 13–10 game, which this model does not price and which no page here claimed to.

None of that is a grade either. The week-1 pricing study put the market's average opening-week error at 10.06 points; a closing number landing on the margin exactly has happened 189 times in the file's 6,967 lined games, one in 37, and ten times in 428 openers, and it says nothing about the number's quality. What the spread column does confirm is the direction of the disagreement I flagged on Tuesday: this site's rating was longer on Seattle than the market by about a point, the market got shorter still by kickoff, and the game landed inside both numbers.

The Miss Budget, One Game In

The miss budget is the frame this site grades itself against, so the accounting should be kept from the first game. Week 1's sixteen frozen probabilities expected 6.09 misses, of which the opener carried .3215 — the probability Seattle would lose. Seattle did not lose, so that .3215 has been converted from expected miss to realised hit, and the fifteen unplayed games still carry 5.77 expected misses between them. Equivalently the week expected 9.91 correct on Tuesday and, conditional on the first hit, expects 10.23 now. Those are the only numbers that changed. The season's 99.1 expected misses are now 98.8.

Two other pre-registered frames get nothing yet. The Week 10 Problem wrote down the bars a single week would have to clear to be surprising — a market error above 11.80 points or a model accuracy under .550 — and those are week-level statistics over sixteen games. After one game the market's week-1 error stands at 0.0 and the model's accuracy at 1.000, and I would ask you to draw no conclusion from either, because I am not going to. The learning-curve page pre-registered 7 to 13 correct as covering 93.6% of the week-1 distribution; one hit moves that window not at all.

What One Game Can and Cannot Tell You

It can tell you the arithmetic works, and it does; a rating system whose published update rule did not reproduce its own board would have a bigger problem than calibration. It can tell you the pick was right, which it was, and which is worth exactly one game of a 272-game ledger. It can tell you that the market moved toward the loser between June and kickoff, which is a fact about the market and not yet about anything else.

It cannot tell you whether 67.9% was the right number, whether New England's inherited 81.8-point deficit was wisdom or dead weight, or whether the rating's memory of 2023 and 2024 is worth carrying. The morning page's 328-game test said that a rating gap this size between two identical records has been worth about two wins in three; Wednesday was one of the two, and the test's standard error did not move. If Seattle goes on to win the division and New England goes 11–6, both descriptions of this game will look right. If New England wins the next six, the model will have been carrying dead weight and this page will have been the first place the evidence did not show up. The number goes up first and the story comes after, and after one game the story is not yet a story.

What This Page Does Not Show

Brier scores on one game rank confidence, not skill. The table above is a mechanical fact about a squared-error rule applied to a single binary outcome. The reverse ordering in the counterfactual column is the whole caution; I put it there so the actual column could not be read alone.

The closing line is the game file's, and it is a snapshot. The nflverse log updated the opener's row to a spread of 3 and moneylines of +140 and −166 once the game was final. I have no independent record of when that number was captured or whether it was the consensus close. The June numbers, +170 and −205 with a 3.5, are the ones the morning page quoted from the bundled schedule.

Nothing about the game itself is used. The file records the starting quarterbacks as Drake Maye and Sam Darnold and the head coaches as Mike Vrabel and Mike Macdonald, and those are the only facts about the football on this page, because the model sees the score and nothing else. A 13–10 game and a 34–31 game are the same three points to the rule.

The percentile scale is a replay. The 7,276 historical moves are recomputed by walking the engine forward through the file with the offseason regression taken out of the measurement, the same procedure Tuesday's page used, and they reproduce its median of 17.05 and maximum of 50.32.

Tonight's game is not on this page. San Francisco and the Rams play in Melbourne this evening; it is unplayed as of this writing, and the companion page prices it in advance the way Tuesday's page priced Wednesday.

Method and Sources

Three files: the nflverse game log at /data/games.csv, the frozen ledger at /data/predictions.json (which nfl_elo.py graded from Thursday morning's pull and which now carries the result on the opener's row), and a replay of explainer_src/nfl_elo.py, imported rather than copied. The harness is explainer_src/make_opener_graded_chart.py. The grading of a probability is three lines:

outcome = 1.0                                  # Seattle won
brier   = (p - outcome) ** 2                   # .6785 -> .1033 ; .6447 -> .1262 ; .5686 -> .1861
counter = (p - 0.0) ** 2                       # had New England won: .4604 ; .4157 ; .3234

The script asserts the opener's row from the file (score, margin, total, kickoff, venue, closing prices), the ledger's scoreboard and graded row, the replay's reproduction of all 32 frozen ratings and of the post-game board, the 8.416 move and its percentile against 7,276 historical moves, the 24-point crossover, both week-2 prices and the absence of pick flips, all four probabilities with their scores both ways, the spread and total arithmetic, the miss-budget accounting, and the .65–.70 calibration band: 125 assertions, all green as of September 10, 2026. Because the site's build republishes the June bundle over the served game file, the harness saves the sixteen week-1 rows as pulled on September 10 to a dated snapshot and reads from it thereafter, so it reproduces after any rebuild.

Sources: the nflverse public game log (games.csv, 1999–2025 results plus the 2026 schedule, with the opener's final score and closing prices as of the September 10 pull). The Brier score is Glenn Brier's, from Verification of Forecasts Expressed in Terms of Probability (1950).

Further reading

About the author

C. B. Zakarian

C. B. Zakarian is an independent analyst who writes about what he can measure. He builds every model, chart, and calculator on this site himself from the public nflverse play-by-play and game-log releases, shows the working, and never invents a number. The dataset behind the exhibits is served openly at /data/, and the method behind every figure is spelled out so you can check it against the same file. When the data can't answer a question, he says so.

More Explainers
Do Favorites Cover? Scoring by Week Most Common Scores Favorite Win Rates Playoff Football Division Games & Home Field DVOA EPA vs. DVOA CPOE Passer Rating vs. QBR Pythagorean Wins Air Yards & YAC Fourth-Down Analytics Strength of Schedule ANY/A RYOE Pass Protection Coverage Metrics Special Teams PROE & Game Script Red Zone Efficiency Explosive Plays Third Down Time of Possession Turnovers & Luck Win Probability YAC Over Expected Snaps & Usage Points Per Drive Success Rate Pressure Rate Play-Action Yards After Contact RPO Two-Point Conversions Yards per Route Run Block Win Rates Target Share & WOPR Home-Field Advantage Expected Points Point Spread Accuracy Weather & Scoring Rest & Scheduling Scoring Trend Overtime Over/Under Accuracy Key Numbers (3 & 7) Thursday & Primetime Grass vs. Turf One-Score Games Stadium Scoring Referee Effects QB Continuity Week 1 Signal Shutouts 2026 Schedule Strength 2026 Schedule Quirks Best Record vs. Super Bowl Win & Loss Streaks Division Repeats Close-Game Luck The Prediction Model The Week 1 Slate AFC East 2026 AFC North 2026 AFC South 2026 AFC West 2026 NFC East 2026 NFC North 2026 NFC South 2026 NFC West 2026 Preseason Signal 2026 Preseason 2026 Win Totals 2026 Playoff Odds 2026 International Games 2026 Miss Budget AFC vs NFC The 17-Game Era The Coach Ledger Week 1 Predictions Opening Night 2026 SB Rematch Effect Road Favorites The Chiefs' Rating Division Leverage The Learning Curve The Board, Sorted September, Priced The Offseason Haircut Fair Prices What One Game Moves The Shortest Lines The Week 10 Problem The New-Coach Bounce Same Record, Different Rating The Opener, Graded Rankings After the Opener All explainers

Go deeper

Want the code behind these metrics? Work through the 45-chapter NFL analytics tutorial.

Browse tutorials Free tools