Stat Explainer

What the First Meeting Tells You About the Second

Every within-season rematch in the game file is a division home-and-home — 1,347 of them since 1999 — which makes them the only matchup priced twice in the same season. The same team won both legs in 768 of the 1,340 decided pairs, .5731, where multiplying this site's own two probabilities together says .5497 and the closing market's say .5433 against an observed .5855. The excess lives almost entirely in quick rematches and blowouts, and reverses in December. The single-game price is fine: meeting-2 results do not lean against the spread with meeting-1's margin, at 0.81 standard errors. What is wrong is the independence assumption the site's own playoff simulation makes 272 times a season.

By C. B. Zakarian · Published September 18, 2026

The One Matchup This Site Prices Twice

Almost every pair of NFL teams meets once a season or not at all. A few meet twice, and in this site's game file the few are always the same few: 1,347 pairs from 1999 through 2025, every one of them a division home-and-home, one game at each ground. That makes them the only population where the same two teams are priced twice in the same year, which allows a question no single game can answer. Not “is each price right?” — but “are the two results independent?”

They are not, by a little. Of the 1,340 pairs in which both meetings were decided, 768 were swept by the same team — .5731. Multiplying this site's own two per-game probabilities together, as though the results were independent draws, says .5497. On the 883 pairs that carry a closing moneyline for both games, the market's two prices multiply to .5433 against an observed .5855.

My position is that this is a real property of the joint outcome and not a flaw in either price. The single-game number is fine: meeting-2 results do not lean against the closing spread with meeting-1's margin at all. What is wrong is the independence assumption — and this site makes that assumption 272 times a season, in the simulation that produces its playoff odds.

What Counts as a Rematch

The harness groups every regular-season game from 1999 to 2025 by season and by the unordered pair of teams, mapping the three relocations so a franchise stays itself. The result is clean: 4,273 pairs met once, 1,347 met twice, and none met three times. All 1,347 are division games, and in all 1,347 the venue flips between the two meetings. So “rematch” here means exactly one thing — the second leg of a division home-and-home — and the team that won the first is, by construction, playing the second at the other ground.

Seven pairs contained a tied game and are set aside, leaving 1,340 with a winner in each leg. Meetings are ordered by week, so meeting 1 and meeting 2 are never ambiguous. Across the 2,680 individual games the home team won .5485, which is the 54.8% the division-games page measured for division hosts generally — a useful sign that this population is not odd in some other way.

The Exhibit

Two panels, each with dark blue bars for the observed rate at which the winner of the first meeting also won the rematch, and a dashed red diamond line for what independence implies. Left panel, by weeks between the meetings: 63.3 per cent observed against 56.4 implied at one to four weeks on 275 pairs with a margin correlation of 0.31; 58.1 against 54.8 at five to eight weeks on 494 pairs, r 0.28; 55.7 against 54.4 at nine to twelve weeks on 366 pairs, r 0.18; and 50.2 against 54.4 at thirteen to seventeen weeks on 205 pairs, r 0.13, where the observed rate falls below the implied line. Right panel, by the margin of the first meeting: 52.1 against 54.4 for one to three points on 334 pairs; 55.5 against 53.9 for four to seven; 55.6 against 54.5 for eight to thirteen; 56.7 against 55.5 for fourteen to twenty; and 70.9 against 57.5 for twenty-one or more on 213 pairs.
Left: the sooner the rematch, the more the first game tells you. Right: only a blowout carries over. In both, the red line is what multiplying the two per-game probabilities together implies. Data: nflverse game log, 1999–2025, with a replay of nfl_elo.py.

The Excess Is Not Spread Evenly

Across all 1,340 pairs the two final margins correlate at .2344, which is the effect in a single number: related, and loosely. But a pooled 2.3 points of excess sweep rate would be a dull finding if it sat uniformly across the population, and it does not. Sorted by how many weeks separate the two meetings:

Weeks between meetingsPairsSweptIndependence impliesExcessMargin correlation
1–4 weeks2750.63270.5639+0.06880.310
5–8 weeks4940.58100.5484+0.03250.280
9–12 weeks3660.55740.5438+0.01360.178
13–17 weeks2050.50240.5444-0.04190.133

The excess runs from nearly seven points when the rematch comes inside a month to minus four points when it comes in December, and the correlation between the two margins falls from .3099 to .1334 over the same span. The independence baseline barely moves across the four windows, so the decay is in the outcomes rather than in the prices. A quick rematch is a genuine repeat; a rematch four months later is not one, and might be slightly less than one.

The same story sorted by how emphatic the first meeting was:

Margin of the first meetingPairsSweptIndependence impliesExcess
1 to 3 points3340.52100.5438-0.0228
4 to 7 points3120.55450.5390+0.0155
8 to 13 points2430.55560.5450+0.0106
14 to 20 points2380.56720.5546+0.0127
21 or more2130.70890.5747+0.1342

Nearly the whole effect lives in the bottom row. A team that won the first meeting by three touchdowns took the rematch .7089 of the time against an implied .5747, an excess of 13.4 points on 213 pairs. A team that won by a field goal or less took it .5210 of the time, below what independence implies. The middle three bins sit between +1.1 and +1.6 points and are indistinguishable from nothing. Whatever persists between two meetings of the same pair, a one-score game does not contain it. For scale at the other end, 135 pairs — 10.1% of the 1,340 — were swept with both legs decided by fourteen points or more.

One other cut is worth reporting and then discounting. When the host won meeting 1, the sweep rate is .5159; when the visitor won it, .6407. That looks dramatic and is mostly mechanical: because the venue flips, a team that won on the road in meeting 1 hosts meeting 2, so the 12.5-point gap is largely home-field advantage arriving in the second leg rather than anything about momentum. I would not read it as a separate finding.

How Sure Can Anyone Be

Two tests, and they do not agree about how strong this is. Averaged season by season, the engine's excess is +0.0239 with a season-clustered standard error of 0.0110, which is 2.18 from zero. On the market subset it is +0.0535 with a standard error of 0.0181, or 2.96. Those are the headline numbers and they look respectable.

The sign test is less impressed. The engine's excess was positive in 16 of 27 seasons, which a fair coin produces 22% of the time; the market's in 14 of 20, at 5.8%. Season-level excesses swing from −0.09 to +0.12. So the pooled estimate is carried by the size of the good seasons rather than by their number, and anyone reading this as a metronome would be reading it wrong. The two eras agree at least: +0.0279 from 1999 to 2011 and +0.0188 from 2012 to 2025, on 675 and 665 pairs.

The market subset deserves one caveat of its own. Closing moneylines are missing before 2006 and patchy until 2010, so those 883 pairs are a later, smaller sample than the engine's 1,340, and its larger excess may be partly that. What survives both readings is the direction and rough size: somewhere around two to four points of sweep rate more than independence implies, concentrated in quick rematches and blowouts.

The Single Game Is Priced Correctly

This is the part that keeps the finding honest, and it is a null. If the first meeting carried information the market failed to use, then meeting-2 results would lean against the closing spread with meeting-1's margin. Regress one on the other across all 1,347 pairs and the pooled slope is +0.0335 points of cover per point of first-meeting margin. Season by season it is +0.0186 with a standard error of 0.0229 — 0.81 from zero.

Banded, it is the same nothing: teams that lost the first meeting by fourteen or more finished 0.80 points under the number in the rematch, teams that won by fourteen or more 1.24 points over, and neither clears two of its own standard errors, nor do the two middle bands. A twenty-point first meeting is worth about a third of a point of the second against the number, which is not a bet.

So both things are true at once, and they are not in tension. Each game's price is right on its own. The two outcomes are still positively correlated, because both are driven by the same unobserved facts — how good these teams actually are, how the matchup suits them, who is healthy — and correlated outcomes are exactly what persistent hidden differences produce even when every marginal probability is perfectly calibrated. That is a statistical commonplace rather than a market failure. The week-2 page ran the general version of the single-game test after every week of the season and found the same null.

Worked Through: Green Bay and Minnesota, 2024

The cleanest illustration in the file is a pair the market could not separate. In 2024 Green Bay and Minnesota met in week 4 and again in week 17, and both closing prices were almost exactly even money:

market, vig-free, from Green Bay's side
  meeting 1 (week 4)                    p = .5514
  meeting 2 (week 17)                   p = .5043

if the two results were independent draws
  P(Green Bay sweeps)  = .5514 x .5043                    = .2781
  P(Minnesota sweeps)  = .4486 x .4957                    = .2224
  P(either sweeps)                                        = .5004     a coin flip

what happened
  Minnesota won meeting 1 by 2, and meeting 2 by 2        -> a sweep

Two one-score games, thirteen weeks apart, both to the same side. That single pair proves nothing; it is on the page because the arithmetic is the whole argument in miniature. The two prices were right — a team that wins two games by two points was not mispriced at even money — and multiplying them still understated how often the same team takes both. Note also that this pair sits in the thirteen-plus-week window and the one-to-three-point margin bin, the two cells where the population shows no excess at all. The average is not a rule for any one pair.

What It Costs This Site

The reason to care is that this site's playoff simulation draws all 272 games independently. The odds page measured its calibration slope at 0.694 and yesterday's repair page fixed the levels with a 120-point rating-uncertainty knob. Independence between division rematches is one specific, measurable part of that overconfidence, and it can be priced rather than assumed.

Link the two games of each division pair with a latent correlation and choose it to reproduce the historical sweep rate. The answer is 0.0945, a small number. Re-running the 2026 simulation with the 43 pairs that still have both games unplayed linked at that value:

sweep rate among the linked pairs   .5482  ->  .5727      (history: .5731)
mean distance of a division title from 25%    19.438  ->  19.305
division favourites (above 50%)         −0.30 of a point on average
division longshots (below 10%)          +0.10
total movement across all 32 teams              4.6 points

It goes the right way and it is small. Linking the pairs widens the average team's season win total by about 0.02 of a win and pulls division odds toward the middle, which is the direction the calibration work says they need to go. But four and a half points of division title spread over 32 teams is not a repair; it is one named contributor to a gap that the 120-point knob moved by an order of magnitude more. I am reporting it as an identified mechanism, not proposing a change to the shipped simulation.

The 2026 schedule carries 96 division games in 48 home-and-home pairs, five of which have already played their first leg. In four of those five the host won, so the rematch is at the loser's ground — the .5159 group. The exception is San Francisco, which won in Melbourne and hosts the return in week 14, landing in the .6407 group. On this page's own reading, that grouping is mostly a statement about home field, and the rematch dates are late enough that the excess has largely decayed anyway.

What This Page Does Not Show

It is not a betting edge. The single-game test is a null at 0.81 standard errors, and the correlation finding is about the joint distribution of two results. It would matter to a parlay or to a season simulation; it does not say either game's price is wrong.

Correlation is what hidden differences produce. Perfectly calibrated marginals plus persistent unobserved team quality generate exactly this pattern. Nothing here identifies a cause, and “revenge”, “familiarity” and “one team is simply better than the file knows” all fit the same numbers.

Many cuts. Four time windows, five margin bins, two eras, a venue split and two forecasters. The claims I would defend are the pooled excess and its season-clustered standard error; the bin-level cells, especially the 205-pair and 213-pair ones, are descriptive.

The two tests disagree in strength. The clustered estimate reaches 2.18 and 2.96; the sign test reaches 22% and 5.8%. Both are reported above because publishing only the flattering one would misdescribe the evidence.

Division pairs are not a random sample of games. They are the matchups the schedule repeats, between teams that know each other and often finish near each other. Whatever is measured here belongs to that population and should not be extended to rematches that do not exist in the file.

The moneyline subset is later and smaller. 883 pairs against 1,340, concentrated from 2010 on, which is one plausible reason its excess is larger.

The latent correlation is fitted in hindsight, on the same 1,340 pairs it is then applied to, and a Gaussian copula is an assumption about shape that the data cannot confirm. The 2026 figures are an illustration of magnitude, not a forecast.

Method and Sources

Two files and one module. The June 2026 nflverse bundle (/data/games.csv, served at /data/games.csv): 7,276 played games from 1999 to 2025, 6,967 of them regular season, with closing spreads and, from 2006, two-way moneylines; its 272 2026 rows are schedule only, so no 2026 result reaches a historical figure here. The 2026 board is the dated snapshot explainer_src/_2026_week2_snapshot_2026-09-18.json, because the build republishes the June bundle over the served copy. And explainer_src/nfl_elo.py, imported rather than copied, fed every played game in file order exactly as its own backtest does. The harness is explainer_src/make_division_rematch_chart.py. The core of it:

pairs = {(season, frozenset(teams)): [meeting1, meeting2]}     # 1,347, all division
swept = (margin1 > 0) == (margin2 > 0)                          # 768 of 1,340 = .5731
implied = mean(p1 * p2 + (1 - p1) * (1 - p2))                   # .5497 engine, .5433 market
# the single game: does meeting 2 beat its number with meeting 1's margin?
ols(result_vs_spread2 ~ margin1)   -> +0.0186 per point, season-clustered SE 0.0229

The script asserts the engine constants, the bundle's row counts and the absence of any 2026 score, the pair census and that all 1,347 are division home-and-homes with a flipped venue, the sweep and split counts, both independence baselines, both season-clustered standard errors and both sign tests, all four time windows and all five margin bins cell by cell with their baselines, the monotone decay and the sign reversal at each end, the venue split, the home-win rate against the division-games page, the era split, the regression against the spread with its bands, the copula calibration, the 2026 simulation before and after linking, the five 2026 pairs that have played a first leg, the worked example down to its moneylines, and this page's own figures: 68 assertions, all green as of September 18, 2026.

Sources: the nflverse public game log (games.csv), including its closing spreads and moneylines. The rating method is Arpad Elo's. The dependence structure is modelled with a Gaussian copula, after Abe Sklar's theorem (1959).

Further reading

About the author

C. B. Zakarian

C. B. Zakarian is an independent analyst who writes about what he can measure. He builds every model, chart, and calculator on this site himself from the public nflverse play-by-play and game-log releases, shows the working, and never invents a number. The dataset behind the exhibits is served openly at /data/, and the method behind every figure is spelled out so you can check it against the same file. When the data can't answer a question, he says so.

More Explainers
Do Favorites Cover? Scoring by Week Most Common Scores Favorite Win Rates Playoff Football Division Games & Home Field DVOA EPA vs. DVOA CPOE Passer Rating vs. QBR Pythagorean Wins Air Yards & YAC Fourth-Down Analytics Strength of Schedule ANY/A RYOE Pass Protection Coverage Metrics Special Teams PROE & Game Script Red Zone Efficiency Explosive Plays Third Down Time of Possession Turnovers & Luck Win Probability YAC Over Expected Snaps & Usage Points Per Drive Success Rate Pressure Rate Play-Action Yards After Contact RPO Two-Point Conversions Yards per Route Run Block Win Rates Target Share & WOPR Home-Field Advantage Expected Points Point Spread Accuracy Weather & Scoring Rest & Scheduling Scoring Trend Overtime Over/Under Accuracy Key Numbers (3 & 7) Thursday & Primetime Grass vs. Turf One-Score Games Stadium Scoring Referee Effects QB Continuity Week 1 Signal Shutouts 2026 Schedule Strength 2026 Schedule Quirks Best Record vs. Super Bowl Win & Loss Streaks Division Repeats Close-Game Luck The Prediction Model The Week 1 Slate AFC East 2026 AFC North 2026 AFC South 2026 AFC West 2026 NFC East 2026 NFC North 2026 NFC South 2026 NFC West 2026 Preseason Signal 2026 Preseason 2026 Win Totals 2026 Playoff Odds 2026 International Games 2026 Miss Budget AFC vs NFC The 17-Game Era The Coach Ledger Week 1 Predictions Opening Night 2026 SB Rematch Effect Road Favorites The Chiefs' Rating Division Leverage The Learning Curve The Board, Sorted September, Priced The Offseason Haircut Fair Prices What One Game Moves The Shortest Lines The Week 10 Problem The New-Coach Bounce Same Record, Different Rating The Opener, Graded Rankings After the Opener The First Miss, Graded What Sunday Can Do SF and LA, Re-Priced No Good Years, Only Lucky Ones Week 2 Overreaction Sunday, Graded Week 1 Scoring, 2026 What 0-2 Costs Monday Night, Graded Playoff Odds After Week 1 Calibrated Playoff Odds Thursday: DET at BUF Game-Pick Calibration The QB Blind Spot DET at BUF, Graded The Second Meeting All explainers

Go deeper

Want the code behind these metrics? Work through the 45-chapter NFL analytics tutorial.

Browse tutorials Free tools