the model

7 standards, written down before a single result was looked at. 5 can be settled today, and this poll clears 0.

  • the bar, fixed before this poll ran
  • this poll · FAIL
  • rating systems somebody else publishes
  • plain baselines, not rating systems
  • earlier versions of this model

  1. Call more games right than anybody else does.

    Straight-up accuracy at or above the floor

    the bar 70.0%
    Colley 68.0%Elo 68.8%Home team wins 56.3%L1 efficiency 69.2%L2 results 69.5%L3 Power 68.7%Random walker 65.1%Résumé 68.7%SRS / Massey 69.1%Win percentage 66.1%
    68.7% FAIL
    54.6%further right is better71.7%
    • along the line, low to high
    • the bar70.0%
  2. Miss the final score by fewer points.

    Mean absolute error at or below the ceiling

    the bar 12.8
    Colley 13.56Elo 13.56Home team wins 15.46L1 efficiency 13.16L2 results 13.23L3 Power 13.04Random walker 14.01Résumé 13.04SRS / Massey 13.20Win percentage 13.82
    13.04 FAIL
    12.5further left is better15.8
    • along the line, low to high
    • the bar12.8
  3. And keep the blowout misses rare.

    Root mean squared error at or below the ceiling

    the bar 15.8
    Colley 17.24Elo 17.10Home team wins 19.86L1 efficiency 16.68L2 results 16.74L3 Power 16.53Random walker 17.98Résumé 16.53SRS / Massey 16.70Win percentage 17.66
    16.53 FAIL
    15.3further left is better20.4
    • along the line, low to high
    • the bar15.8
  4. When it says 70%, be right about seven times in ten.

    Worst decile calibration deviation within tolerance

    the bar 5.0
    7.37 FAIL
    4.7further left is betterno other system publishes this one, so there is nothing to compare7.7
    • along the line, low to high
    • the bar5.0
  5. Put the loser above the winner less often than anybody else does.

    Retrodictive violations at or below every baseline

    the bar is every other system
    Colley 19.8%Elo 22.0%L1 efficiency 24.3%L2 results 21.9%L3 Power 22.8%Random walker 20.2%Résumé 20.2%SRS / Massey 22.1%Win percentage 18.3%
    20.2% FAIL
    17.6%further left is better25.1%
    • along the line, low to high

Every bar sits above the best number any system scored here puts up.

Failing a bar set at somebody else’s best number does not make this ranking worse than the ones you already read. It means it is not better yet, and closing that gap is what the season is for.

2 more standards were written down and one season cannot settle them. They count as neither a pass nor a fail.

  • Brier score beats every baseline
  • Retro-vs-live divergence declines monotonically

Every system, every measure, every number.

To rank week N, every system here is fitted through week N−1 and nothing else.

systemnSU%MAERMSEBrierloglossviol%
This poll (schedule odds)158568.71%13.03816.5310.19710.577220.19%
Colley158567.95%13.56217.2370.20630.598919.76%
Elo158568.77%13.55917.0980.20140.588522.03%
L1 efficiency158569.15%13.16516.6840.19950.582624.34%
L2 results core158569.53%13.22716.7400.19750.579421.86%
L3 Power (efficiency + results blend)158568.71%13.03816.5310.19710.577222.80%
Random walker158565.05%14.01517.9840.21780.624420.23%
Résumé (the ordering it replaced)158568.71%13.03816.5310.19710.577220.15%
SRS / Massey158569.09%13.19616.6950.19950.582122.11%
Win percentage158566.12%13.82417.6600.21180.611818.31%
Home team always wins158556.34%15.45819.8630.24750.6885not published

n games scored · SU% winners called right · MAE average miss in points · RMSE the same with blowout misses weighted heavier · Brier and logloss how honest the stated chances were · viol% how often the loser finished above the winner

The poll and the résumé predict through the same power rating, so only the last column tells them apart, and on that column the résumé is the better of the two.

What each of these systems actually is →

Two teams, the same record, and one of them was much harder.

Every score in the country sets how good each team is. Then the model replays your schedule against those ratings, over and over, and counts how often an ordinary team survives it.

Georgia12-1 · ranked 3

1 in 86

James Madison12-1 · ranked 14

1 in 7

Replay Georgia’s schedule and an ordinary team has that season about once in 86 tries. Replay James Madison’s and it happens about once in 7. That is what the poll sorts on.

Your own margin never lifts your own ranking. Your opponents’ margins are the whole of what your wins are worth.

Why it re-ranks the past every week

0.06.112.218.316.84 places0.91147101316week of the season

A win in week 2 is worth whatever that team turns out to be, and you find out in November. How far each week’s answer moved once the model knew, over all 136 teams. It is supposed to fall like that.

See week 5 of 2025 both ways, every team →

The wall between the two rankings

The Poll reads

  • who played who
  • who won
  • the final score
  • where it was played
  • every play of every game

The Projection reads

  • last season's final ratings, off the poll
  • returning production
  • the transfer portal
  • a head-coach change

The PollThe Projectionand never the other way

If a single one of the projection's inputs turns up anywhere near the poll's math, the build fails and names it.

The projection is graded from week 5 of the season it forecast, by a ranking never allowed to see it. Checked by a machine, not remembered by a person.

How the August projection gets built

  1. 1 Last season's rating

    How good the model had them by the end of last year.

    moved 136 of 136 teams · priced about right

  2. 2 Returning production

    The share of last year's offense that came back.

    moved 134 of 136 teams · priced about right

  3. 3 The transfer portal

    Players in against players out.

    moved 136 of 136 teams · priced about right

  4. 4 A head-coach change

    A new head coach, and only that.

    moved 32 of 136 teams · priced about right

and what it threw out

  • The clean correction for a promoted team made the predictions worse, so the model kept the smaller number those games measured.
  • A setting that won on its own test broke a second one when the two ran together, so it stayed out.

Four ingredients, in the order the model adds them, with what the grading found when it scored each against the season that followed. All four came back dull, which is the answer nobody would have made up.

What one season of grading can and cannot settle about a recipe.the pipeline's own words

Across the 136 teams the poll ranked, every one of the four terms came back priced about right. The furthest from zero was last season's rating, at 1.4 standard errors, and the data cannot tell that apart from the value the recipe already uses. The season did not ask for a different coefficient.

The league-wide attribution is a regression of projection error on each term's contribution, over about 134 teams and four correlated terms. One season is one data point about the recipe. It is suggestive and it is not a verdict.

The three rules the model was built inside, and why it got no opinions →

What the line under each team on the weekly poll means.
GeorgiaRanked 3. The model replayed this schedule 1,000 times, and this team finished between 2 and 44 of 136 in 90% of them.rank 3 · lands between 2 and 44
James MadisonRanked 14. The model replayed this schedule 1,000 times, and this team finished between 3 and 57 of 136 in 90% of them.rank 14 · lands between 3 and 57

The model replays the whole schedule 1,000 times, with every result redrawn, and re-ranks all 136 teams it ranked in 2025 each time. The line is where that team kept landing.

The real range is wider than the line, because the replays keep each team’s efficiency fixed instead of re-running every play. The methodology page explains that limit.

Rebuild every number.

One command on The Data does it, over the same published files, with no clone and no account.

run 61a5fd2c · published 2026-08-19 23:18:03 UTC · code e215160 · config 1ef7cf23…
q_ref 15.64 (Arizona) · β_w 7 · C 32 · h 3.706 · σ 15.884 · λ₁ 200 · λ₂ 0.5 · k 72.96 · w₁ 0.5184 · w₂ 0.3169