the model
7 standards, written down before a single result was looked at. 5 can be settled today, and this poll clears 0.
Every bar sits above the best number any system scored here puts up.
Failing a bar set at somebody else’s best number does not make this ranking worse than the ones you already read. It means it is not better yet, and closing that gap is what the season is for.
2 more standards were written down and one season cannot settle them. They count as neither a pass nor a fail.
- Brier score beats every baseline
- Retro-vs-live divergence declines monotonically
Every system, every measure, every number.
To rank week N, every system here is fitted through week N−1 and nothing else.
| system | n | SU% | MAE | RMSE | Brier | logloss | viol% |
|---|---|---|---|---|---|---|---|
| This poll (schedule odds) | 1585 | 68.71% | 13.038 | 16.531 | 0.1971 | 0.5772 | 20.19% |
| Colley | 1585 | 67.95% | 13.562 | 17.237 | 0.2063 | 0.5989 | 19.76% |
| Elo | 1585 | 68.77% | 13.559 | 17.098 | 0.2014 | 0.5885 | 22.03% |
| L1 efficiency | 1585 | 69.15% | 13.165 | 16.684 | 0.1995 | 0.5826 | 24.34% |
| L2 results core | 1585 | 69.53% | 13.227 | 16.740 | 0.1975 | 0.5794 | 21.86% |
| L3 Power (efficiency + results blend) | 1585 | 68.71% | 13.038 | 16.531 | 0.1971 | 0.5772 | 22.80% |
| Random walker | 1585 | 65.05% | 14.015 | 17.984 | 0.2178 | 0.6244 | 20.23% |
| Résumé (the ordering it replaced) | 1585 | 68.71% | 13.038 | 16.531 | 0.1971 | 0.5772 | 20.15% |
| SRS / Massey | 1585 | 69.09% | 13.196 | 16.695 | 0.1995 | 0.5821 | 22.11% |
| Win percentage | 1585 | 66.12% | 13.824 | 17.660 | 0.2118 | 0.6118 | 18.31% |
| Home team always wins | 1585 | 56.34% | 15.458 | 19.863 | 0.2475 | 0.6885 | not published |
n games scored · SU% winners called right · MAE average miss in points · RMSE the same with blowout misses weighted heavier · Brier and logloss how honest the stated chances were · viol% how often the loser finished above the winner
The poll and the résumé predict through the same power rating, so only the last column tells them apart, and on that column the résumé is the better of the two.
Two teams, the same record, and one of them was much harder.
Every score in the country sets how good each team is. Then the model replays your schedule against those ratings, over and over, and counts how often an ordinary team survives it.
Georgia
1 in 86
James Madison
1 in 7
Replay Georgia’s schedule and an ordinary team has that season about once in 86 tries. Replay James Madison’s and it happens about once in 7. That is what the poll sorts on.
Your own margin never lifts your own ranking. Your opponents’ margins are the whole of what your wins are worth.
Why it re-ranks the past every week
A win in week 2 is worth whatever that team turns out to be, and you find out in November. How far each week’s answer moved once the model knew, over all 136 teams. It is supposed to fall like that.
The wall between the two rankings
The Poll reads
- who played who
- who won
- the final score
- where it was played
- every play of every game
The Projection reads
- last season's final ratings, off the poll
- returning production
- the transfer portal
- a head-coach change
The PollThe Projectionand never the other way
If a single one of the projection's inputs turns up anywhere near the poll's math, the build fails and names it.
The projection is graded from week 5 of the season it forecast, by a ranking never allowed to see it. Checked by a machine, not remembered by a person.
How the August projection gets built
1 Last season's rating
How good the model had them by the end of last year.
moved 136 of 136 teams · priced about right
2 Returning production
The share of last year's offense that came back.
moved 134 of 136 teams · priced about right
3 The transfer portal
Players in against players out.
moved 136 of 136 teams · priced about right
4 A head-coach change
A new head coach, and only that.
moved 32 of 136 teams · priced about right
and what it threw out
- The clean correction for a promoted team made the predictions worse, so the model kept the smaller number those games measured.
- A setting that won on its own test broke a second one when the two ran together, so it stayed out.
Four ingredients, in the order the model adds them, with what the grading found when it scored each against the season that followed. All four came back dull, which is the answer nobody would have made up.
What one season of grading can and cannot settle about a recipe.the pipeline's own words
Across the 136 teams the poll ranked, every one of the four terms came back priced about right. The furthest from zero was last season's rating, at 1.4 standard errors, and the data cannot tell that apart from the value the recipe already uses. The season did not ask for a different coefficient.
The league-wide attribution is a regression of projection error on each term's contribution, over about 134 teams and four correlated terms. One season is one data point about the recipe. It is suggestive and it is not a verdict.
The three rules the model was built inside, and why it got no opinions →
What the line under each team on the weekly poll means.
The model replays the whole schedule 1,000 times, with every result redrawn, and re-ranks all 136 teams it ranked in 2025 each time. The line is where that team kept landing.
The real range is wider than the line, because the replays keep each team’s efficiency fixed instead of re-running every play. The methodology page explains that limit.
Rebuild every number.
One command on The Data does it, over the same published files, with no clone and no account.
- The Data Every file, every hash, and the one command.
- Move a setting A table cannot tell you which of these decides anything. Moving one can.
- Every constant, and where this is weak The settings this run used, and the decision records that say where it falls short.