This is the page with everything on it: every number the model runs on, how it scored against everybody else, and the places it falls short. You do not need it to read the poll. It is here so that nothing about the poll is hidden.
This whole page answers one question from the front of the site: how do you put a number on how hard a season was?
Teams are ranked by −log₁₀ P(W ≥ W_t): how improbable it is that a team of reference quality q_ref would have gone at least this well against this exact schedule.
The stack
The path from a play to a rank, with every constant that enters along the way.
Never enters the poll. Connected to nothing above.
no preseason ranking
no human poll
no committee
no conference
no market lines
no recruiting rankings
no third party's model output
nothing carried over from last season
no logos
The layers, top to bottom, in the order the data moves through them. Each band carries the published constants that enter at it, and every one of them is in the table further down. The bottom band is the ranking.
Margin lives at L2 and L3 and never reaches the rank key, which is the -log10(P(W >= W_t)) under a Poisson-binomial over the exact schedule; margin never enters string in the table below.
The claim has two halves and both are enforced. Your own margin never enters your own rank: the schedule flattener carries opponent power, site and a boolean win, with no margin column to leak from, and a test scrambles every final score while preserving every winner and asserts the ranking is bit-identical. Your opponents’ margins set what your wins are worth: opponent quality is the L3 power rating, fitted on compressed scoring margin, so a second test refits power from those same scrambled scores and asserts the ranking moves. The wider claim, “margin never enters” full stop, was true of the module and false of the poll. An outside review said so first, and both tests exist so it cannot quietly widen again.
Before each fit the harness rebuilds every design matrix from the frame restricted to the allow-list and requires bit-identity, so an input nobody thought of fails closed.
The publication gate
The bar the build-your-own page mentions, written down before any result was seen.
Publication gateFAILEvery decided criterion passesUndecided criteria are reported as undecided, never as passes: brier_beats_all_baselines, retro_vs_live_monotone
FAILStraight-up accuracy at or above the floorsu_accuracybar 0.7 · actual 0.6871
FAILMean absolute error at or below the ceilingmaebar 12.8 · actual 13.04
FAILRoot mean squared error at or below the ceilingrmsebar 15.8 · actual 16.53
FAILWorst decile calibration deviation within tolerancecalibrationbar 5.0 · actual 7.37
FAILRetrodictive violations at or below every baselineviolations_vs_baselinesobserved 0.2019
not yet decidedBrier score beats every baselinebrier_beats_all_baselinesno constant threshold
not yet decidedRetro-vs-live divergence declines monotonicallyretro_vs_live_monotoneno constant threshold
Every threshold above was fixed in the config before anyone had seen a result. I set the bar above every public rating system I could score. Nothing clears it, including this one.
The verdict comes from the published status rather than from the picture, and undecided is never dressed as a pass.
Against every baseline
Strict walk-forward: to predict week N the harness fits through week N−1 of the same season and nothing else. The home-team-always-wins floor is always in the table, because a table without its floor is not a table.
system
n
SU%
MAE
RMSE
Brier
logloss
viol%
This poll (schedule odds)
1585
68.71%
13.038
16.531
0.1971
0.5772
20.19%
Colley
1585
67.95%
13.562
17.237
0.2063
0.5989
19.76%
Elo
1585
68.77%
13.559
17.098
0.2014
0.5885
22.03%
L1 efficiency
1585
69.15%
13.165
16.684
0.1995
0.5826
24.34%
L2 results core
1585
69.53%
13.227
16.740
0.1975
0.5794
21.86%
L3 Power (efficiency + results blend)
1585
68.71%
13.038
16.531
0.1971
0.5772
22.80%
Random walker
1585
65.05%
14.015
17.984
0.2178
0.6244
20.23%
Résumé (the ordering it replaced)
1585
68.71%
13.038
16.531
0.1971
0.5772
20.15%
SRS / Massey
1585
69.09%
13.196
16.695
0.1995
0.5821
22.11%
Win percentage
1585
66.12%
13.824
17.660
0.2118
0.6118
18.31%
Home team always wins
1585
56.34%
15.458
19.863
0.2475
0.6885
not published
n games scored · SU% winners called right · MAE average miss in points · RMSE the same with blowout misses weighted heavier · Brier and logloss how honest the stated chances were · viol% how often the loser finished above the winner
The poll and the résumé predict through the same power rating, so only the last column tells them apart, and on that column the résumé is the better of the two.
Ranking teams purely by win percentage beats this model on one of these measures. That is the price of caring who you played: a system that ignores schedule will always look tidy on a test that also ignores schedule.
The divergence curve
How far the poll’s answer for a given week moved once the season finished, drawn over every team it ranked.
The gap between what the poll published in week N and what it says about that week after the season, over all 136 ranked teams, weeks 1 through 4 included even though the poll never published them. If this line ever stops declining, the weekly poll was never worth reading.
Every constant this run used
Every one of these was published the same week the poll was. There is no second file with the real settings in it.
constant
value
C
32
beta_w
7
bisection_max_iter
60
blend_n_games
714
bootstrap_draws
1000
bootstrap_jobs_requested
4
bootstrap_jobs_used
1
bootstrap_lambda
0.5
bootstrap_n_games
1637
bootstrap_n_ranked_teams
136
bootstrap_seed
20260812
bootstrap_sigma
15.884317060721084
cv_folds
5
gauss_hermite_nodes
20
h_points
3.706281732877181
home_field
4.041520630398333
home_field_home_and_home
4
interval_level
0.9
k_points_per_unit
72.96415994533689
lambda
0.5
lambda_l1
200
lambda_l2
0.5
max_games_one_team
15
min_out_of_sample_games
40
min_tail
1e-300
n_games
1637
n_plays
212826
n_record_games
1637
n_resume_games
1637
n_resume_teams
315
n_saturated_high
3
n_saturated_low
50
n_teams
315
power_home_field_points
3.706281732877181
power_scale_b
1
power_scale_n_games
714
power_se_median_points
1.3981064284088478
power_se_scale
0.3168928694216252
power_sigma
15.884317060721084
power_sigma_n_out_of_sample_games
77
q_ref
15.64109836328906
q_ref_pool_size
136
recency_gamma
1
season
2025
seed
20260812
sigma
15.884317060721084
w1_efficiency
0.5184474738298377
w2_results
0.3168928694216252
weight_bowl_non_cfp
0.25
weight_cfp
1
weight_conference_championship
1
weight_regular_season
1
The named layers and orderings (36 strings)
blend_weight_source
out_of_sample
bootstrap_jobs_note
draws run sequentially; per-draw streams come from SeedSequence.spawn, so parallelising later cannot move a published number (report 03 §9.3 item 2)
bootstrap_lambda_note
held at the value the real data's CV selected; the bootstrap propagates sampling uncertainty at a fixed hyperparameter
bootstrap_note
parametric on the FIXED schedule: outcomes redrawn from the fitted model, refit, re-ranked. Games are edges in the schedule graph and are not exchangeable, so resampling them with replacement - which report 02 §3.3's parenthetical specified - is invalid here; see naive_resample_diagnostic. The L1 efficiency half of Power is held at its point estimate because plays are not resimulated, so these intervals are a LOWER BOUND on total uncertainty.
bootstrap_power_source
L3
bootstrap_scheme
parametric_on_fixed_schedule
bootstrap_sigma_source
walk_forward_residuals
companion_layer
L3_power
cv_group
game_id
estimator
ols
fit_out_of_sample_only
true
headline_decided
2026-08-12, docs/adr/0005-headline-ordering.md
headline_layer
C_schedule_odds
headline_ordering
schedule_odds
hindsight_data_bucket
2025-post-w01
hindsight_is_live
false
hindsight_variant
A_frozen_form
layer
C schedule odds
power_scale_universe
blend_regression_on_actual_margin
power_se_note
ridge sandwich on the L2 half only (report 02 §3.3), carried onto the points scale by w2. THE EFFICIENCY HALF IS HELD AT ITS POINT ESTIMATE: a play-level covariance over ~2,000 coefficients and 170,000 correlated rows is a different object and is not built, so this is a LOWER BOUND on the uncertainty of an L3 rating. The parametric bootstrap inherits the same limitation and says so
power_sigma_source
walk_forward_residuals
power_source
L3
power_version
v1
provisional
false
q_ref_method
power_rank_25
q_ref_team
Arizona
ranking_key
-log10(P(W >= W_t)) under a Poisson-binomial over the exact schedule; margin never enters
raw wins, and raw compressed margin; game weights shape the Power fit, not the accomplishment
resume_version
v0
saturation_tiebreak
margin
sigma_source
walk_forward_residuals
tail_method
exact_poisson_binomial_dp
tie_break
mid_p
version
v0
Where this is weak
The arguments against every claim on the pages above, verbatim from the decision records, including the ones that record being wrong. Nothing here was softened for the website.
Where this decision is weak
docs/adr/0005-headline-ordering.md
Stated so it is not read as stronger than it is, and reproduced from study §10.4:
- **Four seasons.** The forward-accuracy gap between B and the other two (2 pp on
9,433 games) is comfortably significant. **The gap between A and C on violations
(0.1 to 0.5 pp) is not, and nothing here claims it is.** A and C are tied on
that axis, and what decided between them was the structural finding.
- The postseason axis is unusable at current sample sizes (14 CFP games, 4 non-CFP
NY6 games) and 2021-2022 cannot contribute to it at all, because the archive
carries no postseason rows for those seasons.
- 2024 is the validation season. Every conclusion holds on it with the same sign,
but it was consulted once, in one pass, and must not be consulted again for a
re-tune.
- 2025 is untouched and stays locked.
- **No bootstrap intervals anywhere.** `model/bootstrap.py` is still a stub, so the
study reports point estimates and sample sizes and leaves the reader to judge.
Why an unbeaten team can finish behind a one-loss team
The price of C, stated plainly
docs/adr/0005-headline-ordering.md
1. **An unbeaten team can finish behind a one-loss team, and that will require
explaining every year.** It is the direct consequence of the promise. The
explanation is on the page: the tail probability, the reference team it was
measured against by name, and the Power column.
2. **One published constant that A did not have.** `q_ref` is the Power rating of
the 25th-ranked Power team that week, the least flattering defensible reading of
ESPN's "average Top-25 team", and a single team that can be *named* each week.
Study §9 measured the ordering's sensitivity to it rather than asserting it was
fine: across a 16-point swing in reference quality, Kendall's τ never fell below
0.985, the mean rank change never reached one place, and at most one team
entered or left the top 25. 2023 Liberty spans #8 to #12 across that whole swing
and never reaches Georgia at #7 under any choice. The probability *values* move
by orders of magnitude, which is exactly why the value is published rather than
the rank alone.
3. **C loses forward ordering accuracy to B by about 2 points.** Accepted, and for
a stated reason: forward accuracy is a prediction metric, and the headline poll
is not the instrument this project ships for prediction. L3 Power is, and it
beats all three orderings on that axis.
C also inherits the invariance that makes the résumé's zero point harmless: shift
every Power rating by a constant and every rank-derived `q_ref` shifts with it, so
no probability moves at all. Only the `fixed` method breaks that, which is why it
is not the default.
Whether tuning the constants could ever have closed the gate
The two uncomfortable results
docs/adr/0007-tuned-constants.md
### 1. The optimum is a corner solution
`C = 32` is the **top of `c_grid`**. The search did not bracket the optimum; the
data wants to keep going and the published grid stops it. The protocol fixed the
search space as *exactly the config grids* precisely so this campaign could not
widen the net after seeing the numbers, so the boundary is reported rather than
crossed.
It also says something about where the bounds came from. The review's §7 table
lists C's range as Pasteur's cap of 21 and the CFBD SRS walkthrough's ±28. Both
are other people's answers on other people's datasets, and the fitted value is
above both.
**Widening `c_grid` is the first item for the next campaign, and it must be
pre-registered before it is searched.**
### 2. Tuning cannot close the gate, and it is not close
The whole 416-cell factorial spans **0.135 points of MAE**, best cell to worst.
The winner beats the incumbent by **0.0086**. The gate needs **0.219** more than
the incumbent has, and the best cell in the entire searched space is still
**0.210 points above the threshold**.
**The gate gap is built into the design. No amount of tuning closes it.** That is
the most important sentence in this ADR and it is a negative result. Every constant in the pre-
registered space was searched at full resolution and the answer is that these
constants were never what stood between this system and its own thresholds.
The calibration diagnosis: diagnosed, and deliberately unfixed
docs/adr/0007-tuned-constants.md
Four candidate fixes were declared before any was run, with an adoption bar of
**≥ 2.0 pp on tune AND direction holding on 2024**. **None clears it, and the
config does not move on account of the calibration criterion.**
| Candidate | Tune Δ | Clears 2 pp? | 2024 Δ | Verdict |
|---|---:|---|---:|---|
| Student-t margins, at the fitted ν = 92.7 | +2.50 pp | yes | −0.38 pp | direction reverses; not adopted |
| Heteroscedastic σ(\|m̂\|) | −3.22 pp | no | +0.81 pp | fails tune; not adopted |
| Home-and-home `h` | — | — | — | **not runnable** under constraint 2 |
| Favourite-longshot | — | — | — | **the diagnosis** itself, and there is no knob for it |
**Neither named suspect did it.**
- **The normal tail is eliminated.** Maximum likelihood on the walk-forward
residuals puts ν at **92.7**; skew −0.098, excess kurtosis +0.020, Jarque-Bera
p = 0.275. These residuals are not distinguishable from normal. Low-ν rows in
the sweep *do* cut the deviation, and that is a clue rather than an exoneration:
a t with a matched second moment and small ν has a *narrower body*, so what
those rows buy is sharpness rather than tail weight.
- **The single home-field constant is eliminated as the cause.** The residual
mean is −1.13 points at home sites against −0.19 at neutral, an order of
magnitude too small to make a 13.67 pp decile. It is **separately convicted of
something else.** The site coefficient the harness uses averages **6.39 points
with a standard deviation of 4.34** across published weeks, against **1.88 ±
0.34** from 1,113 home-and-home pairs. Only 37 of 1,585 scored games are at
neutral sites, so the intercept and the site term are nearly collinear and `h`
is barely identified. Real defect. Different defect.
**The cause is under-dispersion of the point forecast, tilted toward the home
side**, and it replicates out of sample:
| | Tune 2021-2023 | 2024 |
|---|---:|---:|
| Slope of actual on predicted margin | 1.1428 ± 0.0435 | 1.2442 ± 0.0851 |
| Intercept, points | −1.606 | −2.108 |
| Mean residual, points | −1.104 | −0.927 |
A slope 3.3 standard errors above one means the forecasts are systematically too
timid about mismatches; probabilities built from them sit too close to 0.5, so low
deciles land below their predicted rate and high deciles above. The negative
intercept pushes the whole curve down, which is why the low end misses by 13.67 pp
and the high end by 1.23 pp. **That is the asymmetry.** It is not a distributional
assumption, not a variance function and not a home-field constant, and there is no
key in `configs/default.toml` that sets it.
The mechanism is named in the campaign document: both the affine points
calibration and σ are fitted on the games *accumulated so far* in a season, and the
ratings feeding them improve as the season goes on, so a slope fitted on weeks 2-9
under-scales week 10 and a σ fitted on weeks 2-9 (18.46) over-covers week 10
(16.55). **This does not license relaxing the out-of-sample rule.** Fitting either
estimator on the training window costs L2 0.44 points of MAE and inverts the
ordering against Elo. The defect is the *shape* of the accumulation window. A
trailing window is out of sample too, and it is the next campaign's first
pre-registered item.
What this does not settle
docs/adr/0007-tuned-constants.md
**2025 is untouched and stays that way.** The harness refuses it without
`unlock_holdout=True` and no code path passes it.
**2024 has now been read once.** Every future decision that reads it again must
say so publicly and re-designate the split (report 02 §5.1). This ADR is that
public statement for this campaign.
**The next campaign's pre-registered items**, named here so they are on the record
before they are searched: widen `c_grid` above 32; a trailing-window σ and a
trailing-window points calibration; a win-probability model so `garbage_time.mode
= "leverage"` becomes measurable; and an identification strategy for `h` that does
not depend on 37 neutral-site games.
Consequences, including the uncomfortable ones
docs/adr/0006-fit-universe.md
**The margin is inside the noise, and three of the other criteria point the other
way.** `model` wins the pre-registered rule by 0.055 points of MAE over
`fbs_vs_fbs`, and the honest floor for distinguishing MAE over three seasons is
around 0.3 points. On straight-up accuracy (69.65% vs 69.21%), on calibration
deviation (11.11pp vs 13.67pp) and on retrodictive violations (0.2006 vs 0.2015),
the FBS-only universe is ahead. A pre-registered rule that picks one column while
three others disagree is a rule doing its job, and a reader is entitled to know it
was close.
**It is a dial.** Kendall's tau between the incumbent's ranking and the FBS-only
ranking is **0.9344**, against the 0.985 floor the published `q_ref` sweep never
dipped below. Mean absolute rank change 3.32 places, maximum 17. By the project's
own published standard that is a dial, and the config now says so.
**The review's own example does not reproduce, and that is recorded rather than
quietly dropped.** It reported James Madison moving #7 → #4 under `fbs_vs_fbs`.
On this build JMU is #4 under all three universes. The review measured against a
baseline this repository no longer has: σ was the 15.3 constant rather than an
estimate ([ADR-adjacent, review S6](../analysis/fresh-eyes-review.md)) and its own
§S4 baseline used in-sample blend weights. The *sensitivity* is real and larger
than `q_ref`'s; the particular team it landed on was a property of the
configuration it was measured under. It now shows up on UCF (#66 → #83), Stanford
(#80 → #65) and Northern Illinois (#97 → #80).
**The G5 caveat is now a number.** Mean Power of a fourteen-team P4 sample minus a
fourteen-team G5 sample, 2023 through week 10: 10.83 points under `model`, 9.64
under `fbs_vs_fbs`, 11.23 under `all`. Dropping the non-FBS teams *narrows* the
P4-over-G5 gap by 1.19 points, so the mixed-division universe is mildly favourable
to G5 teams. That is the review's direction, at a smaller magnitude than its
framing suggests. That caveat belongs on the methodology page and is no longer unstated.
**What this does not settle.** 2024 (validate) and 2025 (holdout) are untouched.
If this decision is ever revisited against them, that has to be said publicly and
the split re-designated (report 02 §5.1). And the mechanism numbers in §3 of the
analysis are one week of one season: the direction is stable, the magnitude is not
claimed to be.
Which of these actually decides anything?
A table cannot tell you that. Moving one can, so the ones that matter most are settings you can move yourself, against the same games, with the shift measured for you.