Methodology · 2025 through week 16

How this works, and where it is weak.

This is the page with everything on it: every number the model runs on, how it scored against everybody else, and the places it falls short. You do not need it to read the poll. It is here so that nothing about the poll is hidden.

This whole page answers one question from the front of the site: how do you put a number on how hard a season was?

Teams are ranked by −log₁₀ P(W ≥ W_t): how improbable it is that a team of reference quality q_ref would have gone at least this well against this exact schedule.

The stack

The path from a play to a rank, with every constant that enters along the way.

INTHE INPUTSgames 1,637 · plays 212,826L1EFFICIENCYλ1 200L2RESULTS COREλ 0.500 · C 32 · β 7L3POWERw1 0.518 · w2 0.317L4RÉSUMÉ RATINGHFA 4.042CSCHEDULE ODDSq_ref 15.641
Never enters the poll. Connected to nothing above.
  • no preseason ranking
  • no human poll
  • no committee
  • no conference
  • no market lines
  • no recruiting rankings
  • no third party's model output
  • nothing carried over from last season
  • no logos

The layers, top to bottom, in the order the data moves through them. Each band carries the published constants that enter at it, and every one of them is in the table further down. The bottom band is the ranking.

Margin lives at L2 and L3 and never reaches the rank key, which is the -log10(P(W >= W_t)) under a Poisson-binomial over the exact schedule; margin never enters string in the table below.

The claim has two halves and both are enforced. Your own margin never enters your own rank: the schedule flattener carries opponent power, site and a boolean win, with no margin column to leak from, and a test scrambles every final score while preserving every winner and asserts the ranking is bit-identical. Your opponents’ margins set what your wins are worth: opponent quality is the L3 power rating, fitted on compressed scoring margin, so a second test refits power from those same scrambled scores and asserts the ranking moves. The wider claim, “margin never enters” full stop, was true of the module and false of the poll. An outside review said so first, and both tests exist so it cannot quietly widen again.

Before each fit the harness rebuilds every design matrix from the frame restricted to the allow-list and requires bit-identity, so an input nobody thought of fails closed.

The publication gate

The bar the build-your-own page mentions, written down before any result was seen.

Publication gateFAILEvery decided criterion passesUndecided criteria are reported as undecided, never as passes: brier_beats_all_baselines, retro_vs_live_monotone
  1. FAILStraight-up accuracy at or above the floorsu_accuracy
    bar 0.7 · actual 0.6871
  2. FAILMean absolute error at or below the ceilingmae
    bar 12.8 · actual 13.04
  3. FAILRoot mean squared error at or below the ceilingrmse
    bar 15.8 · actual 16.53
  4. FAILWorst decile calibration deviation within tolerancecalibration
    bar 5.0 · actual 7.37
  5. FAILRetrodictive violations at or below every baselineviolations_vs_baselinesobserved 0.2019
  6. not yet decidedBrier score beats every baselinebrier_beats_all_baselinesno constant threshold
  7. not yet decidedRetro-vs-live divergence declines monotonicallyretro_vs_live_monotoneno constant threshold

Every threshold above was fixed in the config before anyone had seen a result. I set the bar above every public rating system I could score. Nothing clears it, including this one.

The verdict comes from the published status rather than from the picture, and undecided is never dressed as a pass.

Against every baseline

Strict walk-forward: to predict week N the harness fits through week N−1 of the same season and nothing else. The home-team-always-wins floor is always in the table, because a table without its floor is not a table.

systemnSU%MAERMSEBrierloglossviol%
This poll (schedule odds)158568.71%13.03816.5310.19710.577220.19%
Colley158567.95%13.56217.2370.20630.598919.76%
Elo158568.77%13.55917.0980.20140.588522.03%
L1 efficiency158569.15%13.16516.6840.19950.582624.34%
L2 results core158569.53%13.22716.7400.19750.579421.86%
L3 Power (efficiency + results blend)158568.71%13.03816.5310.19710.577222.80%
Random walker158565.05%14.01517.9840.21780.624420.23%
Résumé (the ordering it replaced)158568.71%13.03816.5310.19710.577220.15%
SRS / Massey158569.09%13.19616.6950.19950.582122.11%
Win percentage158566.12%13.82417.6600.21180.611818.31%
Home team always wins158556.34%15.45819.8630.24750.6885not published

n games scored · SU% winners called right · MAE average miss in points · RMSE the same with blowout misses weighted heavier · Brier and logloss how honest the stated chances were · viol% how often the loser finished above the winner

The poll and the résumé predict through the same power rating, so only the last column tells them apart, and on that column the résumé is the better of the two.

Ranking teams purely by win percentage beats this model on one of these measures. That is the price of caring who you played: a system that ignores schedule will always look tidy on a test that also ignores schedule.

The divergence curve

How far the poll’s answer for a given week moved once the season finished, drawn over every team it ranked.

0.06.112.218.316.84 places0.91147101316week of the season

The gap between what the poll published in week N and what it says about that week after the season, over all 136 ranked teams, weeks 1 through 4 included even though the poll never published them. If this line ever stops declining, the weekly poll was never worth reading.

Every constant this run used

Every one of these was published the same week the poll was. There is no second file with the real settings in it.

constantvalue
C32
beta_w7
bisection_max_iter60
blend_n_games714
bootstrap_draws1000
bootstrap_jobs_requested4
bootstrap_jobs_used1
bootstrap_lambda0.5
bootstrap_n_games1637
bootstrap_n_ranked_teams136
bootstrap_seed20260812
bootstrap_sigma15.884317060721084
cv_folds5
gauss_hermite_nodes20
h_points3.706281732877181
home_field4.041520630398333
home_field_home_and_home4
interval_level0.9
k_points_per_unit72.96415994533689
lambda0.5
lambda_l1200
lambda_l20.5
max_games_one_team15
min_out_of_sample_games40
min_tail1e-300
n_games1637
n_plays212826
n_record_games1637
n_resume_games1637
n_resume_teams315
n_saturated_high3
n_saturated_low50
n_teams315
power_home_field_points3.706281732877181
power_scale_b1
power_scale_n_games714
power_se_median_points1.3981064284088478
power_se_scale0.3168928694216252
power_sigma15.884317060721084
power_sigma_n_out_of_sample_games77
q_ref15.64109836328906
q_ref_pool_size136
recency_gamma1
season2025
seed20260812
sigma15.884317060721084
w1_efficiency0.5184474738298377
w2_results0.3168928694216252
weight_bowl_non_cfp0.25
weight_cfp1
weight_conference_championship1
weight_regular_season1
The named layers and orderings (36 strings)
blend_weight_sourceout_of_sample
bootstrap_jobs_notedraws run sequentially; per-draw streams come from SeedSequence.spawn, so parallelising later cannot move a published number (report 03 §9.3 item 2)
bootstrap_lambda_noteheld at the value the real data's CV selected; the bootstrap propagates sampling uncertainty at a fixed hyperparameter
bootstrap_noteparametric on the FIXED schedule: outcomes redrawn from the fitted model, refit, re-ranked. Games are edges in the schedule graph and are not exchangeable, so resampling them with replacement - which report 02 §3.3's parenthetical specified - is invalid here; see naive_resample_diagnostic. The L1 efficiency half of Power is held at its point estimate because plays are not resimulated, so these intervals are a LOWER BOUND on total uncertainty.
bootstrap_power_sourceL3
bootstrap_schemeparametric_on_fixed_schedule
bootstrap_sigma_sourcewalk_forward_residuals
companion_layerL3_power
cv_groupgame_id
estimatorols
fit_out_of_sample_onlytrue
headline_decided2026-08-12, docs/adr/0005-headline-ordering.md
headline_layerC_schedule_odds
headline_orderingschedule_odds
hindsight_data_bucket2025-post-w01
hindsight_is_livefalse
hindsight_variantA_frozen_form
layerC schedule odds
power_scale_universeblend_regression_on_actual_margin
power_se_noteridge sandwich on the L2 half only (report 02 §3.3), carried onto the points scale by w2. THE EFFICIENCY HALF IS HELD AT ITS POINT ESTIMATE: a play-level covariance over ~2,000 coefficients and 170,000 correlated rows is a different object and is not built, so this is a LOWER BOUND on the uncertainty of an L3 rating. The parametric bootstrap inherits the same limitation and says so
power_sigma_sourcewalk_forward_residuals
power_sourceL3
power_versionv1
provisionalfalse
q_ref_methodpower_rank_25
q_ref_teamArizona
ranking_key-log10(P(W >= W_t)) under a Poisson-binomial over the exact schedule; margin never enters
recipe_config_sha256061a6d4eda23eab4ea587cd4d88f9d619c130c5a44b83dabb8e790ef4b76b061
resume_layerL4 resume rating
resume_targetraw wins, and raw compressed margin; game weights shape the Power fit, not the accomplishment
resume_versionv0
saturation_tiebreakmargin
sigma_sourcewalk_forward_residuals
tail_methodexact_poisson_binomial_dp
tie_breakmid_p
versionv0

Where this is weak

The arguments against every claim on the pages above, verbatim from the decision records, including the ones that record being wrong. Nothing here was softened for the website.

Where this decision is weak

docs/adr/0005-headline-ordering.md

Stated so it is not read as stronger than it is, and reproduced from study §10.4:

- **Four seasons.** The forward-accuracy gap between B and the other two (2 pp on
  9,433 games) is comfortably significant. **The gap between A and C on violations
  (0.1 to 0.5 pp) is not, and nothing here claims it is.** A and C are tied on
  that axis, and what decided between them was the structural finding.
- The postseason axis is unusable at current sample sizes (14 CFP games, 4 non-CFP
  NY6 games) and 2021-2022 cannot contribute to it at all, because the archive
  carries no postseason rows for those seasons.
- 2024 is the validation season. Every conclusion holds on it with the same sign,
  but it was consulted once, in one pass, and must not be consulted again for a
  re-tune.
- 2025 is untouched and stays locked.
- **No bootstrap intervals anywhere.** `model/bootstrap.py` is still a stub, so the
  study reports point estimates and sample sizes and leaves the reader to judge.

Why an unbeaten team can finish behind a one-loss team

The price of C, stated plainly

docs/adr/0005-headline-ordering.md

1. **An unbeaten team can finish behind a one-loss team, and that will require
   explaining every year.** It is the direct consequence of the promise. The
   explanation is on the page: the tail probability, the reference team it was
   measured against by name, and the Power column.
2. **One published constant that A did not have.** `q_ref` is the Power rating of
   the 25th-ranked Power team that week, the least flattering defensible reading of
   ESPN's "average Top-25 team", and a single team that can be *named* each week.
   Study §9 measured the ordering's sensitivity to it rather than asserting it was
   fine: across a 16-point swing in reference quality, Kendall's τ never fell below
   0.985, the mean rank change never reached one place, and at most one team
   entered or left the top 25. 2023 Liberty spans #8 to #12 across that whole swing
   and never reaches Georgia at #7 under any choice. The probability *values* move
   by orders of magnitude, which is exactly why the value is published rather than
   the rank alone.
3. **C loses forward ordering accuracy to B by about 2 points.** Accepted, and for
   a stated reason: forward accuracy is a prediction metric, and the headline poll
   is not the instrument this project ships for prediction. L3 Power is, and it
   beats all three orderings on that axis.

C also inherits the invariance that makes the résumé's zero point harmless: shift
every Power rating by a constant and every rank-derived `q_ref` shifts with it, so
no probability moves at all. Only the `fixed` method breaks that, which is why it
is not the default.

Whether tuning the constants could ever have closed the gate

The two uncomfortable results

docs/adr/0007-tuned-constants.md

### 1. The optimum is a corner solution

`C = 32` is the **top of `c_grid`**. The search did not bracket the optimum; the
data wants to keep going and the published grid stops it. The protocol fixed the
search space as *exactly the config grids* precisely so this campaign could not
widen the net after seeing the numbers, so the boundary is reported rather than
crossed.

It also says something about where the bounds came from. The review's §7 table
lists C's range as Pasteur's cap of 21 and the CFBD SRS walkthrough's ±28. Both
are other people's answers on other people's datasets, and the fitted value is
above both.
**Widening `c_grid` is the first item for the next campaign, and it must be
pre-registered before it is searched.**

### 2. Tuning cannot close the gate, and it is not close

The whole 416-cell factorial spans **0.135 points of MAE**, best cell to worst.
The winner beats the incumbent by **0.0086**. The gate needs **0.219** more than
the incumbent has, and the best cell in the entire searched space is still
**0.210 points above the threshold**.

**The gate gap is built into the design. No amount of tuning closes it.** That is
the most important sentence in this ADR and it is a negative result. Every constant in the pre-
registered space was searched at full resolution and the answer is that these
constants were never what stood between this system and its own thresholds.

The calibration diagnosis: diagnosed, and deliberately unfixed

docs/adr/0007-tuned-constants.md

Four candidate fixes were declared before any was run, with an adoption bar of
**≥ 2.0 pp on tune AND direction holding on 2024**. **None clears it, and the
config does not move on account of the calibration criterion.**

| Candidate | Tune Δ | Clears 2 pp? | 2024 Δ | Verdict |
|---|---:|---|---:|---|
| Student-t margins, at the fitted ν = 92.7 | +2.50 pp | yes | −0.38 pp | direction reverses; not adopted |
| Heteroscedastic σ(\|m̂\|) | −3.22 pp | no | +0.81 pp | fails tune; not adopted |
| Home-and-home `h` | — | — | — | **not runnable** under constraint 2 |
| Favourite-longshot | — | — | — | **the diagnosis** itself, and there is no knob for it |

**Neither named suspect did it.**

- **The normal tail is eliminated.** Maximum likelihood on the walk-forward
  residuals puts ν at **92.7**; skew −0.098, excess kurtosis +0.020, Jarque-Bera
  p = 0.275. These residuals are not distinguishable from normal. Low-ν rows in
  the sweep *do* cut the deviation, and that is a clue rather than an exoneration:
  a t with a matched second moment and small ν has a *narrower body*, so what
  those rows buy is sharpness rather than tail weight.
- **The single home-field constant is eliminated as the cause.** The residual
  mean is −1.13 points at home sites against −0.19 at neutral, an order of
  magnitude too small to make a 13.67 pp decile. It is **separately convicted of
  something else.** The site coefficient the harness uses averages **6.39 points
  with a standard deviation of 4.34** across published weeks, against **1.88 ±
  0.34** from 1,113 home-and-home pairs. Only 37 of 1,585 scored games are at
  neutral sites, so the intercept and the site term are nearly collinear and `h`
  is barely identified. Real defect. Different defect.

**The cause is under-dispersion of the point forecast, tilted toward the home
side**, and it replicates out of sample:

| | Tune 2021-2023 | 2024 |
|---|---:|---:|
| Slope of actual on predicted margin | 1.1428 ± 0.0435 | 1.2442 ± 0.0851 |
| Intercept, points | −1.606 | −2.108 |
| Mean residual, points | −1.104 | −0.927 |

A slope 3.3 standard errors above one means the forecasts are systematically too
timid about mismatches; probabilities built from them sit too close to 0.5, so low
deciles land below their predicted rate and high deciles above. The negative
intercept pushes the whole curve down, which is why the low end misses by 13.67 pp
and the high end by 1.23 pp. **That is the asymmetry.** It is not a distributional
assumption, not a variance function and not a home-field constant, and there is no
key in `configs/default.toml` that sets it.

The mechanism is named in the campaign document: both the affine points
calibration and σ are fitted on the games *accumulated so far* in a season, and the
ratings feeding them improve as the season goes on, so a slope fitted on weeks 2-9
under-scales week 10 and a σ fitted on weeks 2-9 (18.46) over-covers week 10
(16.55). **This does not license relaxing the out-of-sample rule.** Fitting either
estimator on the training window costs L2 0.44 points of MAE and inverts the
ordering against Elo. The defect is the *shape* of the accumulation window. A
trailing window is out of sample too, and it is the next campaign's first
pre-registered item.

What this does not settle

docs/adr/0007-tuned-constants.md

**2025 is untouched and stays that way.** The harness refuses it without
`unlock_holdout=True` and no code path passes it.

**2024 has now been read once.** Every future decision that reads it again must
say so publicly and re-designate the split (report 02 §5.1). This ADR is that
public statement for this campaign.

**The next campaign's pre-registered items**, named here so they are on the record
before they are searched: widen `c_grid` above 32; a trailing-window σ and a
trailing-window points calibration; a win-probability model so `garbage_time.mode
= "leverage"` becomes measurable; and an identification strategy for `h` that does
not depend on 37 neutral-site games.

Consequences, including the uncomfortable ones

docs/adr/0006-fit-universe.md

**The margin is inside the noise, and three of the other criteria point the other
way.** `model` wins the pre-registered rule by 0.055 points of MAE over
`fbs_vs_fbs`, and the honest floor for distinguishing MAE over three seasons is
around 0.3 points. On straight-up accuracy (69.65% vs 69.21%), on calibration
deviation (11.11pp vs 13.67pp) and on retrodictive violations (0.2006 vs 0.2015),
the FBS-only universe is ahead. A pre-registered rule that picks one column while
three others disagree is a rule doing its job, and a reader is entitled to know it
was close.

**It is a dial.** Kendall's tau between the incumbent's ranking and the FBS-only
ranking is **0.9344**, against the 0.985 floor the published `q_ref` sweep never
dipped below. Mean absolute rank change 3.32 places, maximum 17. By the project's
own published standard that is a dial, and the config now says so.

**The review's own example does not reproduce, and that is recorded rather than
quietly dropped.** It reported James Madison moving #7 → #4 under `fbs_vs_fbs`.
On this build JMU is #4 under all three universes. The review measured against a
baseline this repository no longer has: σ was the 15.3 constant rather than an
estimate ([ADR-adjacent, review S6](../analysis/fresh-eyes-review.md)) and its own
§S4 baseline used in-sample blend weights. The *sensitivity* is real and larger
than `q_ref`'s; the particular team it landed on was a property of the
configuration it was measured under. It now shows up on UCF (#66 → #83), Stanford
(#80 → #65) and Northern Illinois (#97 → #80).

**The G5 caveat is now a number.** Mean Power of a fourteen-team P4 sample minus a
fourteen-team G5 sample, 2023 through week 10: 10.83 points under `model`, 9.64
under `fbs_vs_fbs`, 11.23 under `all`. Dropping the non-FBS teams *narrows* the
P4-over-G5 gap by 1.19 points, so the mixed-division universe is mildly favourable
to G5 teams. That is the review's direction, at a smaller magnitude than its
framing suggests. That caveat belongs on the methodology page and is no longer unstated.

**What this does not settle.** 2024 (validate) and 2025 (holdout) are untouched.
If this decision is ever revisited against them, that has to be said publicly and
the split re-designated (report 02 §5.1). And the mechanism numbers in §3 of the
analysis are one week of one season: the direction is stable, the magnitude is not
claimed to be.

Which of these actually decides anything?

A table cannot tell you that. Moving one can, so the ones that matter most are settings you can move yourself, against the same games, with the shift measured for you.

Move one and watch the poll answer →

run 61a5fd2c · published 2026-08-19 23:18:03 UTC · code e215160 · config 1ef7cf23…
q_ref 15.64 (Arizona) · β_w 7 · C 32 · h 3.706 · σ 15.884 · λ₁ 200 · λ₂ 0.5 · k 72.96 · w₁ 0.5184 · w₂ 0.3169