Kickoff model
Walk-forward logistic regression on nflverse pregame features.2019–2025, 1,865 decided regular-season games. Not betting advice. Regular-season evaluation only.
v0.2 adds reliability-weighted entering-QB passing efficiency. It was promoted because Brier and log loss improved without collapsing accuracy. Calibration error is slightly worse than v0.1.
Left is v0.1, right is v0.2. v0.1-qb-efficiency improved Brier by 0.0016 and log loss by 0.0034 with accuracy 0.48pp vs v0.1. Brier improved in 4 seasons. The entering-record baseline remains 63.8% SU — still ahead of Kickoff straight-up.
Straight-up accuracy on the same 1,865 games. A simple entering-record rule is still ahead straight-up; Kickoff v0.2 is the probabilistic model.
Kickoff is reasonably calibrated near coin-flip games but has historically been overconfident in some stronger-favorite ranges, especially around 70–74%.
Filled: v0.2. Open: v0.1. Diagonal is perfect calibration.
| Bucket | n | Predicted | Observed |
|---|---|---|---|
| 50–54% | 324 | 52.4% | 53.4% |
| 55–59% | 370 | 57.5% | 57.8% |
| 60–64% | 306 | 62.5% | 61.8% |
| 65–69% | 238 | 67.6% | 61.3% |
| 70–74% | 218 | 72.4% | 62.4% |
| 75%+ | 409 | 81.5% | 77.5% |
Accuracy by the favorite’s stated probability, not by the confidence score.
Same 1,865 stored walk-forward games, sliced. Slices under 80 games are not shown here. QB reliability is high in the experiment archive overall (95% coverage) but is not stored per game in the public history file, so Low/Medium/High QB buckets are not displayed.
Fav 65% · obs 61% · Brier 0.236
Fav 65% · obs 63% · Brier 0.227
Fav 66% · obs 65% · Brier 0.221
Fav 52% · obs 53% · Brier 0.248
Fav 60% · obs 60% · Brier 0.239
Fav 70% · obs 62% · Brier 0.241
Fav 82% · obs 78% · Brier 0.175
Fav 67% · obs 63% · Brier 0.223
Fav 64% · obs 63% · Brier 0.234
Fav 67% · obs 63% · Brier 0.228
Fav 65% · obs 63% · Brier 0.227
Highest-confidence incorrect calls from the historical archive. These are stored walk-forward predictions, not reconstructed after the fact.
Separate ridge regression on the same features. Mean absolute error 7.7 points per team. Projected scores are not implied by win probability.
Opponent differentials the logistic model sees.
Production is Model v0.2: the v0.1 logistic plus reliability-weighted historical QB passing efficiency. Frozen v0.1 remains the baseline. Official evaluation is 2019–2025 only.
v0.2 adds reliability-weighted entering-QB passing efficiency from prior games. Identity is derived from team passing, not from a named Week 1 starter list.
confidence = clip(0.35 * 2*|p-0.5| + 0.65 * dataCompleteness, 0.08, 0.92). Edge is how far the probability sits from 50%. Completeness is how much current-season (vs prior-season) information was available. Confidence is not the win probability, and it is not a betting recommendation. Average confidence on the backtest: 60%.
temperature improved overall Brier by 0.0008 and log loss by 0.0022, but worsened at least one season’s Brier by 0.0036. Raw v0.2 remained preferable overall. Production remains raw v0.2. ECE looking prettier is not a promotion standard.
272 published 2026 snapshots. No 2026 games have been played, so these rely on 2025 priors, home field, and schedule rest. They are labeled preseason model snapshots.