goal-line-calibration
How close do statistical and machine-learning models get to the bookmakers' probabilities for home win, draw and away win? Final test 2021-22 to 2025-26, 8906 first-division matches, scored once with hyperparameters frozen in commit 9d04047d5f764eef4de8a43179fa2ff99f111393, by code at commit 33ea98909b60ecc5ef4717846e13724bb9d1fcb3.
No model beats the de-margined Bet365 probabilities. The closest, Logistic regression, has a log loss higher by 0.0179 (95% CI 0.0145 to 0.0213).
Final test
Mean per match; lower is better. The difference is model minus Bet365 (Shin) log loss, with a 95% interval from a paired bootstrap over matchdays.
| Model | Log loss | RPS | Brier | Δ log loss vs Shin | 95% CI |
|---|---|---|---|---|---|
| Uniform | 1.0986 | 0.2354 | 0.6667 | +0.1283 | +0.1190 to +0.1378 |
| Frequency | 1.0752 | 0.2305 | 0.6507 | +0.1049 | +0.0968 to +0.1128 |
| Elo | 0.9939 | 0.2025 | 0.5932 | +0.0235 | +0.0197 to +0.0272 |
| Dixon–Coles | 0.9886 | 0.2009 | 0.5891 | +0.0182 | +0.0147 to +0.0215 |
| Logistic regression | 0.9882 | 0.2006 | 0.5891 | +0.0179 | +0.0145 to +0.0213 |
| XGBoost | 0.9901 | 0.2011 | 0.5902 | +0.0197 | +0.0162 to +0.0235 |
| Bet365, Shin | 0.9703 | 0.1952 | 0.5771 | — | — |
Serie A
| Model | Matches | Log loss | RPS | Brier |
|---|---|---|---|---|
| Uniform | 1899 | 1.0986 | 0.2324 | 0.6667 |
| Frequency | 1899 | 1.0911 | 0.2318 | 0.6618 |
| Elo | 1899 | 0.9998 | 0.2001 | 0.5983 |
| Dixon–Coles | 1899 | 0.9846 | 0.1958 | 0.5873 |
| Logistic regression | 1899 | 0.9886 | 0.1970 | 0.5904 |
| XGBoost | 1899 | 0.9891 | 0.1974 | 0.5908 |
| Bet365, Shin | 1899 | 0.9712 | 0.1914 | 0.5785 |
By league and by season
| Log loss | Serie A | Premier League | La Liga | Bundesliga | Ligue 1 |
|---|---|---|---|---|---|
| Uniform | 1.0986 | 1.0986 | 1.0986 | 1.0986 | 1.0986 |
| Frequency | 1.0911 | 1.0692 | 1.0659 | 1.0710 | 1.0785 |
| Elo | 0.9998 | 0.9854 | 0.9827 | 0.9958 | 1.0077 |
| Dixon–Coles | 0.9846 | 0.9852 | 0.9829 | 0.9925 | 0.9998 |
| Logistic regression | 0.9886 | 0.9795 | 0.9822 | 0.9890 | 1.0037 |
| XGBoost | 0.9891 | 0.9815 | 0.9846 | 0.9914 | 1.0059 |
| Bet365, Shin | 0.9712 | 0.9597 | 0.9656 | 0.9743 | 0.9831 |
Calibration
Ten equal-count bins per outcome. Points on the dashed line are perfectly calibrated.
| Expected calibration error | Home | Draw | Away |
|---|---|---|---|
| Elo | 0.0366 | 0.0139 | 0.0394 |
| Dixon–Coles | 0.0152 | 0.0085 | 0.0136 |
| Logistic regression | 0.0254 | 0.0125 | 0.0219 |
| XGBoost | 0.0247 | 0.0139 | 0.0222 |
| Bet365, Shin | 0.0137 | 0.0118 | 0.0138 |
Sensitivity
The same comparison against proportionally de-margined Bet365 odds and against the best price reported in the data (its MaxHome, MaxDraw and MaxAway columns). 2323 test matches have best prices whose implied probabilities sum to less than one; they are normalised anyway.
| Model | Reference | Matches | Δ log loss | 95% CI |
|---|---|---|---|---|
| Elo | Bet365, proportional | 8906 | +0.0230 | +0.0194 to +0.0266 |
| Dixon–Coles | Bet365, proportional | 8906 | +0.0177 | +0.0144 to +0.0210 |
| Logistic regression | Bet365, proportional | 8906 | +0.0173 | +0.0140 to +0.0207 |
| XGBoost | Bet365, proportional | 8906 | +0.0192 | +0.0157 to +0.0228 |
| Elo | Best odds, proportional | 8906 | +0.0237 | +0.0199 to +0.0273 |
| Dixon–Coles | Best odds, proportional | 8906 | +0.0184 | +0.0149 to +0.0217 |
| Logistic regression | Best odds, proportional | 8906 | +0.0180 | +0.0148 to +0.0214 |
| XGBoost | Best odds, proportional | 8906 | +0.0199 | +0.0164 to +0.0234 |
Models against each other
| Model | Against | Δ log loss | 95% CI |
|---|---|---|---|
| Elo | Dixon–Coles | +0.0053 | +0.0015 to +0.0088 |
| Elo | Logistic regression | +0.0057 | +0.0035 to +0.0078 |
| Elo | XGBoost | +0.0038 | +0.0018 to +0.0058 |
| Dixon–Coles | Logistic regression | +0.0004 | -0.0028 to +0.0035 |
| Dixon–Coles | XGBoost | -0.0015 | -0.0048 to +0.0018 |
| Logistic regression | XGBoost | -0.0019 | -0.0035 to -0.0003 |
Intervals, not p-values: with this many comparisons, an interval that excludes zero is evidence, not proof.
Development and frozen hyperparameters
Chosen by mean log loss over the walk-forward development seasons 2005-06 to 2020-21, then frozen before the final test, in commit 9d04047d5f764eef4de8a43179fa2ff99f111393. The final test ran with code at commit 33ea98909b60ecc5ef4717846e13724bb9d1fcb3.
| Model | Hyperparameters | Development log loss |
|---|---|---|
| Elo | K 20, H 100, reversion 0, draw peak 0.3, scale 300 | 0.9921 |
| Dixon–Coles | ξ 0.0019 per day | 0.9876 |
| Logistic regression | C 0.1 | 0.9902 |
| XGBoost | depth 2, learning rate 0.1, min child weight 10 | 0.9911 |
Home advantage and 2020-21
Share of home wins per season. In 2020-21 most matches were played without crowds; the models use one home advantage for every league and season and do not adjust for it.
Data
Matches and odds come from xgabora/Club-Football-Match-Data (MIT licence), which collects results and odds from Football-Data.co.uk and Elo snapshots from ClubElo. Only results, dates and odds are used; Elo and form are recomputed here. File SHA-256 ef224cf2c252f07a842b3bcfd4ba5c718c25cedd8937ffa174a74b86b5ba4221.
| Quality check | Rows |
|---|---|
| b365_odds_blanked | 6 |
| no_result | 1 |
Limitations
- The bookmakers know things the models do not: injuries, line-ups, news.
- The Bet365 columns are documented only as match odds; they need not be closing odds, which would be a harder benchmark.
- Elo and Dixon–Coles see only domestic league matches; cups and European games are missing.
- One home advantage for all leagues, and no special treatment of the 2020-21 season.
- This is a forecast evaluation, not a betting strategy.
Reproduce
pip install -r requirements-lock.txt && pip install -e . --no-deps
goalline download # pinned file, SHA-256 checked
goalline validate
goalline develop # slow: tunes every model on 2005-06 to 2020-21
goalline final # uses config/selected.json
goalline report
Generated 2026-10-07 from code commit 33ea98909b60ecc5ef4717846e13724bb9d1fcb3, hyperparameters from commit 9d04047d5f764eef4de8a43179fa2ff99f111393. Environment: python 3.13.15, numpy 2.5.3, pandas 2.3.3, scipy 1.18.1, scikit-learn 1.7.2, xgboost 3.4.1. Dixon–Coles fits that did not converge: 0 of 260.