goal-line-calibration

How close do statistical and machine-learning models get to the bookmakers' probabilities for home win, draw and away win? Final test 2021-22 to 2025-26, 8906 first-division matches, scored once with hyperparameters frozen in commit 9d04047d5f764eef4de8a43179fa2ff99f111393, by code at commit 33ea98909b60ecc5ef4717846e13724bb9d1fcb3.

No model beats the de-margined Bet365 probabilities. The closest, Logistic regression, has a log loss higher by 0.0179 (95% CI 0.0145 to 0.0213).

Final test

Mean per match; lower is better. The difference is model minus Bet365 (Shin) log loss, with a 95% interval from a paired bootstrap over matchdays.

ModelLog lossRPSBrierΔ log loss vs Shin95% CI
Uniform1.09860.23540.6667 +0.1283 +0.1190 to +0.1378
Frequency1.07520.23050.6507 +0.1049 +0.0968 to +0.1128
Elo0.99390.20250.5932 +0.0235 +0.0197 to +0.0272
Dixon–Coles0.98860.20090.5891 +0.0182 +0.0147 to +0.0215
Logistic regression0.98820.20060.5891 +0.0179 +0.0145 to +0.0213
XGBoost0.99010.20110.5902 +0.0197 +0.0162 to +0.0235
Bet365, Shin0.97030.19520.5771 — —

Serie A

ModelMatchesLog lossRPSBrier
Uniform18991.09860.23240.6667
Frequency18991.09110.23180.6618
Elo18990.99980.20010.5983
Dixon–Coles18990.98460.19580.5873
Logistic regression18990.98860.19700.5904
XGBoost18990.98910.19740.5908
Bet365, Shin18990.97120.19140.5785

By league and by season

Log lossSerie APremier LeagueLa LigaBundesligaLigue 1
Uniform1.09861.09861.09861.09861.0986
Frequency1.09111.06921.06591.07101.0785
Elo0.99980.98540.98270.99581.0077
Dixon–Coles0.98460.98520.98290.99250.9998
Logistic regression0.98860.97950.98220.98901.0037
XGBoost0.98910.98150.98460.99141.0059
Bet365, Shin0.97120.95970.96560.97430.9831
2026-10-07T20:34:14.032209 image/svg+xml Matplotlib v3.11.2, https://matplotlib.org/ 2021.0 2021.5 2022.0 2022.5 2023.0 2023.5 2024.0 2024.5 2025.0 Season (start year) 0.96 0.97 0.98 0.99 1.00 Mean log loss Elo Dixon–Coles Logistic regression XGBoost Bet365, Shin

Calibration

Ten equal-count bins per outcome. Points on the dashed line are perfectly calibrated.

2026-10-07T20:34:14.207282 image/svg+xml Matplotlib v3.11.2, https://matplotlib.org/ 0.0 0.2 0.4 0.6 0.8 1.0 Predicted probability 0.0 0.2 0.4 0.6 0.8 1.0 Observed frequency Home win 0.0 0.2 0.4 0.6 0.8 1.0 Predicted probability Draw 0.0 0.2 0.4 0.6 0.8 1.0 Predicted probability Away win Elo Dixon–Coles Logistic regression XGBoost Bet365, Shin
Expected calibration errorHomeDrawAway
Elo0.03660.01390.0394
Dixon–Coles0.01520.00850.0136
Logistic regression0.02540.01250.0219
XGBoost0.02470.01390.0222
Bet365, Shin0.01370.01180.0138

Sensitivity

The same comparison against proportionally de-margined Bet365 odds and against the best price reported in the data (its MaxHome, MaxDraw and MaxAway columns). 2323 test matches have best prices whose implied probabilities sum to less than one; they are normalised anyway.

ModelReferenceMatchesΔ log loss95% CI
EloBet365, proportional8906+0.0230+0.0194 to +0.0266
Dixon–ColesBet365, proportional8906+0.0177+0.0144 to +0.0210
Logistic regressionBet365, proportional8906+0.0173+0.0140 to +0.0207
XGBoostBet365, proportional8906+0.0192+0.0157 to +0.0228
EloBest odds, proportional8906+0.0237+0.0199 to +0.0273
Dixon–ColesBest odds, proportional8906+0.0184+0.0149 to +0.0217
Logistic regressionBest odds, proportional8906+0.0180+0.0148 to +0.0214
XGBoostBest odds, proportional8906+0.0199+0.0164 to +0.0234

Models against each other

ModelAgainstΔ log loss95% CI
EloDixon–Coles+0.0053+0.0015 to +0.0088
EloLogistic regression+0.0057+0.0035 to +0.0078
EloXGBoost+0.0038+0.0018 to +0.0058
Dixon–ColesLogistic regression+0.0004-0.0028 to +0.0035
Dixon–ColesXGBoost-0.0015-0.0048 to +0.0018
Logistic regressionXGBoost-0.0019-0.0035 to -0.0003

Intervals, not p-values: with this many comparisons, an interval that excludes zero is evidence, not proof.

Development and frozen hyperparameters

Chosen by mean log loss over the walk-forward development seasons 2005-06 to 2020-21, then frozen before the final test, in commit 9d04047d5f764eef4de8a43179fa2ff99f111393. The final test ran with code at commit 33ea98909b60ecc5ef4717846e13724bb9d1fcb3.

ModelHyperparametersDevelopment log loss
EloK 20, H 100, reversion 0, draw peak 0.3, scale 3000.9921
Dixon–Colesξ 0.0019 per day0.9876
Logistic regressionC 0.10.9902
XGBoostdepth 2, learning rate 0.1, min child weight 100.9911

Home advantage and 2020-21

Share of home wins per season. In 2020-21 most matches were played without crowds; the models use one home advantage for every league and season and do not adjust for it.

2026-10-07T20:34:14.330846 image/svg+xml Matplotlib v3.11.2, https://matplotlib.org/ 2000 2005 2010 2015 2020 2025 Season (start year) 0.38 0.40 0.42 0.44 0.46 0.48 0.50 0.52 Home-win share 2020-21 Bundesliga Premier League Ligue 1 Serie A La Liga

Data

Matches and odds come from xgabora/Club-Football-Match-Data (MIT licence), which collects results and odds from Football-Data.co.uk and Elo snapshots from ClubElo. Only results, dates and odds are used; Elo and form are recomputed here. File SHA-256 ef224cf2c252f07a842b3bcfd4ba5c718c25cedd8937ffa174a74b86b5ba4221.

Quality checkRows
b365_odds_blanked6
no_result1

Limitations

Reproduce

pip install -r requirements-lock.txt && pip install -e . --no-deps
goalline download    # pinned file, SHA-256 checked
goalline validate
goalline develop     # slow: tunes every model on 2005-06 to 2020-21
goalline final       # uses config/selected.json
goalline report

Generated 2026-10-07 from code commit 33ea98909b60ecc5ef4717846e13724bb9d1fcb3, hyperparameters from commit 9d04047d5f764eef4de8a43179fa2ff99f111393. Environment: python 3.13.15, numpy 2.5.3, pandas 2.3.3, scipy 1.18.1, scikit-learn 1.7.2, xgboost 3.4.1. Dixon–Coles fits that did not converge: 0 of 260.