Ranked Probability Score
0.2092
vs 0.2371 equal-probability baseline
11.8% lower error
Lower is better. RPS rewards well-calibrated probabilities across home, draw and away outcomes.
EPL model validation
A transparent record of what the FiveStat model gets right, how it compares with simple baselines, and where it still struggles.
Headline scorecard
Backtest file updated 19 Sep 2026 at 23:21 BST
These are historical out-of-sample backtest results, not a live betting record. Each fixture was predicted using only information available before that game.
Ranked Probability Score
0.2092
vs 0.2371 equal-probability baseline
11.8% lower error
Lower is better. RPS rewards well-calibrated probabilities across home, draw and away outcomes.
1X2 accuracy
51.0%
vs 43.6% always-home baseline
+7.4 percentage points
How often the single most likely match outcome was correct.
Brier score
0.2188
Probability error score
Lower is better. Penalises confident probabilities when the predicted event does not occur.
Test sample
1,101
2023/24–2025/26 (combined)
40 fixtures were skipped where the model lacked sufficient prior information.
Gameweek stability
Lower is better
Each point combines the same-numbered gameweek across the three evaluated seasons. Variation is expected in small gameweek samples; the overall scorecard above is the more reliable summary.
Supporting metrics
How it was tested
Ratings use only matches completed before the gameweek being predicted.
Probabilities and expected goals are generated before revealing that gameweek's results.
Forecasts are compared with results using proper probability scores and simple benchmarks.
Known limitation
The aggregate draw probability is reasonably close to the observed rate, but a draw was never the model's single highest-probability outcome in this test. That is why FiveStat publishes probability scores alongside headline classification accuracy.
Scope boundary
It is not yet an immutable live prediction record.
It does not claim realised betting profit or ROI.
It does not yet compare every forecast with closing market prices.
The next validation layer is to archive pre-match production forecasts and score them after results arrive. Until then, live and backtested performance remain clearly separated.
Want the mechanics?
See how attack and defence ratings, recent form, scoreline probabilities and season simulations are produced.