EPL model validation

Model Performance

A transparent record of what the FiveStat model gets right, how it compares with simple baselines, and where it still struggles.

Historical walk-forward backtest 2023/24–2025/26 (combined) 1,101 predictions

Headline scorecard

Probability quality first

Backtest file updated 19 Sep 2026 at 23:21 BST

These are historical out-of-sample backtest results, not a live betting record. Each fixture was predicted using only information available before that game.

1X2 accuracy

51.0%

vs 43.6% always-home baseline

+7.4 percentage points

How often the single most likely match outcome was correct.

Brier score

0.2188

Probability error score

Lower is better. Penalises confident probabilities when the predicted event does not occur.

Test sample

1,101

2023/24–2025/26 (combined)

40 fixtures were skipped where the model lacked sufficient prior information.

Gameweek stability

RPS through the test window

Lower is better

Gameweek RPS values are available in the backtest data.

Each point combines the same-numbered gameweek across the three evaluated seasons. Variation is expected in small gameweek samples; the overall scorecard above is the more reliable summary.

Supporting metrics

Other ways we test the model

Decisive-match accuracy 67.4% 832 matches, draws excluded
Over / Under 2.5 57.3% 1101 evaluated matches
Top correct score 12.0% Exact result matched the highest-probability scoreline
xG mean absolute error 0.694 Average absolute team-goal forecast error

How it was tested

Walk-forward, not in-sample

  1. 1
    Train on the past

    Ratings use only matches completed before the gameweek being predicted.

  2. 2
    Freeze the forecast

    Probabilities and expected goals are generated before revealing that gameweek's results.

  3. 3
    Score every prediction

    Forecasts are compared with results using proper probability scores and simple benchmarks.

Known limitation

Draws are hard to rank first

22.0%average model draw probability
24.4%actual draw rate

The aggregate draw probability is reasonably close to the observed rate, but a draw was never the model's single highest-probability outcome in this test. That is why FiveStat publishes probability scores alongside headline classification accuracy.

Scope boundary

What this page does not claim

It is not yet an immutable live prediction record.

It does not claim realised betting profit or ROI.

It does not yet compare every forecast with closing market prices.

The next validation layer is to archive pre-match production forecasts and score them after results arrive. Until then, live and backtested performance remain clearly separated.

Want the mechanics?

Read the full methodology

See how attack and defence ratings, recent form, scoreline probabilities and season simulations are produced.

Explore methodology