Skip to content
SPOLYTICS
Back to board

Model validation report

Model probabilities are scored against the actual results of finished games. Brier score, calibration curve, recent accuracy.

Model accuracy summary

How honest are our probabilities? These are scored results against finished matches, shown next to published benchmarks.

Soccer 3-way RPS
Top tier ≤0.205 · Good ≤0.21
Pre-game Brier
Good ≤0.22 · Constant 50%: 0.25
Calibration error (ECE)
≤ 0.05

Only current-engine forecasts saved before kickoff are scored. Fewer than 30 games is provisional. Brier 0.25 is the constant-50% reference; absolute scores across market mixes are not directly comparable.

Soccer — per-league scored results

LeagueMatches scored3-way RPSBrier scoreDraw calibration
EPL13Provisional—0.244no draw sample
Championship15Provisional—0.223no draw sample
LaLiga23Provisional—0.199no draw sample
Bundesliga11Provisional—0.191no draw sample
Serie A16Provisional—0.204no draw sample
Ligue 112Provisional—0.202no draw sample
UCLNo finished matches scored yet — first results appear after 30 games.
UEL17Provisional—0.234no draw sample
MLS26Provisional—0.216no draw sample
J League12Provisional—0.194no draw sample
K League 18Provisional—0.228no draw sample
Leagues CupNo finished matches scored yet — first results appear after 30 games.

What RPS means: Ranked Probability Score rewards being confident when right and humble when wrong; lower is better, and strong public soccer models sit near 0.205.

Baseball, basketball, football & esports scored results

LeagueGamesSamplesBrier (raw)
MLB147120540.194 (0.202)Good
WNBA27162Provisional0.220 (0.220)
NBA00Provisional— (—)
NFL18108Provisional0.241 (0.241)
LoL832Provisional0.203 (0.209)
CS2903040.237 (0.239)Watch

Weighted by sample size across the last 20 morning grading runs. Overall calibration error (ECE) is 0.019 (raw 0.063); the benchmark is 0.05 or lower.Top tier

Per-league metric charts

RPS & Brier (lower is better)

Dashed lines are the good benchmarks (RPS 0.21, Brier 0.22); the red line at 0.25 is the base-rate baseline.

Calibration curve (3-way outcome)

No finished matches scored yet — first results appear after 30 games.

X axis is the predicted probability bucket, Y axis the observed frequency. Closer to the dashed line means more honest probabilities. Switch leagues to inspect a single league's curve.

Draw honesty

No finished matches scored yet — first results appear after 30 games.

Draws are the hardest outcome to get right. Observed league draw rates usually fall between 23% and 30%. A plain scoring distribution underrates low-scoring scorelines, so the Spolytics Engine applies its own low-score correction lifting 0-0 and 1-1. Leagues where predicted and actual diverge by more than 4pp get refit on the rolling window.

Last updated 2026-10-03 This summary is fully free. Download the raw data as CSV

Daily learning trend

Last 13 days · 12,713 graded samples · 180 refits

Every morning the previous day's finished games are graded and the per-league, per-market calibrators are refitted on those results. The lines below show how that retraining moved actual performance.

.250.150

7-day rolling Brier Daily Brier (calibrated) Daily Brier (raw) Calibrators refitted that day

Brier, last 7d

0.2023

vs prior 7d -0.0009

Accuracy, last 7d

69.2%

vs prior 7d +0.3pp

Last retrain

2026-10-03

12 calibrators updated · 151 samples

ECE

11.8%

raw 9.0%

Show daily table
DateSamplesAccuracyBrierRawECERefits
2026-10-0315174.8%0.20200.201111.8%12
2026-10-02475.0%0.21350.273824.9%0
2026-09-2221771.0%0.19490.20254.3%0
2026-09-201,77368.9%0.20330.20593.1%0
2026-09-1152471.4%0.18570.19907.7%17
2026-09-101,44467.8%0.21020.21384.2%15
2026-09-091,33469.3%0.20020.20784.4%14
2026-09-081,03869.5%0.20430.20513.4%17
2026-09-071,45870.5%0.19930.20203.5%18
2026-09-061,64867.1%0.20990.21084.1%19
2026-09-0593967.5%0.20120.20514.2%21
2026-09-0479972.6%0.18980.20176.5%22
2026-09-031,38467.8%0.20770.20844.5%25

Pick weight retraining

Every night the pick-selection weights (probability, bucket calibration, sport performance) are refit on the last 30 days of graded results. The first two thirds train, the last third is held out for scoring. The probability models themselves are unchanged.

Active weights

prob 0.80 · calib 0.13 · perf 0.07

Adopted in last 45 days: 1

Last refit

2026-10-03

Improvement below threshold · previous weights kept

train 8,556 · eval 571 · top 15

MetricBeforeAfterChange
Top-pick hit rate68.4%68.4%0.0%p
Top-pick Brier0.19550.19550.0000
DateStatusHit beforeHit afterBrier beforeBrier after
2026-10-03kept68.4%68.4%0.19550.1955
2026-10-02kept68.9%68.9%0.19960.1996
2026-09-22adopted60.0%64.4%0.21540.1770
2026-09-21kept71.1%71.1%0.17310.1731
2026-09-20kept77.8%77.8%0.16120.1612

Scored 2026-09-26 to 2026-10-02 · 1 finished games · n = 90 markets · stored Oct 3 (Sat) · 8:00 AM UTC

Overall summary

2026-09-26 ~ 2026-10-02 · 1 finished games · 90 markets

Brier score

0.2066

0.25 = blindly guessing 50%

Accuracy

70.0%

Directional at 50% threshold

Log loss

0.599

0.693 = coin flip

ECE

10.2%

Calibration error

Equal-mass ECE (ACE) 7.0% — fixed 10-bin ECE distorts when predictions cluster.

Better than a coin flip, but not by much.

Calibration effect (before → after)

  • Brier 0.2047 → 0.2066
  • ECE 6.7% → 10.2%
  • Accuracy 70.0% → 70.0%
  • Mean predicted 36.8% → 29.7% · actual rate 36.7%

Deep validation is Pro

Summary metrics and the first calibration curve are free. Per-market curves, walk-forward validation, the daily table, and windows longer than 7 days are computed for Pro.

See Pro

$7.99/mo or $59/yr · cancel anytime

Combo pick record

Loading…

NFL featured combo record

Loading…

This report evaluates only model snapshots saved before each game. Post-game recalculations and older samples of unknown origin are excluded from training and official evaluation. It is a model evaluation, separate from the published locked-pick record, and results may be empty early in a new collection period.