Model validation report
Model probabilities are scored against the actual results of finished games. Brier score, calibration curve, recent accuracy.
Model accuracy summary
How honest are our probabilities? These are scored results against finished matches, shown next to published benchmarks.
Only current-engine forecasts saved before kickoff are scored. Fewer than 30 games is provisional. Brier 0.25 is the constant-50% reference; absolute scores across market mixes are not directly comparable.
Soccer — per-league scored results
| League | Matches scored | 3-way RPS | Brier score | Draw calibration |
|---|---|---|---|---|
| EPL | 13Provisional | — | 0.244 | no draw sample |
| Championship | 15Provisional | — | 0.223 | no draw sample |
| LaLiga | 23Provisional | — | 0.199 | no draw sample |
| Bundesliga | 11Provisional | — | 0.191 | no draw sample |
| Serie A | 16Provisional | — | 0.204 | no draw sample |
| Ligue 1 | 12Provisional | — | 0.202 | no draw sample |
| UCL | No finished matches scored yet — first results appear after 30 games. | |||
| UEL | 17Provisional | — | 0.234 | no draw sample |
| MLS | 26Provisional | — | 0.216 | no draw sample |
| J League | 12Provisional | — | 0.194 | no draw sample |
| K League 1 | 8Provisional | — | 0.228 | no draw sample |
| Leagues Cup | No finished matches scored yet — first results appear after 30 games. | |||
What RPS means: Ranked Probability Score rewards being confident when right and humble when wrong; lower is better, and strong public soccer models sit near 0.205.
Baseball, basketball, football & esports scored results
| League | Games | Samples | Brier (raw) |
|---|---|---|---|
| MLB | 146 | 11964 | 0.194 (0.202)Good |
| WNBA | 27 | 162Provisional | 0.220 (0.220) |
| NBA | 0 | 0Provisional | — (—) |
| NFL | 18 | 108Provisional | 0.241 (0.241) |
| LoL | 8 | 32Provisional | 0.203 (0.209) |
| CS2 | 90 | 304 | 0.237 (0.239)Watch |
Weighted by sample size across the last 20 morning grading runs. Overall calibration error (ECE) is 0.020 (raw 0.063); the benchmark is 0.05 or lower.Top tier
Per-league metric charts
RPS & Brier (lower is better)
Dashed lines are the good benchmarks (RPS 0.21, Brier 0.22); the red line at 0.25 is the base-rate baseline.
Calibration curve (3-way outcome)
No finished matches scored yet — first results appear after 30 games.
X axis is the predicted probability bucket, Y axis the observed frequency. Closer to the dashed line means more honest probabilities. Switch leagues to inspect a single league's curve.
Draw honesty
No finished matches scored yet — first results appear after 30 games.
Draws are the hardest outcome to get right. Observed league draw rates usually fall between 23% and 30%. A plain scoring distribution underrates low-scoring scorelines, so the Spolytics Engine applies its own low-score correction lifting 0-0 and 1-1. Leagues where predicted and actual diverge by more than 4pp get refit on the rolling window.
Last updated 2026-10-03 This summary is fully free. Download the raw data as CSV
Daily learning trend
Last 13 days · 12,713 graded samples · 180 refitsEvery morning the previous day's finished games are graded and the per-league, per-market calibrators are refitted on those results. The lines below show how that retraining moved actual performance.
7-day rolling Brier Daily Brier (calibrated) Daily Brier (raw) Calibrators refitted that day
Brier, last 7d
0.2023
vs prior 7d -0.0009
Accuracy, last 7d
69.2%
vs prior 7d +0.3pp
Last retrain
2026-10-03
12 calibrators updated · 151 samples
ECE
11.8%
raw 9.0%
Show daily table
| Date | Samples | Accuracy | Brier | Raw | ECE | Refits |
|---|---|---|---|---|---|---|
| 2026-10-03 | 151 | 74.8% | 0.2020 | 0.2011 | 11.8% | 12 |
| 2026-10-02 | 4 | 75.0% | 0.2135 | 0.2738 | 24.9% | 0 |
| 2026-09-22 | 217 | 71.0% | 0.1949 | 0.2025 | 4.3% | 0 |
| 2026-09-20 | 1,773 | 68.9% | 0.2033 | 0.2059 | 3.1% | 0 |
| 2026-09-11 | 524 | 71.4% | 0.1857 | 0.1990 | 7.7% | 17 |
| 2026-09-10 | 1,444 | 67.8% | 0.2102 | 0.2138 | 4.2% | 15 |
| 2026-09-09 | 1,334 | 69.3% | 0.2002 | 0.2078 | 4.4% | 14 |
| 2026-09-08 | 1,038 | 69.5% | 0.2043 | 0.2051 | 3.4% | 17 |
| 2026-09-07 | 1,458 | 70.5% | 0.1993 | 0.2020 | 3.5% | 18 |
| 2026-09-06 | 1,648 | 67.1% | 0.2099 | 0.2108 | 4.1% | 19 |
| 2026-09-05 | 939 | 67.5% | 0.2012 | 0.2051 | 4.2% | 21 |
| 2026-09-04 | 799 | 72.6% | 0.1898 | 0.2017 | 6.5% | 22 |
| 2026-09-03 | 1,384 | 67.8% | 0.2077 | 0.2084 | 4.5% | 25 |
Pick weight retraining
Every night the pick-selection weights (probability, bucket calibration, sport performance) are refit on the last 30 days of graded results. The first two thirds train, the last third is held out for scoring. The probability models themselves are unchanged.
Active weights
prob 0.80 · calib 0.13 · perf 0.07
Adopted in last 45 days: 1
Last refit
2026-10-03
Improvement below threshold · previous weights kept
train 8,556 · eval 571 · top 15
| Metric | Before | After | Change |
|---|---|---|---|
| Top-pick hit rate | 68.4% | 68.4% | 0.0%p |
| Top-pick Brier | 0.1955 | 0.1955 | 0.0000 |
| Date | Status | Hit before | Hit after | Brier before | Brier after |
|---|---|---|---|---|---|
| 2026-10-03 | kept | 68.4% | 68.4% | 0.1955 | 0.1955 |
| 2026-10-02 | kept | 68.9% | 68.9% | 0.1996 | 0.1996 |
| 2026-09-22 | adopted | 60.0% | 64.4% | 0.2154 | 0.1770 |
| 2026-09-21 | kept | 71.1% | 71.1% | 0.1731 | 0.1731 |
| 2026-09-20 | kept | 77.8% | 77.8% | 0.1612 | 0.1612 |
Scored 2026-09-26 to 2026-10-02 · 1 finished games · n = 90 markets · stored Oct 3 (Sat) · 8:00 AM UTC
Overall summary
2026-09-26 ~ 2026-10-02 · 1 finished games · 90 marketsBrier score
0.2066
0.25 = blindly guessing 50%
Accuracy
70.0%
Directional at 50% threshold
Log loss
0.599
0.693 = coin flip
ECE
10.2%
Calibration error
Equal-mass ECE (ACE) 7.0% — fixed 10-bin ECE distorts when predictions cluster.
Better than a coin flip, but not by much.
Calibration effect (before → after)
- Brier 0.2047 → 0.2066
- ECE 6.7% → 10.2%
- Accuracy 70.0% → 70.0%
- Mean predicted 36.8% → 29.7% · actual rate 36.7%
Deep validation is Pro
Summary metrics and the first calibration curve are free. Per-market curves, walk-forward validation, the daily table, and windows longer than 7 days are computed for Pro.
See Pro$7.99/mo or $59/yr · cancel anytime
Combo pick record
Loading…
NFL featured combo record
Loading…
This report evaluates only model snapshots saved before each game. Post-game recalculations and older samples of unknown origin are excluded from training and official evaluation. It is a model evaluation, separate from the published locked-pick record, and results may be empty early in a new collection period.