The Ultimate Fight Fan

Under the hood · Walk-forward honest

Moneyline models

The winner model: a walk-forward backtest on the most recent 3033 UFC fights. Every prediction below was trained on data strictly older than the fight it predicted.

Model scoreboard
Walk-forward out-of-sample on 3033 fights. Lower log-loss is better; AUC and accuracy are higher-is-better. ECE is the average reliability-bin miscalibration.
Accuracy
64.6%
Log-loss
0.6313
AUC
0.695
Brier
0.2208
ModelFamilyFeaturesn OOSLog-lossAccuracyBrierAUCECE
Ensembleensemblediff30330.632164.5%0.22100.69640.0290
Logistic Elo Onlylogisticelo_only30330.677157.4%0.24210.60070.0181
Histgb Diffhistgbdiff30330.644162.4%0.22650.67510.0199
Logistic L1Best LLlogisticdiff30330.631364.6%0.22080.69490.0293
Calibration
Predicted probability vs observed win rate, deciles of the served logistic. Diagonal = perfectly calibrated.
EnsembleServed logisticPerfect calibration
Feature importance
Logistic-regression magnitude of standardized coefficients. Right = pushes prediction toward fighter A; left = toward B.
Steady Elo (long-memory)
-0.441
Upset avg
-0.404
Pound-for-pound Elo edge
0.333
Perf raw career
0.316
Age (late-career decline)
-0.302
Recent activity (fights)
0.250
Ring-rust-adjusted Elo
0.210
Recent opponent quality
0.182
UFC experience edge
-0.134
Submission attempts (per 15 min)
-0.121
Recent strikes absorbed (per min)
-0.118
Reach edge in heavy divisions
0.116
Preufc beaten wpct
0.108
Grappling-offense rating edge
0.106
Recent form (win %)
0.101
Takedowns (per 15 min)
0.086
Perf raw recent3
0.079
Win/loss-streak edge
0.078
Partial dependence
How the model's mean prediction shifts when one feature moves, holding others at their data distribution.
Per-segment performance
Where the model is strong, weak, or unstable. Choose a slice:
Skill gapnLog-lossAccuracyAUC
<2516960.653961.0%0.6534
25-7511940.610168.3%0.7344
75-1501430.555972.7%0.7933

Want to run this model on a matchup of your own? Interactive predictor — it runs in your browser, no API call.