Why do you need a baseline model? The two lines that expose your 94%
The dumb answers every model must beat, the two-line sklearn baselines, the class-imbalance accuracy trap, and how to read the gap that is your actual value: one page.
Get the free PDF
One page, print-ready, free to share. No signup needed.
A 94% model can still be useless. Before celebrating any score, run the dumb answer through the same split and look at the gap: that gap is your entire contribution. One page on baselines, the accuracy trap, and reporting honestly. The print-ready A4 PDF is at the bottom.
The rule
- Score the dumb answer before any model.
- Your value = model score minus baseline score, same split, same metric.
- No gap? Ship the rule, not the model.
The dumb answers
- Classification: predict the majority class.
- Regression: predict the mean (or median).
- Time series: tomorrow = today, or seasonal naive (= same day last week). Brutal to beat.
The two-line reality check
from sklearn.dummy import DummyClassifier
dumb = DummyClassifier(strategy="most_frequent")
cross_val_score(dumb, X, y, cv=5).mean()
# 0.94 <- fraud is 6% of rows
cross_val_score(model, X, y, cv=5).mean()
# 0.94 <- the model learned nothing
Same 0.94, zero value. Switch the scoring to recall or PR-AUC on imbalanced data: the dummy collapses to 0 there and the comparison becomes honest.
The accuracy trap
- 94% legit rows makes 94% accuracy free.
- Accuracy hides it; precision and recall expose it.
- PR-AUC is the imbalance-proof summary score.
Better baselines
- The current rule: whatever ops does today. Beating the dummy but losing to the incumbent ships nothing.
- One tiny model: logistic regression on 3 features. If the big model wins by a rounding error, ship the small one.
- Seasonal naive for anything with a weekly rhythm.
The trap: which baseline for which task
| Task | Baseline | Beats more models than |
|---|---|---|
| churn, fraud, spam | majority class + recall | you would hope |
| price, demand | mean, or median | linear reg on bad features |
| forecasting | same day last week | most first Prophets |
Gotchas
- Run the baseline in the same CV split, or the comparison is fake.
- Report both numbers, always: the model AND the baseline.
- 0.94 vs a 0.93 baseline is one point. Say whether one point is worth the infra.
- Baselines drift too: re-score them at every retrain.
Interview phrasing worth memorizing: never present a model score alone. Present the gap over a named baseline, on the same split, in the metric the business feels.
Frequently asked questions
What is a baseline model?
Why can a model with 94% accuracy be useless?
How do you build a baseline in scikit-learn?
What baseline should you use besides a dummy model?
Get the free PDF
One page, print-ready, free to share. No signup needed.