How to choose an ML algorithm in production
There is no best algorithm. Here’s how teams actually pick one.
The idea
Teams don’t pick a model from a leaderboard. They start by exploring the data: how many rows, which columns, what’s missing, how balanced the labels are. Then they pin down the goal (which mistake is more expensive?) and the constraints: does someone need an explanation, how fast must it answer, how often will it retrain?
Only then do they compare a few candidates fairly against a simple baseline, and pick the simplest model that meets the goal. Often that’s not the fanciest one.
What the lesson will build
- Exploratory data analysis on a customer churn dataset
- A baseline that shows why accuracy misleads on imbalanced data
- Logistic regression, a decision tree and gradient boosting compared with cross-validation
- A decision you can defend
Key ideas
- Exploratory data analysis
- Choosing the metric from the business goal
- Constraints: explainability, latency, cost
- Baselines and cross-validation
The video
- Long form: A path from data to decision: EDA, goal, constraints, baseline, candidates, choice, with real scores growing in as bars.
- Short: No best algorithm. Just the right one for your data and goal.
The lesson is ready: code and a line-by-line walkthrough in machine-learning/11-choosing-a-model/.