Machine Learning Models Performance Evaluation
Datasets for evaluation
The story: A driving student practices on the same streets every week, takes mock tests on new routes while still learning, and finally takes the real driving test on a route they've never driven.
In AI/AWS terms:
- Training set trains the model (the practice streets).
- Validation set checks generalization while you're still improving the model (the mock tests).
- Test set is the final check before production (the real driving test).
For the exam: Validation is for tuning during development. Test is the final check before production.
Model fit and the bullseye
The story: Four archers shoot at a target.
- One's arrows land tightly together, but far from the center.
- One's arrows are scattered all over the target.
- One is off-center and scattered.
- One lands tightly in the center.
In AI/AWS terms: Bias is how far the shots land from the center. Variance is how spread out they are. Compare errors on training data and evaluation data:
- Underfitting: poor even on training data (high bias).
- Overfitting: good on training data, poor on evaluation data. The model memorized instead of generalizing (high variance).
- Balanced: low bias and low variance, the archer hitting the center tightly.
For the exam: Bias = distance from the center. Variance = spread. The goal is low bias and low variance.
The confusion matrix
The story: A smoke alarm can do four things. It rings when there's a fire (right). It stays quiet when there's no fire (right). It rings when you're only making toast (false alarm). It stays quiet during a real fire (a miss, the dangerous one).
In AI/AWS terms: Those four outcomes are the confusion matrix:
- True positive (TP): rings for a real fire.
- True negative (TN): quiet when there's no fire.
- False positive (FP): rings for toast.
- False negative (FN): quiet during a real fire.
For the exam: False positive = a false alarm. False negative = a miss.
Classification metrics
The story: Judging smoke alarms:
- "How often was it right overall?" sounds good, but an alarm that never rings is "right" on every quiet day. It looks great and is useless.
- "When it rings, how often is there really a fire?" matters if false alarms are expensive, like evacuating a whole hospital.
- "Of all the real fires, how many did it catch?" matters most when a miss can kill.
- Sometimes you need a balance of both.
- And you can compare alarms at every sensitivity setting to pick the best alarm and setting.
In AI/AWS terms:
| Question | Metric | Formula | Use it when |
|---|---|---|---|
| Right overall? | Accuracy | (TP + TN) / all predictions | Classes are balanced. Misleading when there are many true negatives |
| When it rings, is it real? | Precision | TP / (TP + FP) | False positives are costly, like a spam filter hiding real email |
| Did it catch every fire? | Recall (sensitivity) | TP / (TP + FN) | False negatives are costly, like missing a serious illness |
| Both | F1 score | Harmonic mean of precision and recall | You need a balance of precision and recall |
| Every setting | AUC-ROC | Area under the true positive rate versus false positive rate curve across thresholds | Comparing models and choosing a threshold |
For the exam: Costly false alarms → precision. Costly misses → recall. Balance → F1. Imbalanced data makes accuracy misleading.
Regression metrics
The story: A weather forecaster predicts tomorrow's temperature. You can score them by how far off they usually are, punishing big misses heavily. Or you can ask how much of the day-to-day ups and downs in temperature they manage to explain.
In AI/AWS terms:
- Mean squared error (MSE): the average of squared prediction errors, so big misses count extra. Lower is better.
- R squared: the share of variance the model explains, from 0 to 1. Closer to 1 is better.
For the exam: MSE: lower is better. R squared: closer to 1 is better.
Business metrics
The story: A new oven in a bakery can be technically excellent, but the owner judges it by whether it sells more bread or cuts costs. Some mistakes (a slightly pale loaf) are fine; others (burning the wedding cake) are very costly. Before switching every oven, the owner tries the new one in one shop first, or on a few batches.
In AI/AWS terms:
- Tie model metrics to the KPIs set during business goal identification, such as more sales, lower costs, or less churn.
- Check that the metrics reflect the business's tolerance for errors, and consider a cost function for the economic impact of each kind of error.
- Compare model variants in production with A/B testing or canary deployments (trying it in one shop first).
For the exam: Tie model metrics to business KPIs. Compare variants in production with A/B testing or canary deployments.