Robot AI -- Model Evaluation

Memory tricks that make precision, recall, F1 and AUC click

From the confusion matrix to cross-validation -- these memory tricks lock in every model evaluation metric, when to use each one, and how to honestly report your results.

Community
📋
Model Evaluation Forum
Ask questions · Share tricks
💬
Model Evaluation Study Room
Live · Study together now

Or continue to the sub-topics below for more specialized Study Rooms and Forums

Model Evaluation

Memory Tricks

Proven Mnemonics & Acronyms — fast to learn, hard to forget.

🎥 How Flashcards Work
A quick walkthrough of tap-to-flip, rating, and how card colors track what you're struggling with.
← Back Next →
Model Evaluation deck1 of 20
Tap to flip
← →
How well do YOU think you know this?
Easy Medium Hard Harder
Tap to flip back
Model Evaluation deck
Easy0
Medium0
Hard0
Harder0
Confusion Matrix
CONFUSE the matrix -- TP, FP, FN, TN are the four boxes everything else comes from
TRUE POSITIVE AND FALSE POSITIVE AND FALSE NEGATIVE AND TRUE NEGATIVE
Draw the matrix -- predicted on columns, actual on rows. Diagonal = correct predictions.
True Positive (TP): predicted positive, actually positive. False Positive (FP): predicted positive, actually negative -- Type I error, false alarm. False Negative (FN): predicted negative, actually positive -- Type II error, missed detection. True Negative (TN): predicted negative, actually negative. Every classification metric -- accuracy, precision, recall, F1, AUC -- derives from these four numbers.
TP
Predicted positive, actually positive -- correct detection
FP
Predicted positive, actually negative -- false alarm (Type I error)
FN
Predicted negative, actually positive -- missed detection (Type II error)
TN
Predicted negative, actually negative -- correct rejection
🎥 Watch Instead
▶
Video coming soon
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 Confusion Matrix
Confusion matrix — what are TP, FP, FN and TN?
Tap to flip
🃏 Answer
CONFUSE the matrix -- TP, FP, FN, TN are the four boxes everything else comes from
TPPredicted positive, actually positive -- correct detection
FPPredicted positive, actually negative -- false alarm (Type I error)
FNPredicted negative, actually positive -- missed detection (Type II error)
TNPredicted negative, actually negative -- correct rejection
Tap to flip back
Precision and Recall
Precision is PICKY -- Recall REMEMBERS to catch everything
PICKY FOR PRECISION AND REMEMBERS FOR RECALL -- IMPOSSIBLE TO FORGET
P for Picky -- few false alarms. R for Remembers -- few misses. Two words for life.
Precision is PICKY: only says yes when really sure -- few false alarms. TP divided by (TP + FP). Recall REMEMBERS everything: sweeps wide to catch every real positive -- few misses. TP divided by (TP + FN). Tradeoff: raising classification threshold raises precision but lowers recall. Spam filter: be PICKY (don't block real email). Cancer screening: REMEMBER everything (don't miss a cancer).
Precision = TP/(TP+FP)
Picky -- minimize false alarms
Recall = TP/(TP+FN)
Remembers -- minimize missed detections
Tradeoff
Higher threshold = more precise but misses more
Cancer screening
Prioritize recall -- missing a cancer is catastrophic
🎥 Watch Instead
▶
Video coming soon
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 Precision and Recall
Precision vs recall — the formulas, and when to favor each?
Tap to flip
🃏 Answer
Precision is PICKY -- Recall REMEMBERS to catch everything
Precision = TP/(TP+FP)Picky -- minimize false alarms
Recall = TP/(TP+FN)Remembers -- minimize missed detections
TradeoffHigher threshold = more precise but misses more
Cancer screeningPrioritize recall -- missing a cancer is catastrophic
Tap to flip back
F1 Score
F1 = Harmonic mean of Precision and Recall -- punishes extreme imbalance between the two
F1 = 2 TIMES (P TIMES R) DIVIDED BY (P + R)
Harmonic mean is low if EITHER value is low -- you cannot compensate low recall with high precision
F1 combines precision and recall into one metric. Harmonic mean punishes extreme imbalance -- you cannot hide 10% recall behind 99% precision. F1 = 1.0 is perfect, 0 is worst. Use when class imbalance exists and you care about both FP and FN. F-beta: beta greater than 1 emphasizes recall (medical). Beta less than 1 emphasizes precision.
Harmonic mean
Low if EITHER P or R is low -- no hiding bad recall
Use F1 when
Class imbalance exists and both FP and FN matter
F-beta > 1
Emphasizes recall -- medical diagnosis
F-beta < 1
Emphasizes precision -- spam filtering
🎥 Watch Instead
▶
Video coming soon
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 F1 Score
F1 score — what is it, and when do you use it?
Tap to flip
🃏 Answer
F1 = Harmonic mean of Precision and Recall -- punishes extreme imbalance between the two
Harmonic meanLow if EITHER P or R is low -- no hiding bad recall
Use F1 whenClass imbalance exists and both FP and FN matter
F-beta > 1Emphasizes recall -- medical diagnosis
F-beta < 1Emphasizes precision -- spam filtering
Tap to flip back
ROC and AUC
ROC curve plots TPR vs FPR at every threshold -- AUC = area under that curve
AUC 0.5 = RANDOM AND AUC 1.0 = PERFECT AND AUC 0.8+ = GOOD
AUC = probability that model ranks a random positive higher than a random negative
ROC plots True Positive Rate (recall) vs False Positive Rate at every classification threshold. AUC = 0.5: no better than random. AUC = 1.0: perfect. AUC = 0.8+: good. Threshold-independent -- evaluates ranking ability. Use PR curve instead of ROC when positive class is very rare -- ROC can be optimistic with severe class imbalance.
TPR (Y-axis)
True Positive Rate = Recall = TP/(TP+FN)
FPR (X-axis)
False Positive Rate = FP/(FP+TN)
AUC meaning
Probability that model ranks a random positive above a random negative
PR curve
Better than ROC when positive class is very rare
🎥 Watch Instead
▶
Video coming soon
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 ROC and AUC
ROC curve and AUC — what do they show?
Tap to flip
🃏 Answer
ROC curve plots TPR vs FPR at every threshold -- AUC = area under that curve
TPR (Y-axis)True Positive Rate = Recall = TP/(TP+FN)
FPR (X-axis)False Positive Rate = FP/(FP+TN)
AUC meaningProbability that model ranks a random positive above a random negative
PR curveBetter than ROC when positive class is very rare
Tap to flip back
Regression Metrics
MAE in original units -- MSE squares the error -- RMSE back to original units -- R-squared explains variance
FOUR WAYS TO MEASURE REGRESSION ERROR
R-squared can be negative -- model is worse than simply predicting the mean for every example
MAE (Mean Absolute Error): average of absolute differences -- robust to outliers, original units. MSE: average of squared differences -- penalizes large errors, not in original units. RMSE: square root of MSE -- back in original units, most commonly reported. R-squared: proportion of variance explained -- 1.0 perfect, 0 no better than mean, can be negative.
MAE
Average absolute error -- robust to outliers, original units
MSE
Average squared error -- penalizes large errors heavily
RMSE
Square root of MSE -- back in original units
R-squared
Variance explained: 1=perfect, 0=no better than mean, negative=worse
🎥 Watch Instead
▶
Video coming soon
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 Regression Metrics
MAE, MSE, RMSE, R² — how do they differ?
Tap to flip
🃏 Answer
MAE in original units -- MSE squares the error -- RMSE back to original units -- R-squared explains variance
MAEAverage absolute error -- robust to outliers, original units
MSEAverage squared error -- penalizes large errors heavily
RMSESquare root of MSE -- back in original units
R-squaredVariance explained: 1=perfect, 0=no better than mean, negative=worse
Tap to flip back
Cross-Validation
K-FOLD -- K times: train on K-1 folds, test on 1 fold, rotate, average all scores
MORE RELIABLE THAN A SINGLE TRAIN/TEST SPLIT
Golden rule: NEVER use the test set for anything except final evaluation -- no tuning on test
Split data into K equal folds. For each fold: train on K-1 folds, evaluate on remaining fold. Average performance across all K folds. 5-fold and 10-fold standard. Stratified K-fold: ensures each fold has same class distribution -- important for imbalanced data. Time series: always split chronologically -- never random.
5-fold standard
5 rounds: each fold serves as test set once
Stratified K-fold
Each fold maintains class distribution -- use for imbalanced data
Time series CV
Walk-forward validation -- always train on past, test on future
Golden rule
Never use test set for anything except final evaluation
🎥 Watch Instead
▶
Video coming soon
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 Cross-Validation
K-fold cross-validation — how does it work?
Tap to flip
🃏 Answer
K-FOLD -- K times: train on K-1 folds, test on 1 fold, rotate, average all scores
5-fold standard5 rounds: each fold serves as test set once
Stratified K-foldEach fold maintains class distribution -- use for imbalanced data
Time series CVWalk-forward validation -- always train on past, test on future
Golden ruleNever use test set for anything except final evaluation
Tap to flip back
Class Imbalance
SMOTE -- Synthetic Minority Oversampling TEchnique -- create synthetic rare-class examples
WHEN 99% IS ONE CLASS -- ACCURACY IS USELESS
Apply resampling to training data ONLY -- never to test set
A model that always predicts the majority class gets 99% accuracy but 0% recall on minority -- useless. SMOTE: creates synthetic minority examples by interpolating between existing ones. Class weights: tell algorithm to penalize minority class errors more (class_weight=balanced). Use F1 or AUC, not accuracy, for evaluation.
Accuracy trap
99% accuracy with all-majority predictor = useless
SMOTE
Synthetic minority oversampling -- apply to training data ONLY
Class weights
class_weight=balanced -- equivalent to oversampling
Metrics to use
F1, AUC-ROC -- not accuracy with imbalanced data
🎥 Watch Instead
▶
Video coming soon
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 Class Imbalance
Class imbalance — what is SMOTE?
Tap to flip
🃏 Answer
SMOTE -- Synthetic Minority Oversampling TEchnique -- create synthetic rare-class examples
Accuracy trap99% accuracy with all-majority predictor = useless
SMOTESynthetic minority oversampling -- apply to training data ONLY
Class weightsclass_weight=balanced -- equivalent to oversampling
Metrics to useF1, AUC-ROC -- not accuracy with imbalanced data
Tap to flip back
Data Leakage
LEAK = test set information contaminating training -- model looks great until it hits the real world
THE MOST COMMON REASON ML MODELS FAIL IN PRODUCTION
Always split FIRST then transform -- never fit transformers on the full dataset
Data leakage: information from outside training set contaminates the model. Sources: fitting scaler on full dataset before splitting (never do this), including features that wouldn't be available at inference time, using future information for historical prediction. Detection: unrealistically high performance, feature importances that don't make causal sense, dramatic performance drop in production.
Train-test contamination
Fitting transformers on full dataset before splitting
Target leakage
Feature correlated with target only because of how target was defined
Temporal leakage
Using future data to predict past
Split FIRST
Always split data BEFORE fitting any transformers
🎥 Watch Instead
▶
Video coming soon
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 Data Leakage
Data leakage — what is it, and why is it dangerous?
Tap to flip
🃏 Answer
LEAK = test set information contaminating training -- model looks great until it hits the real world
Train-test contaminationFitting transformers on full dataset before splitting
Target leakageFeature correlated with target only because of how target was defined
Temporal leakageUsing future data to predict past
Split FIRSTAlways split data BEFORE fitting any transformers
Tap to flip back
Calibration
A CALIBRATED model -- when it says 80% confident, it is right 80% of the time
HIGH ACCURACY DOES NOT EQUAL WELL CALIBRATED
Modern neural networks are typically overconfident -- temperature scaling is the fix
Calibration measures whether predicted probabilities match actual frequencies. Reliability diagram: predicted probability vs actual frequency -- diagonal line = perfect calibration. Overconfident: predictions cluster near 0 and 1. Underconfident: cluster near 0.5. Temperature scaling: divide logits by T before softmax -- T greater than 1 softens probabilities. Critical for medical AI, financial risk, any application where you act on confidence.
Perfect calibration
When model says 80%, it is correct 80% of the time
Reliability diagram
Plot predicted probability vs actual frequency -- diagonal = perfect
Overconfidence
Common in neural networks -- predictions too extreme
Temperature scaling
Divide logits by T before softmax -- T>1 softens overconfidence
🎥 Watch Instead
▶
Video coming soon
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 Calibration
Calibration — what does a well-calibrated model mean?
Tap to flip
🃏 Answer
A CALIBRATED model -- when it says 80% confident, it is right 80% of the time
Perfect calibrationWhen model says 80%, it is correct 80% of the time
Reliability diagramPlot predicted probability vs actual frequency -- diagonal = perfect
OverconfidenceCommon in neural networks -- predictions too extreme
Temperature scalingDivide logits by T before softmax -- T>1 softens overconfidence
Tap to flip back
MLOps
MLOps = DevOps + Data + Models -- CI/CD for machine learning systems
DATA DRIFT AND CONCEPT DRIFT AND MODEL MONITORING AND RETRAINING
Model code is 5% of ML system code -- the other 95% is pipelines, monitoring, and infrastructure
MLOps applies software engineering to ML systems. Model monitoring: track data drift (input distribution changes), concept drift (label relationship changes), model performance degradation. Version control: code (Git), data (DVC), models (MLflow). CI/CD pipelines: automatically retrain and test models when data or code changes. Feature stores: centralized reusable feature computation.
Data drift
Input feature distribution changes over time
Concept drift
Relationship between features and labels changes
Model monitoring
Track data drift, performance, and prediction distributions
Version control
Track code, data, and model versions -- enable rollback
🎥 Watch Instead
▶
Video coming soon
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 MLOps
MLOps — what is it, and what does it monitor?
Tap to flip
🃏 Answer
MLOps = DevOps + Data + Models -- CI/CD for machine learning systems
Data driftInput feature distribution changes over time
Concept driftRelationship between features and labels changes
Model monitoringTrack data drift, performance, and prediction distributions
Version controlTrack code, data, and model versions -- enable rollback
Tap to flip back
A/B Testing
A/B test = controlled experiment -- one change, random assignment, measure the right metric
NEVER STOP EARLY -- WAIT FOR THE FULL PRE-SPECIFIED DURATION
A/B testing compares two versions by randomly assigning users and measuring outcomes
Split live traffic: 50% current model (A), 50% new model (B). Measure business metric (conversion, revenue) not just ML metric (accuracy). Key principles: random assignment (eliminates selection bias), single change (isolate cause), sufficient sample size, pre-specified duration. Common mistake: stopping test when p<0.05 appears -- leads to inflated false positive rates.
Random assignment
Eliminates selection bias -- essential for valid comparison
Single change
Isolate exactly what is causing any difference
Pre-specified duration
Never stop early -- peeking inflates false positive rate
Business metric
Measure what actually matters -- not just ML accuracy
🎥 Watch Instead
▶
Video coming soon
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 A/B Testing
A/B testing — the rules of a good test?
Tap to flip
🃏 Answer
A/B test = controlled experiment -- one change, random assignment, measure the right metric
Random assignmentEliminates selection bias -- essential for valid comparison
Single changeIsolate exactly what is causing any difference
Pre-specified durationNever stop early -- peeking inflates false positive rate
Business metricMeasure what actually matters -- not just ML accuracy
Tap to flip back
🎯 Exam Favorite
PRECISION = Of all I PREDICTED positive, how many were right? (The sniper — only fires when sure)
TRUE POSITIVES / ALL PREDICTED POSITIVES
Precision — the sniper metric
Precision measures how trustworthy your positive predictions are. Formula: TP / (TP + FP). If your spam filter flags 100 emails as spam and 90 are actually spam, precision = 90%. High precision = when you say "spam," you're almost always right. Low precision = lots of false alarms. Use precision when false positives are costly — like flagging innocent people as criminals or legitimate emails as spam.
🎥 Watch Instead
▶
Video coming soon
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 🎯 Exam Favorite
Precision — the sniper
Tap to flip
🃏 Answer
PRECISION = Of all I PREDICTED positive, how many were right? (The sniper — only fires when sure)
TRUE POSITIVES / ALL PREDICTED POSITIVES
Tap to flip back
🧠 Vivid Story
RECALL = Of all ACTUAL positives, how many did I catch? (The net — tries to catch everything)
TRUE POSITIVES / ALL ACTUAL POSITIVES
Recall — the fishing net metric
Recall (sensitivity) measures how many actual positives you caught. Formula: TP / (TP + FN). If 100 patients actually have cancer and your model catches 90, recall = 90%. High recall = you miss very few real cases. Low recall = dangerous missed detections. Use recall when false negatives are costly — missing a cancer diagnosis, missing a fraud transaction. The sniper (precision) vs the net (recall) — choose based on which error costs more.
🎥 Watch Instead
▶
Video coming soon
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 🧠 Vivid Story
Recall — the fishing net
Tap to flip
🃏 Answer
RECALL = Of all ACTUAL positives, how many did I catch? (The net — tries to catch everything)
TRUE POSITIVES / ALL ACTUAL POSITIVES
Tap to flip back
🔑 Key Distinction
F1 SCORE = The HARMONIC MEAN of precision and recall — punishes extreme imbalance
2 × (P × R) / (P + R)
F1 Score — when you need both precision and recall
F1 is the harmonic mean of precision and recall: 2×(P×R)/(P+R). If precision=1.0 and recall=0.0, F1=0 — no credit for being perfect at one while failing the other. F1 punishes extreme imbalance between the two. Use F1 when both false positives and false negatives matter and classes are imbalanced. A model that predicts "no cancer" for everyone gets 0% recall — F1 catches this while raw accuracy might look fine.
🎥 Watch Instead
▶
Video coming soon
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 🔑 Key Distinction
F1 score — why does it punish imbalance?
Tap to flip
🃏 Answer
F1 SCORE = The HARMONIC MEAN of precision and recall — punishes extreme imbalance
F1 is the harmonic mean of precision and recall: 2×(P×R)/(P+R). If precision=1.0 and recall=0.0, F1=0 — no credit for being perfect at one while failing the other. F1 punishes extreme imbalance between the two. Use F1 when both false positives and false negatives matter and classes are imbalanced. A model that predicts "no cancer" for everyone gets 0% recall — F1 catches this while raw accuracy might look fine.
Tap to flip back
💡 Concept Anchor
ROC CURVE = Plotting the TRADEOFF between catching criminals and arresting innocents
TRUE POSITIVE RATE vs FALSE POSITIVE RATE
ROC curve and AUC — comparing classifiers fairly
The ROC curve plots True Positive Rate (recall) on the Y-axis vs False Positive Rate on the X-axis at every possible threshold. A perfect classifier hugs the top-left corner (high TPR, low FPR). A random classifier is a diagonal line. AUC (Area Under the Curve) summarizes the whole curve in one number: AUC=1.0 is perfect, AUC=0.5 is random guessing. AUC lets you compare classifiers regardless of threshold — it answers "how good is this model overall?"
🎥 Watch Instead
▶
Video coming soon
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 💡 Concept Anchor
ROC curve — what tradeoff does it plot?
Tap to flip
🃏 Answer
ROC CURVE = Plotting the TRADEOFF between catching criminals and arresting innocents
TRUE POSITIVE RATE vs FALSE POSITIVE RATE
Tap to flip back
📅 Quick Reference
CONFUSION MATRIX = A 2×2 REPORT CARD — TP, FP, FN, TN tell the whole story
FOUR CELLS, FOUR OUTCOMES
The confusion matrix — reading all four cells
The confusion matrix shows all four prediction outcomes: True Positive (predicted yes, was yes ✓), False Positive (predicted yes, was no ✗ — Type I error), False Negative (predicted no, was yes ✗ — Type II error), True Negative (predicted no, was no ✓). All evaluation metrics — precision, recall, F1, accuracy — are calculated from these four numbers. Memorize the layout: TP is top-left when positive is the "interesting" class.
🎥 Watch Instead
▶
Video coming soon
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 📅 Quick Reference
Confusion matrix — what are the four cells?
Tap to flip
🃏 Answer
CONFUSION MATRIX = A 2×2 REPORT CARD — TP, FP, FN, TN tell the whole story
FOUR CELLS, FOUR OUTCOMES
Tap to flip back
⭐ Most Important
ACCURACY LIES on imbalanced data — always check Precision, Recall, and F1
THE ACCURACY TRAP — MOST TESTED EVALUATION PITFALL
Why accuracy is misleading and what to use instead
If 99% of emails are not spam, a model that predicts "not spam" for everything achieves 99% accuracy — and catches zero spam. Accuracy hides this completely. Use: Precision (when false positives are costly — spam filter flagging good emails), Recall (when false negatives are costly — missing cancer diagnoses), F1 (when both matter and classes are imbalanced). On balanced datasets accuracy is fine. On real-world imbalanced data, always report all three.
Accuracy trap
99% accuracy on 99:1 imbalanced data — useless model looks perfect
Use Precision
When false alarms are costly — spam filter, fraud alerts, legal flags
Use Recall
When missed cases are costly — cancer screening, fraud detection, safety
Use F1
When both false positives and negatives matter equally
🐍 Code
from sklearn.metrics import classification_report
print(classification_report(y_test, y_pred))
# Shows precision, recall, F1 for EACH class
# Always run this on imbalanced datasets — never just accuracy
🎥 Watch Instead
▶
Video coming soon
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 ⭐ Most Important
Why is accuracy misleading on imbalanced data?
Tap to flip
🃏 Answer
ACCURACY LIES on imbalanced data — always check Precision, Recall, and F1
Accuracy trap99% accuracy on 99:1 imbalanced data — useless model looks perfect
Use PrecisionWhen false alarms are costly — spam filter, fraud alerts, legal flags
Use RecallWhen missed cases are costly — cancer screening, fraud detection, safety
Use F1When both false positives and negatives matter equally
🐍 Codefrom sklearn.metrics import classification_report
print(classification_report(y_test, y_pred))
# Shows precision, recall, F1 for EACH class
# Always run this on imbalanced datasets — never just accuracy
Tap to flip back
🎯 Exam Favorite
MSE SQUARES errors — big mistakes PUNISHED MORE · MAE treats all errors EQUALLY
REGRESSION METRICS — MATCH TO YOUR BUSINESS PROBLEM
MSE vs MAE vs RMSE — which regression metric to use when
MSE (Mean Squared Error): squares each error, so large errors are penalized much more than small ones. Use when large errors are especially bad (predicting structural safety). MAE (Mean Absolute Error): average of absolute errors — all errors weighted equally, robust to outliers. RMSE: square root of MSE — back in original units like MAE but still penalizes large errors. R² (coefficient of determination): 1.0 = perfect, 0.0 = model no better than predicting the mean, can be negative.
MSE
Mean squared error — penalizes large errors heavily, sensitive to outliers
MAE
Mean absolute error — robust to outliers, all errors weighted equally
RMSE
Square root of MSE — interpretable in original units
R²
Proportion of variance explained — 1.0=perfect, 0=baseline
🐍 Code
from sklearn.metrics import mean_squared_error, mean_absolute_error, r2_score
mse = mean_squared_error(y_test, y_pred)
mae = mean_absolute_error(y_test, y_pred)
rmse = mean_squared_error(y_test, y_pred, squared=False)
r2 = r2_score(y_test, y_pred)
🎥 Watch Instead
▶
Video coming soon
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 🎯 Exam Favorite
MSE vs MAE — how does each treat errors?
Tap to flip
🃏 Answer
MSE SQUARES errors — big mistakes PUNISHED MORE · MAE treats all errors EQUALLY
MSEMean squared error — penalizes large errors heavily, sensitive to outliers
MAEMean absolute error — robust to outliers, all errors weighted equally
RMSESquare root of MSE — interpretable in original units
R²Proportion of variance explained — 1.0=perfect, 0=baseline
🐍 Codefrom sklearn.metrics import mean_squared_error, mean_absolute_error, r2_score
mse = mean_squared_error(y_test, y_pred)
mae = mean_absolute_error(y_test, y_pred)
rmse = mean_squared_error(y_test, y_pred, squared=False)
r2 = r2_score(y_test, y_pred)
Tap to flip back
🔑 Key Distinction
STRATIFIED K-FOLD preserves CLASS RATIOS in every fold — essential for imbalanced data
REGULAR K-FOLD vs STRATIFIED K-FOLD
When and why to use stratified cross-validation
Regular K-fold splits data randomly. With imbalanced classes (e.g., 5% positive), some folds might have 0% or 15% positives by chance — giving wildly variable and unreliable scores. Stratified K-fold ensures each fold has the same class proportions as the full dataset. Always use stratified K-fold for classification, especially with class imbalance. For regression, regular K-fold is fine. Rule: if you are classifying, use StratifiedKFold.
Regular K-fold
Random splits — class proportions may vary widely between folds
Stratified K-fold
Preserves class proportions in each fold — essential for classification
🐍 Code
from sklearn.model_selection import StratifiedKFold, cross_val_score
skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
scores = cross_val_score(model, X, y, cv=skf, scoring="f1")
print(f"F1: {scores.mean():.3f} ± {scores.std():.3f}")
🎥 Watch Instead
▶
Video coming soon
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 🔑 Key Distinction
Stratified K-fold — what does it preserve, and when?
Tap to flip
🃏 Answer
STRATIFIED K-FOLD preserves CLASS RATIOS in every fold — essential for imbalanced data
Regular K-foldRandom splits — class proportions may vary widely between folds
Stratified K-foldPreserves class proportions in each fold — essential for classification
🐍 Codefrom sklearn.model_selection import StratifiedKFold, cross_val_score
skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
scores = cross_val_score(model, X, y, cv=skf, scoring="f1")
print(f"F1: {scores.mean():.3f} ± {scores.std():.3f}")
Tap to flip back
💡 Concept Anchor
CALIBRATION = When the model says 80% confident, it should be RIGHT 80% of the time
PREDICTED PROBABILITY vs ACTUAL FREQUENCY
Why probability calibration matters for real decisions
A model might say "70% probability of rain" — but does it actually rain 70% of the time in those cases? A well-calibrated model's predicted probabilities match real-world frequencies. Poorly calibrated: model says 90% confident but is only right 60% of the time — dangerous for medical diagnosis or financial risk. Check with reliability diagrams (calibration curves). Fix with Platt scaling (logistic regression on top) or isotonic regression.
Well calibrated
Predicted 0.8 probability → actually happens ~80% of the time
Overconfident
Predicted 0.9 but actually only 60% — dangerous for high-stakes decisions
Fix with
Platt scaling or isotonic regression post-hoc calibration
🐍 Code
from sklearn.calibration import CalibratedClassifierCV, calibration_curve
calibrated = CalibratedClassifierCV(base_model, method="sigmoid")
calibrated.fit(X_train, y_train)
fraction_pos, mean_pred = calibration_curve(y_test, y_prob, n_bins=10)
🎥 Watch Instead
▶
Video coming soon
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 💡 Concept Anchor
Calibration — what should an 80% prediction mean?
Tap to flip
🃏 Answer
CALIBRATION = When the model says 80% confident, it should be RIGHT 80% of the time
Well calibratedPredicted 0.8 probability → actually happens ~80% of the time
OverconfidentPredicted 0.9 but actually only 60% — dangerous for high-stakes decisions
Fix withPlatt scaling or isotonic regression post-hoc calibration
🐍 Codefrom sklearn.calibration import CalibratedClassifierCV, calibration_curve
calibrated = CalibratedClassifierCV(base_model, method="sigmoid")
calibrated.fit(X_train, y_train)
fraction_pos, mean_pred = calibration_curve(y_test, y_prob, n_bins=10)
Tap to flip back
🎓 Common Exam Questions
0
Correct
0
Wrong
0
Remaining
🔗 Related Sub-Subjects
📊 Machine Learning
The ML training pipeline — cross-validation, train/test splits, and overfitting connect directly here.
Machine Learning →
⚙️ AI Algorithms
Decision trees, SVMs, K-NN — the algorithms whose performance you are evaluating.
AI Algorithms →
⚖️ AI Ethics
Fairness metrics, bias detection, and the ethical implications of evaluation choices.
AI Ethics →