Robot AI -- Machine Learning

Memory tricks that make machine learning algorithms click

From supervised to reinforcement learning -- these memory tricks lock in every major ML algorithm, the bias-variance tradeoff, and the workflow that makes models actually work.

Community
📋
Machine Learning Forum
Ask questions · Share tricks
💬
Machine Learning Study Room
Live · Study together now

Or continue to the sub-topics below for more specialized Study Rooms and Forums

Machine Learning

Memory Tricks

Proven Mnemonics & Acronyms — fast to learn, hard to forget.

🎥 How Flashcards Work
A quick walkthrough of tap-to-flip, rating, and how card colors track what you're struggling with.
← Back Next →
Machine Learning deck1 of 21
Tap to flip
← →
How well do YOU think you know this?
Easy Medium Hard Harder
Tap to flip back
Machine Learning deck
Easy0
Medium0
Hard0
Harder0
🎯 Exam Favorite
Supervised = Memorization. Unsupervised = Searching for Patterns and Deviations.
THE CORE DISTINCTION — NEVER MIX THEM UP AGAIN
Supervised vs Unsupervised learning in one sentence
Supervised learning is memorization: every training example has a known correct answer (a label) attached, and the model memorizes the relationship between question and answer — like a flashcard with the answer on the back. Unsupervised learning is a search for patterns and deviations: no known answers exist, so the model finds natural groupings on its own and flags anything that deviates from them. This matters beyond just cost — supervised learning can only ever replicate patterns humans already identified, while unsupervised learning can uncover genuinely new patterns nobody has found yet, giving it real potential to solve mysteries and support proving theoretical ideas once a pattern shows up in real data. Supervised's edge is precise, measurable accuracy since the answers are known; the tradeoff is the cost of labeling every example, and it can only ever recognize categories it was explicitly trained on.
🎥 Watch Instead
Supervised vs Unsupervised Learning — with AI-Mee — 4:58.
Flashcard
🃏 🎯 Exam Favorite
Supervised vs unsupervised learning?
Tap to flip
🃏 Answer
Supervised = Memorization. Unsupervised = Searching for Patterns and Deviations.
THE CORE DISTINCTION — NEVER MIX THEM UP AGAIN
Tap to flip back
🧠 Vivid Story
Gradient Descent = Rolling a BALL DOWN A HILL to find the lowest point
THE HILL METAPHOR — HOW MODELS ACTUALLY LEARN
Gradient descent explained with one unforgettable image
Imagine you're blindfolded on a hilly landscape. Your goal is to reach the lowest valley (minimum loss). You feel the slope under your feet and take a small step downhill. Repeat thousands of times. That's gradient descent — the model checks which direction reduces error, takes a small step that direction, and repeats until it can't go lower. Learning rate = how big each step is. Too big = overshoot. Too small = takes forever. It's also the same underlying process nearly every ML model uses to learn, from simple linear regression up to massive neural networks.
🎥 Watch Instead
Gradient Descent — with AI-Mee — 3:02.
Flashcard
🃏 🧠 Vivid Story
Gradient descent — how do models learn?
Tap to flip
🃏 Answer
Gradient Descent = Rolling a BALL DOWN A HILL to find the lowest point
THE HILL METAPHOR — HOW MODELS ACTUALLY LEARN
Tap to flip back
Star Most Tested
SPLIT -- Select, Prepare, Learn, Interpret, Test -- the 5-step ML workflow
THE ML PROJECT WORKFLOW
Every ML project follows this repeatable process -- know it cold
Select the right algorithm for your problem type. Prepare data: clean, normalize, handle missing values, encode categoricals, split train/test. Learn: train the model -- algorithm finds optimal parameters. Interpret: examine what was learned. Test: evaluate on held-out test set -- never seen during training. Performance on test data is the only honest measure of model quality.
Select
Match algorithm to problem type (classification, regression, clustering)
Prepare
~80% of ML work -- clean, normalize, encode, split
Learn
Train on training set -- find optimal parameters
Test
Held-out test set -- the only honest performance measure
🎥 Watch Instead
▶
Video coming soon
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 Star Most Tested
The ML workflow (SPLIT) — the five steps?
Tap to flip
🃏 Answer
SPLIT -- Select, Prepare, Learn, Interpret, Test -- the 5-step ML workflow
SelectMatch algorithm to problem type (classification, regression, clustering)
Prepare~80% of ML work -- clean, normalize, encode, split
LearnTrain on training set -- find optimal parameters
TestHeld-out test set -- the only honest performance measure
Tap to flip back
Regression
LINE -- Linear regression predicts a NUMBER. Logistic predicts a CATEGORY.
THE NAMING TRAP -- LOGISTIC REGRESSION IS A CLASSIFIER
Linear predicts a number. Logistic predicts a category. Despite the name.
Linear Regression: y = mx + b, minimizes squared errors, predicts continuous values (house price, temperature). Logistic Regression: despite the name, used for CLASSIFICATION -- uses sigmoid function to output probability 0-1. Decision boundary at 0.5. Above = positive class. Below = negative class. This distinction appears on virtually every ML exam.
Linear Regression
Continuous output -- house price, temperature, salary
Logistic Regression
CLASSIFICATION despite the name -- outputs probability 0-1
Sigmoid function
Squashes to 0-1: s(z) = 1/(1+e^-z)
Decision boundary
0.5 threshold -- above=positive, below=negative
🎥 Watch Instead
▶
Video coming soon
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 Regression
Linear vs logistic regression — what does each predict?
Tap to flip
🃏 Answer
LINE -- Linear regression predicts a NUMBER. Logistic predicts a CATEGORY.
Linear RegressionContinuous output -- house price, temperature, salary
Logistic RegressionCLASSIFICATION despite the name -- outputs probability 0-1
Sigmoid functionSquashes to 0-1: s(z) = 1/(1+e^-z)
Decision boundary0.5 threshold -- above=positive, below=negative
Tap to flip back
Decision Trees
TREE -- Test feature, Recurse on subsets, End at leaf, Ensemble to improve
DECISION TREE to RANDOM FOREST to XGBOOST
Single trees overfit. Ensembles are among the most powerful ML algorithms.
Decision Trees split data by feature questions using Gini impurity or Information Gain. Prone to overfitting alone. Random Forest (bagging): parallel trees on random data and feature subsets -- reduces variance. XGBoost (boosting): sequential trees, each corrects previous errors -- dominant for tabular data competitions. No feature scaling needed for tree-based methods.
Gini impurity
Probability of misclassifying a random sample -- used to pick splits
Random Forest
Bagging -- parallel independent trees, majority vote
XGBoost
Boosting -- sequential, each tree corrects prior errors
No scaling needed
Tree-based methods are scale-invariant
🎥 Watch Instead
▶
Video coming soon
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 Decision Trees
Decision trees and ensembles — how do they work?
Tap to flip
🃏 Answer
TREE -- Test feature, Recurse on subsets, End at leaf, Ensemble to improve
Gini impurityProbability of misclassifying a random sample -- used to pick splits
Random ForestBagging -- parallel independent trees, majority vote
XGBoostBoosting -- sequential, each tree corrects prior errors
No scaling neededTree-based methods are scale-invariant
Tap to flip back
K-Nearest Neighbors
K-NN -- K Neighbors vote, No training needed, Nearest wins
LAZY LEARNER -- ALL COMPUTATION AT PREDICTION TIME
K-NN has no training phase -- it simply memorizes the training set
Find K closest training examples (Euclidean distance), take majority vote (classification) or average (regression). Always normalize -- large-range features dominate distance. Curse of dimensionality: too many features makes distance meaningless. Choose K with cross-validation. K=1 overfits. Large K underfits.
Always normalize
Distance is meaningless without feature scaling
Curse of dimensionality
Many features = sparse space = nearest neighbor meaningless
K=1
Overfits -- very sensitive to noise
Choose K
Cross-validate -- odd K for binary classification avoids ties
🎥 Watch Instead
▶
Video coming soon
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 K-Nearest Neighbors
K-NN — how does it classify, and why is it a lazy learner?
Tap to flip
🃏 Answer
K-NN -- K Neighbors vote, No training needed, Nearest wins
Always normalizeDistance is meaningless without feature scaling
Curse of dimensionalityMany features = sparse space = nearest neighbor meaningless
K=1Overfits -- very sensitive to noise
Choose KCross-validate -- odd K for binary classification avoids ties
Tap to flip back
SVM
SVM = finds the WIDEST STREET between classes
MAXIMUM MARGIN CLASSIFIER
Support vectors are the training points closest to the boundary -- they define everything
SVM finds the hyperplane that maximally separates classes. Wider margin = better generalization. Kernel trick: maps data to higher dimensions where linearly separable -- RBF kernel handles nonlinear data. C parameter: low C = wider margin, more misclassifications tolerated. High C = narrow margin, fewer misclassifications. Best for high-dimensional sparse data.
Maximum margin
Wider margin = better generalization to new data
Kernel trick
Map to higher dimension where data IS linearly separable
C parameter
Low C = wide margin (tolerant). High C = narrow margin (strict).
Best for
Text classification and high-dimensional sparse data
🎥 Watch Instead
▶
Video coming soon
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 SVM
SVM — the maximum margin and the kernel trick?
Tap to flip
🃏 Answer
SVM = finds the WIDEST STREET between classes
Maximum marginWider margin = better generalization to new data
Kernel trickMap to higher dimension where data IS linearly separable
C parameterLow C = wide margin (tolerant). High C = narrow margin (strict).
Best forText classification and high-dimensional sparse data
Tap to flip back
K-Means Clustering
K-MEANS -- K clusters, Move centroids, Assign nearest, Repeat until stable
UNSUPERVISED CLUSTERING -- GROUP WITHOUT LABELS
Use the elbow method to find optimal K -- plot inertia vs K, pick the bend
Steps: choose K, randomly initialize K centroids, assign each point to nearest centroid, recalculate centroids as cluster means, repeat until stable. Elbow method: plot inertia vs K, choose where improvement slows. DBSCAN alternative: finds arbitrary-shaped clusters and outliers automatically without specifying K.
Elbow method
Plot inertia vs K -- choose where improvement slows
Weakness
Needs K in advance, sensitive to outliers, assumes spherical clusters
DBSCAN
No K needed, finds arbitrary shapes, detects outliers
Scale first
K-Means uses distance -- normalize all features
🎥 Watch Instead
▶
Video coming soon
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 K-Means Clustering
K-means — the steps, and how to pick K?
Tap to flip
🃏 Answer
K-MEANS -- K clusters, Move centroids, Assign nearest, Repeat until stable
Elbow methodPlot inertia vs K -- choose where improvement slows
WeaknessNeeds K in advance, sensitive to outliers, assumes spherical clusters
DBSCANNo K needed, finds arbitrary shapes, detects outliers
Scale firstK-Means uses distance -- normalize all features
Tap to flip back
Feature Engineering
SCALE before you TRAIN -- unscaled features ruin distance-based algorithms
NORMALIZE and STANDARDIZE and ENCODE and REDUCE
Always split data BEFORE fitting any transformers to prevent data leakage
Normalization (Min-Max): scales to 0-1. Standardization (Z-score): mean=0, std=1. Scale for: K-NN, SVM, neural networks, PCA. NOT for: tree-based methods. One-hot encoding: categoricals to binary columns. Label encoding: integers to categories -- only when ordinal. PCA: reduce dimensions by keeping directions of maximum variance. Golden rule: fit transformers on training data only.
Scale for
K-NN, SVM, neural networks, PCA, logistic regression
Skip scaling for
Decision trees, Random Forest, XGBoost -- scale-invariant
One-hot encoding
Red/Green/Blue becomes 3 binary columns
PCA
Keep maximum variance directions -- reduce noise and computation
🎥 Watch Instead
▶
Video coming soon
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 Feature Engineering
Feature scaling — which algorithms need it?
Tap to flip
🃏 Answer
SCALE before you TRAIN -- unscaled features ruin distance-based algorithms
Scale forK-NN, SVM, neural networks, PCA, logistic regression
Skip scaling forDecision trees, Random Forest, XGBoost -- scale-invariant
One-hot encodingRed/Green/Blue becomes 3 binary columns
PCAKeep maximum variance directions -- reduce noise and computation
Tap to flip back
Naive Bayes
NAIVE = assumes all features INDEPENDENT -- wrong but works surprisingly well
BAYES THEOREM WITH AN INDEPENDENCE ASSUMPTION -- FAST AND EFFECTIVE
For text classification the independence assumption holds approximately -- that's enough
P(class|features) proportional to P(class) times product of P(feature|class). Despite independence assumption being almost always violated, it works well for text -- word frequencies are roughly independent given class. Gaussian NB: continuous features. Multinomial NB: word counts. Laplace smoothing: add 1 to all counts to prevent zero probabilities.
Gaussian NB
Assumes continuous features follow Gaussian distribution
Multinomial NB
For word counts -- text classification, spam detection
Laplace smoothing
Add 1 to counts -- prevents zero probability killing the product
Why it works
For text, word independence is approximately true
🎥 Watch Instead
▶
Video coming soon
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 Naive Bayes
Naive Bayes — what's naive about it?
Tap to flip
🃏 Answer
NAIVE = assumes all features INDEPENDENT -- wrong but works surprisingly well
Gaussian NBAssumes continuous features follow Gaussian distribution
Multinomial NBFor word counts -- text classification, spam detection
Laplace smoothingAdd 1 to counts -- prevents zero probability killing the product
Why it worksFor text, word independence is approximately true
Tap to flip back
Reinforcement Learning
AGENT SPAR -- State, Policy, Action, Reward -- maximize cumulative reward
TRIAL ERROR AND REWARDS -- THE THIRD ML PARADIGM
RLHF is how ChatGPT and Claude are aligned -- human feedback trains a reward model
Agent: learner. State: current situation. Action: choice made. Reward: feedback signal. Policy: strategy mapping states to actions. Goal: maximize expected cumulative reward. Exploration vs exploitation. Q-learning: learns value of state-action pairs. RLHF (Reinforcement Learning from Human Feedback): human raters rank outputs, reward model trained, RL maximizes it -- standard for aligning LLMs.
Agent
Learner that takes actions and receives rewards
Policy
Strategy: given state s, what action a to take
Q-learning
Learns Q(s,a) = expected cumulative reward from state s taking action a
RLHF
Human rankings + reward model + RL = aligned ChatGPT/Claude
🎥 Watch Instead
▶
Video coming soon
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 Reinforcement Learning
Reinforcement learning — agent, state, policy, action, reward?
Tap to flip
🃏 Answer
AGENT SPAR -- State, Policy, Action, Reward -- maximize cumulative reward
AgentLearner that takes actions and receives rewards
PolicyStrategy: given state s, what action a to take
Q-learningLearns Q(s,a) = expected cumulative reward from state s taking action a
RLHFHuman rankings + reward model + RL = aligned ChatGPT/Claude
Tap to flip back
Time Series
TASC -- Trend, Autocorrelation, Seasonality, Cycle
NEVER SHUFFLE TIME SERIES DATA -- ALWAYS SPLIT CHRONOLOGICALLY
Future data cannot predict the past -- chronological order must be preserved
Trend: long-term direction. Seasonality: regular repeating patterns (daily, weekly, yearly). Autocorrelation: current value correlates with its own past values. Cycle: irregular multi-year patterns. NEVER shuffle before splitting -- always chronological. Classic models: ARIMA. Modern: LSTMs, Temporal CNNs, Transformer-based forecasters.
Trend
Long-term direction (upward, downward, flat)
Seasonality
Regular repeating patterns -- daily, weekly, yearly
Autocorrelation
Current value correlates with past values (lag-1, lag-2...)
Never shuffle
Always split chronologically -- future cannot predict past
🎥 Watch Instead
▶
Video coming soon
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 Time Series
Time series — the four components (TASC)?
Tap to flip
🃏 Answer
TASC -- Trend, Autocorrelation, Seasonality, Cycle
TrendLong-term direction (upward, downward, flat)
SeasonalityRegular repeating patterns -- daily, weekly, yearly
AutocorrelationCurrent value correlates with past values (lag-1, lag-2...)
Never shuffleAlways split chronologically -- future cannot predict past
Tap to flip back
Anomaly Detection
RARE events are ISOLATED -- Isolation Forest finds anomalies with fewer random splits
DETECT THE UNUSUAL WITHOUT LABELED ANOMALY EXAMPLES
Train on normal data only -- anomalies stand out by being hard to reconstruct
Isolation Forest: anomalies isolated by fewer random splits (short path = anomaly) -- fast, scalable, no distance metric needed. Statistical: z-score, flag beyond N standard deviations. Autoencoder: train on normal data only, flag high reconstruction error points as anomalies. Applications: fraud detection, network intrusion, manufacturing defects, medical outliers.
Isolation Forest
Short path to isolate = anomaly. Fast, scalable.
Z-score
Simple -- flag points beyond 2-3 standard deviations
Autoencoder
Train on normal data -- anomalies cannot be reconstructed well
Applications
Fraud, network intrusion, defects, medical outliers
🎥 Watch Instead
▶
Video coming soon
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 Anomaly Detection
Anomaly detection — the main methods?
Tap to flip
🃏 Answer
RARE events are ISOLATED -- Isolation Forest finds anomalies with fewer random splits
Isolation ForestShort path to isolate = anomaly. Fast, scalable.
Z-scoreSimple -- flag points beyond 2-3 standard deviations
AutoencoderTrain on normal data -- anomalies cannot be reconstructed well
ApplicationsFraud, network intrusion, defects, medical outliers
Tap to flip back
AutoML
AutoML = automate algorithm selection, hyperparameter tuning, feature engineering
AI BUILDING AI -- NEURAL ARCHITECTURE SEARCH FOUND EFFICIENTNET
AutoML automates the tuning, not the thinking -- still need good problem framing
AutoML tools: Google AutoML (cloud, no-code), H2O AutoML (open-source), Auto-sklearn. NAS (Neural Architecture Search): discovered EfficientNet -- outperformed human-designed networks. Active Learning: model queries the most uncertain examples for labeling -- reduces labeling cost by 5-10x. Limitations: still requires clean data, correct evaluation metric, good problem definition.
Google AutoML
Cloud-based, no-code, requires little ML expertise
H2O AutoML
Open-source, competitive with cloud options on tabular data
NAS
Found EfficientNet -- better than human-designed CNN architectures
Active Learning
Query uncertain examples -- 5-10x more label-efficient
🎥 Watch Instead
▶
Video coming soon
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 AutoML
AutoML — what does it automate?
Tap to flip
🃏 Answer
AutoML = automate algorithm selection, hyperparameter tuning, feature engineering
Google AutoMLCloud-based, no-code, requires little ML expertise
H2O AutoMLOpen-source, competitive with cloud options on tabular data
NASFound EfficientNet -- better than human-designed CNN architectures
Active LearningQuery uncertain examples -- 5-10x more label-efficient
Tap to flip back
🔑 Key Distinction
REWARD DOG — Reinforcement Learning is a dog learning tricks for treats
AGENT · ENVIRONMENT · REWARD — the three RL components
Reinforcement learning — the reward dog story
Reinforcement learning has three parts: an Agent (the dog/AI), an Environment (the world it acts in), and a Reward signal (treat or punishment). The agent takes actions, receives rewards or penalties, and learns which actions maximize total reward over time. No labels needed — just trial, error, and feedback. This is how AlphaGo beat world chess champions and how robots learn to walk.
🎥 Watch Instead
▶
Video coming soon
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 🔑 Key Distinction
Reinforcement learning — the three components?
Tap to flip
🃏 Answer
REWARD DOG — Reinforcement Learning is a dog learning tricks for treats
AGENT · ENVIRONMENT · REWARD — the three RL components
Tap to flip back
💡 Concept Anchor
CROSS-VALIDATION = Taking the SAME EXAM multiple times with different questions each round
K-FOLD — ROTATE THE TEST SET EVERY ROUND
K-fold cross-validation — why one test isn't enough
K-fold cross-validation splits data into K equal chunks. In each round, one chunk is the test set and the rest train the model. Rotate K times so every chunk gets tested once. Average the K scores for a reliable estimate. Why? One test split might be lucky or unlucky. K-fold gives a fair average across many splits. K=5 or K=10 are standard. Think of it as rotating who grades your exam to get a fair average score.
🎥 Watch Instead
▶
Video coming soon
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 💡 Concept Anchor
Cross-validation — how does K-fold work?
Tap to flip
🃏 Answer
CROSS-VALIDATION = Taking the SAME EXAM multiple times with different questions each round
K-FOLD — ROTATE THE TEST SET EVERY ROUND
Tap to flip back
📅 Quick Reference
FOREST BEATS TREE — Ensemble methods usually outperform single models
WISDOM OF CROWDS IN MACHINE LEARNING
Why ensemble methods like Random Forest win
A single decision tree can overfit badly. A Random Forest grows hundreds of trees on random data subsets and averages their predictions. More diverse opinions = more stable result. This is ensemble learning — combining many weak learners into one strong learner. Bagging (Random Forest) averages parallel models. Boosting (XGBoost) trains models sequentially, each fixing the last one's errors. Either way: the forest always beats the tree.
🎥 Watch Instead
▶
Video coming soon
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 📅 Quick Reference
Ensembles — why does a forest usually beat a tree?
Tap to flip
🃏 Answer
FOREST BEATS TREE — Ensemble methods usually outperform single models
WISDOM OF CROWDS IN MACHINE LEARNING
Tap to flip back
⭐ Most Important
BULLSEYE — Low Bias + Low Variance = dead center every shot
BIAS = AIM · VARIANCE = SPREAD
Bias-Variance Tradeoff — the most tested ML theory concept
Picture a dartboard. Bias = how far from the bullseye your average shot lands (systematic error — wrong aim). Variance = how spread out your shots are (inconsistency). High Bias = underfitting — model too simple, misses the pattern every time. High Variance = overfitting — model too complex, nails training data but wild on new data. The goal: low bias AND low variance. Increasing model complexity reduces bias but raises variance — the unavoidable tradeoff.
High Bias
Underfitting — model too simple, misses real patterns (linear model on curved data)
High Variance
Overfitting — model memorizes training noise, fails on new data
Sweet spot
Right complexity — generalizes well to unseen data
🐍 Code
from sklearn.model_selection import validation_curve
train_scores, val_scores = validation_curve(model, X, y, param_name='max_depth', param_range=range(1,20))
# Plot to see the bias-variance tradeoff visually
🎥 Watch Instead
▶
Video coming soon
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 ⭐ Most Important
Bias vs variance — what does the bullseye show?
Tap to flip
🃏 Answer
BULLSEYE — Low Bias + Low Variance = dead center every shot
High BiasUnderfitting — model too simple, misses real patterns (linear model on curved data)
High VarianceOverfitting — model memorizes training noise, fails on new data
Sweet spotRight complexity — generalizes well to unseen data
🐍 Codefrom sklearn.model_selection import validation_curve
train_scores, val_scores = validation_curve(model, X, y, param_name='max_depth', param_range=range(1,20))
# Plot to see the bias-variance tradeoff visually
Tap to flip back
🎯 Exam Favorite
OVERFIT = Memorized the TEXTBOOK · UNDERFIT = Didn't study at all
TRAINING ERROR vs VALIDATION ERROR GAP
Overfitting vs Underfitting — spotted by the training/validation gap
Overfitting: training accuracy very high, validation accuracy much lower — the model memorized the training set including its noise. Like a student who memorized every practice exam verbatim but can't handle new questions. Underfitting: both training AND validation accuracy are low — the model hasn't learned the pattern at all. The diagnostic: plot learning curves. A big gap between training and validation = overfitting. Both curves low and flat = underfitting.
Overfitting fix
More data, regularization (L1/L2), dropout, simpler model, pruning
Underfitting fix
More complex model, more features, longer training, remove regularization
Diagnostic
Learning curve: plot train vs validation score against training set size
🐍 Code
from sklearn.model_selection import learning_curve
train_sizes, train_scores, val_scores = learning_curve(model, X, y, cv=5)
# Big gap between curves = overfitting. Both low = underfitting.
🎥 Watch Instead
▶
Video coming soon
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 🎯 Exam Favorite
Overfitting vs underfitting — signs and fixes?
Tap to flip
🃏 Answer
OVERFIT = Memorized the TEXTBOOK · UNDERFIT = Didn't study at all
Overfitting fixMore data, regularization (L1/L2), dropout, simpler model, pruning
Underfitting fixMore complex model, more features, longer training, remove regularization
DiagnosticLearning curve: plot train vs validation score against training set size
🐍 Codefrom sklearn.model_selection import learning_curve
train_sizes, train_scores, val_scores = learning_curve(model, X, y, cv=5)
# Big gap between curves = overfitting. Both low = underfitting.
Tap to flip back
🔑 Key Distinction
BAGGING = PARALLEL trees vote · BOOSTING = SEQUENTIAL trees correct each other
RANDOM FOREST vs XGBOOST — know both cold
Bagging vs Boosting — ensemble methods compared
Bagging (Bootstrap Aggregating): trains many models IN PARALLEL on random data subsets, averages their predictions. Reduces variance. Random Forest is the gold standard. Boosting: trains models SEQUENTIALLY — each new model focuses on examples the previous one got wrong. Reduces bias. XGBoost, LightGBM, and AdaBoost are the most powerful implementations. Rule of thumb: Random Forest when you need speed and reliability. XGBoost when you need maximum accuracy on tabular data.
Bagging
Parallel, reduces variance, Random Forest — fast, robust, hard to overfit
Boosting
Sequential, reduces bias, XGBoost — highest accuracy on tabular data
🐍 Bagging code
from sklearn.ensemble import RandomForestClassifier
rf = RandomForestClassifier(n_estimators=100, random_state=42)
rf.fit(X_train, y_train)
🐍 Boosting code
from xgboost import XGBClassifier
xgb = XGBClassifier(n_estimators=100, learning_rate=0.1)
xgb.fit(X_train, y_train)
🎥 Watch Instead
▶
Video coming soon
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 🔑 Key Distinction
Bagging vs boosting — parallel or sequential?
Tap to flip
🃏 Answer
BAGGING = PARALLEL trees vote · BOOSTING = SEQUENTIAL trees correct each other
BaggingParallel, reduces variance, Random Forest — fast, robust, hard to overfit
BoostingSequential, reduces bias, XGBoost — highest accuracy on tabular data
🐍 Bagging codefrom sklearn.ensemble import RandomForestClassifier
rf = RandomForestClassifier(n_estimators=100, random_state=42)
rf.fit(X_train, y_train)
🐍 Boosting codefrom xgboost import XGBClassifier
xgb = XGBClassifier(n_estimators=100, learning_rate=0.1)
xgb.fit(X_train, y_train)
Tap to flip back
💡 Concept Anchor
L1 KILLS features · L2 SHRINKS features — regularization is a tax on complexity
LASSO vs RIDGE — two flavors of the same cure
L1 vs L2 Regularization — preventing overfitting by penalizing weights
Regularization adds a penalty to the loss function for large weights — taxing the model for complexity. L1 (Lasso): penalty = sum of absolute values of weights. Drives unimportant weights to exactly zero — built-in feature selection. L2 (Ridge): penalty = sum of squared weights. Shrinks all weights toward zero but rarely to exactly zero. Elastic Net combines both. Use L1 when you suspect only a few features matter. Use L2 when all features contribute something. Lambda controls the penalty strength.
L1 (Lasso)
Kills irrelevant features — sparse solution, automatic feature selection
L2 (Ridge)
Shrinks all features — smooth solution, no feature elimination
Elastic Net
L1 + L2 combined — best of both when unsure
🐍 Code
from sklearn.linear_model import Lasso, Ridge, ElasticNet
lasso = Lasso(alpha=0.1) # L1 — kills features
ridge = Ridge(alpha=1.0) # L2 — shrinks features
enet = ElasticNet(alpha=0.1, l1_ratio=0.5) # both
🎥 Watch Instead
▶
Video coming soon
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 💡 Concept Anchor
L1 vs L2 regularization — what does each do?
Tap to flip
🃏 Answer
L1 KILLS features · L2 SHRINKS features — regularization is a tax on complexity
L1 (Lasso)Kills irrelevant features — sparse solution, automatic feature selection
L2 (Ridge)Shrinks all features — smooth solution, no feature elimination
Elastic NetL1 + L2 combined — best of both when unsure
🐍 Codefrom sklearn.linear_model import Lasso, Ridge, ElasticNet
lasso = Lasso(alpha=0.1) # L1 — kills features
ridge = Ridge(alpha=1.0) # L2 — shrinks features
enet = ElasticNet(alpha=0.1, l1_ratio=0.5) # both
Tap to flip back
🎓 Common Exam Questions
0
Correct
0
Wrong
0
Remaining
🔗 Related Sub-Subjects
🕸️ Neural Networks
How neurons, layers, backpropagation, and activation functions build on ML fundamentals.
Neural Networks →
📈 Model Evaluation
Precision, recall, F1, confusion matrix, ROC-AUC — how to measure if your model is actually good.
Model Evaluation →
🔬 Deep Learning
CNNs, RNNs, Transformers — the deep neural network architectures powering modern AI.
Deep Learning →