NLP Pipeline
NLP PIPELINE -- Tokenize, Normalize, Represent, Model, Evaluate
RAW TEXT TO PREDICTIONS IN FIVE STEPS
Modern LLMs skip most preprocessing -- trained end-to-end with BPE subword tokenization
🎥 Watch Instead
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 NLP Pipeline
NLP pipeline — the five steps?
Tap to flip
🃏 Answer
NLP PIPELINE -- Tokenize, Normalize, Represent, Model, Evaluate
TokenizationSplit text into words, subwords, or characters
NormalizationLowercase, punctuation handling, contraction expansion
LemmatizationReturns actual dictionary base form (better than stemming)
Modern LLMsSkip most preprocessing -- end-to-end with BPE tokenization
Tap to flip back
Text Representation
BOW to TF-IDF to Word2Vec to BERT -- from counting words to understanding meaning
FOUR GENERATIONS OF TEXT REPRESENTATION -- EACH MORE POWERFUL
BERT gives same word different vectors based on context -- bank in finance vs bank of river
🎥 Watch Instead
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 Text Representation
Text representation — BoW to TF-IDF to Word2Vec to BERT?
Tap to flip
🃏 Answer
BOW to TF-IDF to Word2Vec to BERT -- from counting words to understanding meaning
BoWWord counts -- sparse, ignores order and context
TF-IDFWeights rare words higher -- still sparse but smarter
Word2VecDense semantic vectors -- arithmetic works, but context-free
BERT / contextualSame word, different vector based on context -- most powerful
Tap to flip back
BERT vs GPT
BERT reads the whole sentence (bidirectional). GPT reads left-to-right only.
ENCODER FOR UNDERSTANDING -- DECODER FOR GENERATION
BERT cannot generate text. GPT cannot see the right context. Both are excellent at what they do.
🎥 Watch Instead
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 BERT vs GPT
BERT vs GPT — how does each read text?
Tap to flip
🃏 Answer
BERT reads the whole sentence (bidirectional). GPT reads left-to-right only.
BERT encoder-onlyBidirectional, masked LM, understanding tasks: classification, NER, QA
GPT decoder-onlyLeft-context only, next-token prediction, generation tasks
T5 and BARTEncoder-Decoder -- translation, summarization, seq2seq tasks
Why GPT dominates nowScale + emergent few-shot abilities -- can do BERT tasks via prompting
Tap to flip back
Tokenization (BPE)
BPE -- Byte Pair Encoding: merge the most frequent character pair, repeat until vocabulary size reached
NO TRUE OUT-OF-VOCABULARY WORDS WITH SUBWORD TOKENIZATION
750 English words is approximately 1000 tokens -- other languages use more tokens per word
🎥 Watch Instead
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 Tokenization (BPE)
BPE tokenization — how does it build a vocabulary?
Tap to flip
🃏 Answer
BPE -- Byte Pair Encoding: merge the most frequent character pair, repeat until vocabulary size reached
BPE algorithmMerge most frequent character pairs iteratively until target vocab size
No OOV wordsAny word can be split into known subword units
Token counting750 English words ~ 1000 tokens. Other languages: more tokens per word.
Vocabulary sizeGPT models: 50K-100K token vocabulary
Tap to flip back
Prompt Engineering
CLEAR -- Context, Length, Examples, Ask specifically, Role assignment
BETTER PROMPT = BETTER OUTPUT -- PROMPT ENGINEERING IS A SKILL
Chain-of-Thought: adding Let's think step by step dramatically improves reasoning accuracy
🎥 Watch Instead
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 Prompt Engineering
Prompting (CLEAR) — the five elements?
Tap to flip
🃏 Answer
CLEAR -- Context, Length, Examples, Ask specifically, Role assignment
Zero-shotJust ask -- no examples needed for simple clear tasks
Few-shotGive 2-5 examples -- improves consistency and format
Chain-of-ThoughtLet's think step by step -- dramatic improvement on reasoning
RAGRetrieve documents, inject as context -- reduces hallucination
Tap to flip back
NLP Evaluation
BLEU scores translation -- ROUGE scores summaries -- Perplexity scores language models
DIFFERENT TASKS NEED DIFFERENT EVALUATION METRICS
BLEU measures n-gram overlap -- it misses semantic equivalents like automobile vs car
🎥 Watch Instead
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 NLP Evaluation
BLEU vs ROUGE vs perplexity — what does each score?
Tap to flip
🃏 Answer
BLEU scores translation -- ROUGE scores summaries -- Perplexity scores language models
BLEUN-gram overlap for translation -- misses semantic equivalents
ROUGEN-gram overlap for summarization -- ROUGE-1, ROUGE-2, ROUGE-L
PerplexityHow surprised LM is by test text -- lower = better LM
BERTScoreSemantic similarity via embeddings -- better than BLEU/ROUGE
Tap to flip back
Machine Translation
ENCODE source then ATTEND to relevant parts then DECODE to target language
FROM RULE-BASED TO STATISTICAL TO NEURAL TO TRANSFORMER
Parallel corpora are needed for training -- back-translation helps low-resource languages
🎥 Watch Instead
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 Machine Translation
Machine translation — how did it evolve?
Tap to flip
🃏 Answer
ENCODE source then ATTEND to relevant parts then DECODE to target language
Rule-basedHand-crafted grammar rules -- limited coverage
Statistical (SMT)Phrase-based, learns from parallel corpora
Neural (NMT)Encoder-decoder + attention -- first major DL NLP breakthrough
Transformer (2017)Surpassed RNN systems immediately -- now the standard
Tap to flip back
Speech Recognition
ASR = Acoustic Signal to Recognition to transcript -- convert audio to text
WHISPER IS OPEN SOURCE AND NEAR-HUMAN ACCURACY ON ENGLISH
WER = (Substitutions + Deletions + Insertions) / Total words -- lower is better
🎥 Watch Instead
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 Speech Recognition
ASR
Tap to flip
🃏 Answer
ASR = Acoustic Signal to Recognition to transcript -- convert audio to text
Traditional ASRAcoustic model + pronunciation dictionary + language model
End-to-end (CTC)Directly map audio features to character sequences
Whisper680K training hours, multilingual, near-human WER, open-source
WER formula(Substitutions + Deletions + Insertions) / Total reference words
Tap to flip back
Coreference
COREFERENCE -- John told Mary he liked her -- who is he and who is her?
IDENTIFY WHICH WORDS IN A TEXT REFER TO THE SAME REAL-WORLD ENTITY
Winograd Schema Challenge: The trophy does not fit because it is too big -- what is too big?
🎥 Watch Instead
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 Coreference
Coreference resolution — what does it solve?
Tap to flip
🃏 Answer
COREFERENCE -- John told Mary he liked her -- who is he and who is her?
PronounsHe, she, it, they -- must resolve to named entities
Noun phrasesThe president, the company -- track across long documents
Winograd SchemaTests common-sense reasoning for coreference resolution
ApplicationsReading comprehension, information extraction, summarization
Tap to flip back
Text Generation Decoding
GREEDY picks top token -- SAMPLING picks randomly -- TEMPERATURE controls creativity
TOP-P NUCLEUS SAMPLING IS THE STANDARD IN PRODUCTION LLM APPLICATIONS
Top-P (nucleus sampling): sample from smallest set of tokens whose cumulative probability exceeds P
🎥 Watch Instead
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 Text Generation Decoding
Decoding — greedy vs sampling vs temperature?
Tap to flip
🃏 Answer
GREEDY picks top token -- SAMPLING picks randomly -- TEMPERATURE controls creativity
GreedyAlways pick highest prob -- deterministic but repetitive
Temperature < 1Sharper distribution -- more conservative, predictable
Temperature > 1Flatter distribution -- more creative, sometimes incoherent
Top-P (nucleus)Most widely used -- adapts to actual distribution shape
Tap to flip back
🎯 Exam Favorite
TOKENIZATION = Chopping a SENTENCE into LEGO BRICKS before the model reads it
TEXT IN → TOKENS OUT → MODEL READS
Tokenization — the first step in all NLP
🎥 Watch Instead
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 🎯 Exam Favorite
Tokenization — what does it do?
Tap to flip
🃏 Answer
TOKENIZATION = Chopping a SENTENCE into LEGO BRICKS before the model reads it
TEXT IN → TOKENS OUT → MODEL READS
Tap to flip back
🧠 Vivid Story
WORD EMBEDDINGS = Giving every word an ADDRESS in a city where similar words live close together
KING − MAN + WOMAN ≈ QUEEN
Word embeddings — meaning as location in space
🎥 Watch Instead
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 🧠 Vivid Story
Word embeddings — the city of addresses
Tap to flip
🃏 Answer
WORD EMBEDDINGS = Giving every word an ADDRESS in a city where similar words live close together
KING − MAN + WOMAN ≈ QUEEN
Tap to flip back
🔑 Key Distinction
BERT reads BOTH DIRECTIONS — GPT reads LEFT TO RIGHT only
BIDIRECTIONAL vs AUTOREGRESSIVE
BERT vs GPT — two fundamentally different Transformer approaches
🎥 Watch Instead
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 🔑 Key Distinction
BERT vs GPT — which direction does each read?
Tap to flip
🃏 Answer
BERT reads BOTH DIRECTIONS — GPT reads LEFT TO RIGHT only
BIDIRECTIONAL vs AUTOREGRESSIVE
Tap to flip back
💡 Concept Anchor
NAMED ENTITY RECOGNITION = The model HIGHLIGHTS the nouns that ARE something — people, places, orgs
NER — FIND THE WHO, WHERE, AND WHAT
Named Entity Recognition (NER) — finding the important nouns
🎥 Watch Instead
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 💡 Concept Anchor
Named entity recognition — what does it find?
Tap to flip
🃏 Answer
NAMED ENTITY RECOGNITION = The model HIGHLIGHTS the nouns that ARE something — people, places, orgs
NER — FIND THE WHO, WHERE, AND WHAT
Tap to flip back
📅 Quick Reference
RAG = OPEN BOOK EXAM — the model looks things up before answering instead of relying on memory
RETRIEVE · AUGMENT · GENERATE
Retrieval-Augmented Generation — why it reduces hallucinations
🎥 Watch Instead
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 📅 Quick Reference
RAG — the open-book exam
Tap to flip
🃏 Answer
RAG = OPEN BOOK EXAM — the model looks things up before answering instead of relying on memory
RETRIEVE · AUGMENT · GENERATE
Tap to flip back
⭐ Most Important
TOKENIZE → EMBED → ENCODE → DECODE → GENERATE — the complete NLP pipeline
FIVE STAGES FROM RAW TEXT TO FINAL OUTPUT
The modern NLP pipeline end to end
🎥 Watch Instead
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 ⭐ Most Important
The NLP pipeline — from tokenize to generate?
Tap to flip
🃏 Answer
TOKENIZE → EMBED → ENCODE → DECODE → GENERATE — the complete NLP pipeline
TokenizeRaw text → token IDs using a learned vocabulary (BPE, WordPiece)
EmbedToken IDs → dense vectors + positional encoding
Encode/DecodeTransformer layers build or generate representations
GenerateSampling strategy converts logits to final token choices
🐍 Codefrom transformers import AutoTokenizer, AutoModel
tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")
tokens = tokenizer("Hello world", return_tensors="pt")
model = AutoModel.from_pretrained("bert-base-uncased")
outputs = model(**tokens) # contextual embeddings
Tap to flip back
🎯 Exam Favorite
BLEU measures PRECISION of n-grams · ROUGE measures RECALL of n-grams
BLEU FOR TRANSLATION · ROUGE FOR SUMMARIZATION
NLP evaluation metrics — BLEU vs ROUGE
🎥 Watch Instead
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 🎯 Exam Favorite
BLEU vs ROUGE — precision or recall?
Tap to flip
🃏 Answer
BLEU measures PRECISION of n-grams · ROUGE measures RECALL of n-grams
BLEUPrecision of n-grams — machine output vs reference. Used for translation.
ROUGERecall of n-grams — reference vs machine output. Used for summarization.
PerplexityHow surprised is the model by test text? Lower = better language model.
🐍 Codefrom nltk.translate.bleu_score import sentence_bleu
reference = [["the", "cat", "sat"]]; hypothesis = ["the", "cat", "sat"]
bleu = sentence_bleu(reference, hypothesis)
# For ROUGE: pip install rouge-score
from rouge_score import rouge_scorer
Tap to flip back
🔑 Key Distinction
SEMANTIC SEARCH = MEANING · KEYWORD SEARCH = EXACT WORDS — embeddings make semantic possible
SPARSE vs DENSE RETRIEVAL
Why semantic search beats keyword search for complex queries
🎥 Watch Instead
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 🔑 Key Distinction
Semantic vs keyword search?
Tap to flip
🃏 Answer
SEMANTIC SEARCH = MEANING · KEYWORD SEARCH = EXACT WORDS — embeddings make semantic possible
Keyword (sparse)Exact word matching — fast, interpretable, fails on synonyms
Semantic (dense)Embedding similarity — handles synonyms, paraphrases, concepts
HybridCombine both — best of precision (keywords) and recall (semantics)
🐍 Codefrom sentence_transformers import SentenceTransformer
model = SentenceTransformer("all-MiniLM-L6-v2")
embeddings = model.encode(["Heart attack", "myocardial infarction"])
from sklearn.metrics.pairwise import cosine_similarity
print(cosine_similarity([embeddings[0]], [embeddings[1]]))
Tap to flip back
💡 Concept Anchor
FINE-TUNING vs PROMPTING — two ways to customize an LLM for your task
WHEN TO FINE-TUNE · WHEN TO PROMPT
Choosing between fine-tuning and prompt engineering
🎥 Watch Instead
This lesson's animated video hasn't been made yet — check back soon.
Flashcard
🃏 💡 Concept Anchor
Fine-tuning vs prompting — when to use each?
Tap to flip
🃏 Answer
FINE-TUNING vs PROMPTING — two ways to customize an LLM for your task
PromptingFast, cheap, flexible — try this first for every task
Fine-tuningBetter consistency and performance — requires labeled data and compute
LoRAEfficient fine-tuning — update only 1% of parameters, get 99% of the benefit
🐍 Code# Prompting — no training required
response = client.messages.create(model="claude-sonnet-4-6",
messages=[{"role":"user","content":"Classify sentiment: I love this!"}])
# Fine-tuning with LoRA (parameter-efficient)
# from peft import LoraConfig, get_peft_model
Tap to flip back