AI — 300 free lessons

ml · neural nets · transformers · llms · agents · rag · eval

Start lesson 1 →

AI Tutorial

1. What is AI · The Field2. AI History · The Three Winters3. Symbolic vs Statistical · The Two Camps4. Machine Learning · Learning from Data

AI Learning Types

5. Supervised Learning · Labeled Data6. Unsupervised Learning · Finding Structure7. Reinforcement Learning · Reward Signals8. Self-Supervised Learning · The Modern Shift

Classical ML

9. Classification vs Regression10. Clustering · K-Means · Dbscan11. Dimensionality Reduction · Pca · T-Sne · Umap12. Linear Regression · The Simplest Model13. Logistic Regression · Binary Classifier14. Decision Trees · Gini · Entropy15. Random Forests · Ensemble of Trees16. Gradient Boosting · Xgboost · Lightgbm17. Svm · Support Vector Machines18. Naive Bayes · Probability Over Features19. Knn · K-Nearest Neighbors

Neural Networks

20. Neural Networks · The Perceptron Origin21. Activation Functions · Relu · Sigmoid · Tanh22. Loss Functions · Mse · Cross-Entropy23. Gradient Descent · The Optimizer24. Backpropagation · How Gradients Flow25. Learning Rate · The Most Important Knob26. Optimizers · Sgd · Adam · Adamw27. Regularization · L1 · L2 · Dropout28. Batch Norm · Layer Norm · Group Norm29. Overfitting · The Bias-Variance Trade

AI Training Methodology

30. Cross-Validation · K-Fold · Stratified31. Train · Val · Test Splits32. Data Augmentation · Flips · Crops · Mixup33. Feature Engineering · The Old Craft

AI Embeddings

34. Embeddings · Turning Things Into Vectors35. Word2vec · The Breakthrough36. Glove · Subword Embeddings

AI Architectures

37. Cnns · Convolutional Neural Networks38. Cnns · Filters · Pooling · Stride39. Rnns · Recurrence · Vanishing Gradients40. Lstms · Gates · Long-Term Memory41. Grus · Simpler Gated Recurrence42. Seq2seq · Encoder-Decoder43. Attention · The Seq2seq Fix44. Transformers · Attention is all you Need45. Self-Attention · Queries · Keys · Values46. Multi-Head Attention · Parallel Views47. Positional Encoding · Giving Order Back48. Encoder-Only · Bert Family49. Decoder-Only · GPT Family50. Encoder-Decoder · T5 · Bart

Tokens & Scale

51. Tokenization · Bpe · Wordpiece · Unigram52. Context Windows · How Much Fits at Once53. Scaling Laws · Params · Data · Compute54. Emergent Abilities · The Surprise Phase

Pretraining & Tuning

55. Pretraining · Next-Token Prediction56. Fine-Tuning · Adapting a Base Model57. Lora · Parameter-Efficient Tuning58. Rlhf · Reinforcement Learning from Human Feedback59. Dpo · Direct Preference Optimization60. Constitutional AI · Self-Critique Training61. Instruction Tuning · Making Models Follow

Prompting

62. Prompt Engineering · The New Programming63. Zero-Shot · One-Shot · Few-Shot64. Chain of Thought · Think Before Answering65. Tree of Thoughts · Branching Reasoning66. React · Reason + Act Loops67. Self-Consistency · Sample and Vote

RAG & Retrieval

68. Retrieval Augmented Generation · RAG Basics69. RAG · Vector Stores · Embeddings · Chunking70. RAG · Hybrid Search · Rerank

Agents

71. Agents · LLMS with Tools72. Tool Use · Function Calling73. Multi-Agent Systems · Roles + Handoffs74. Memory · Short Term · Long Term · Episodic

Evaluation

75. Evaluation · Benchmarks vs Vibes76. Benchmarks · Mmlu · Gsm8k · Humaneval77. Eval Harnesses · Lm-Eval · Helm · Big-Bench

Safety & Alignment

78. Hallucinations · Why Models Invent79. Groundedness · Citations · Attribution80. Red Teaming · Finding Failures on Purpose81. Jailbreaks · Prompt Injection · Countermeasures82. Safety · Refusals · Over-Refusals83. Alignment · The Meta-Objective

Inference

84. Inference · Greedy · Sampling · Temperature85. Top-K · Top-P · Repetition Penalty86. Speculative Decoding · Faster Inference87. Quantization · Int8 · Int4 · Ggml · Gguf88. Distillation · Teacher to Student89. Moe · Mixture of Experts

Multimodal

90. Multimodal · Vision + Language91. Clip · Contrastive Image-Text Training92. Image Generation · Diffusion Models93. Diffusion · Forward · Reverse · Sampling94. Controlnet · Guiding Generation95. Text-To-Speech · Vits · Xtts · Elevenlabs96. Speech Recognition · Whisper97. Video Generation · Sora · Veo98. Embeddings Models · OpenAI · Cohere · Bge99. Vector Dbs · Pinecone · Weaviate · Chroma

MLOps

100. Mlops · Pipelines · Monitoring · Drift101. Data Labeling · The Human Bottleneck102. Synthetic Data · Self-Instruct · Distill103. Context Caching · Cutting Cost104. Structured Outputs · JSON Mode · Constrained105. Function Calling · Tools as Schemas

Models & Computer Use

106. Computer Use · Operating the Desktop107. Claude · GPT · Gemini · Llama · The Lineup108. Open Weights vs Closed · The Divide109. Local Inference · Ollama · Llama.cpp · Lmstudio110. Fine-Tuning Hosted · OpenAI · Anthropic Platforms

Compute & Interpretability

111. Compute · Gpus · Tpus · Accelerators112. Interpretability · Features · Circuits113. Mechanistic Interpretability · Sae · Probing

AI Industry

114. AI Safety Institute · Evals · Frontier Rules115. Agi · The Moving Target116. Cost · Per-Token Economics117. Latency · Ttft · Tps · Streaming118. Carbon · The Energy Footprint119. Ethics · Bias · Privacy · Displacement

AI References

120. References · Papers with Code · Arxiv · Sota

Transformer Internals Deep

121. Attention Math · Scaled Dot-Product from Scratch122. The Kv Cache · Why Generation is Autoregressive123. Flash Attention · IO-Aware Exact Attention124. Rotary Embeddings · Rope · Relative Position125. Alibi · Attention with Linear Biases126. Grouped-Query Attention · Gqa · Mqa127. Rmsnorm vs Layernorm · Pre-Norm vs Post-Norm128. Swiglu · Gated Feed-Forward Networks129. The Residual Stream · The Model’s Working Memory130. Weight Tying · Input and Output Embeddings131. Attention Sinks · The First-Token Anchor132. Logit Soft-Capping · Taming the Output133. The Transformer Block · End to End134. Parameter Counting · Where the Weights Live

Attention & Long Context

135. Sliding-Window Attention · Local Context136. Sparse Attention · Longformer · Bigbird137. Linear Attention · Kernel Tricks138. Ring Attention · Context Across Devices139. Yarn · Ntk-Aware Context Extension140. Position Interpolation · Stretching the Window141. Lost in the Middle · Recency and Primacy Bias142. Retrieval Heads · How Models Recall from Context143. Streaming LLMS · Infinite-Context Attention144. Kv Cache Compression · H2o · Scissorhands145. Paged Attention · Virtual Memory for the Cache146. Prompt Caching · Reusing the Prefix

Pretraining at Scale

147. The Pretraining Data Pipeline · Crawl to Tokens148. Deduplication · Why Duplicates Hurt149. Data Mixtures · Domain Weighting150. Data Curriculum · Ordering the Corpus151. Tokenizer Training · Vocabulary Size Trade-Offs152. Chinchilla · Compute-Optimal Scaling153. Maximal Update Parametrization · MuP154. Learning-Rate Schedules · Warmup · Cosine · Wsd155. Gradient Clipping · Taming Exploding Grads156. Loss Spikes · Detection and Recovery157. Checkpointing · Resuming a Giant Run158. Training Stability · The Dark Art

Distributed Training

159. Data Parallelism · Replicate and All-Reduce160. Tensor Parallelism · Splitting the Matmul161. Pipeline Parallelism · Staged Layers162. Zero & Fsdp · Sharding Optimizer State163. 3D Parallelism · Combining the Three164. Gradient Accumulation · Big Batches on Small Gpus165. Mixed Precision · Bf16 · Fp16 · Loss Scaling166. Fp8 Training · The H100 Frontier167. Activation Checkpointing · Recompute to Save Memory168. Collective Communication · Nccl · All-Reduce …169. Megatron & Deepspeed · The Training Stacks170. Fault Tolerance · Surviving Node Failures

Post-Training Deep

171. Sft · Supervised Fine-Tuning Data Curation172. The Rlhf Pipeline · Sft → Rm → Ppo173. Reward Modeling · Learning Human Preferences174. Ppo for LLMS · The Policy-Gradient Core175. Dpo · Skipping the Reward Model176. Ipo · Kto · Preference-Optimization Variants177. Rejection Sampling · Best-Of-N Distillation178. Rlaif · AI Feedback at Scale179. Process Reward Models · Grading Each Step180. Rlvr · Reinforcement Learning from Verifiable…181. Online vs Offline Preference Learning182. Reward Hacking · Gaming the Proxy183. Length Bias · The Verbosity Trap184. Preference Data · Collection and Quality

Reasoning & Test-Time Compute

185. Chain of Thought · Reasoning in Tokens186. Self-Consistency · Sample and Majority-Vote187. Tree of Thoughts · Search Over Reasoning188. Best-Of-N · Sampling with a Verifier189. Verifier Models · Checking the Answer190. Process Supervision · Rewarding the Path191. O1-Style Reasoning · Long Internal Thought192. Test-Time Scaling Laws · Compute at Inference193. Mcts + LLM · Guided Search194. Self-Refine · Critique and Revise195. Multi-Agent Debate · Arguing to Truth196. Budget-Aware Reasoning · When to Stop Thinking

Agent Architectures Deep

197. The Agent Loop · Perceive · Plan · Act · Observe198. React Deep · Interleaving Reason and Action199. Plan-And-Execute · Decompose Then Run200. Reflexion · Learning from Failed Attempts201. The Model Context Protocol · Mcp Tools202. Function Calling Internals · Schemas to Calls203. Agent Memory Systems · Scratchpad to Vector Store204. Multi-Agent Orchestration · Supervisor + Workers205. Agent Routing · Picking the Right Specialist206. Environment Design · The Agent’s Sandbox207. Computer-Use Agents · Pixels to Clicks208. Browser Agents · Navigating the Web209. Coding Agents · The Swe-Bench Frontier210. Agent Evaluation · Trajectories not Just Outputs

RAG Deep

211. Chunking Strategies · Fixed · Semantic · Recursive212. Embedding Models Deep · Choosing a Retriever213. Hybrid Search · Bm25 + Dense Fusion214. Rerankers · Cross-Encoders for Precision215. Query Rewriting · Expansion and Decomposition216. Hyde · Hypothetical Document Embeddings217. Multi-Hop RAG · Chaining Retrievals218. Graph RAG · Retrieval Over Knowledge Graphs219. Contextual Retrieval · Prepending Context220. Late Interaction · Colbert Token-Level Match221. Agentic RAG · The Model Drives Retrieval222. RAG Evaluation · Faithfulness and Relevance

Embeddings & Vector Search

223. Contrastive Learning · Pulling Pairs Together224. Sentence Transformers · Sbert225. Matryoshka Embeddings · Nested Dimensions226. Hnsw · Navigable Small-World Graphs227. Ivf · Inverted File Indexes228. Product Quantization · Compressing Vectors229. Ann Trade-Offs · Recall vs Latency230. Vector Database Internals · Sharding and Filtering231. Multimodal Embeddings · Shared Spaces232. Embedding Drift · When to Re-Index

Inference Optimization Deep

233. Prefill vs Decode · The Two Inference Phases234. Continuous Batching · Keeping the Gpu Full235. Paged Attention · The Vllm Memory Engine236. Speculative Decoding · Draft and Verify237. Medusa & Eagle · Multi-Token Heads238. Gptq & Awq · Post-Training Quantization239. Smoothquant · Activation-Aware Quant240. Fp8 Inference · The Throughput Frontier241. Tensor-Parallel Inference · Sharding at Serve Time242. Chunked Prefill · Interleaving Long Prompts243. Disaggregated Serving · Prefill / Decode Split244. Structured Decoding · Grammars and JSON245. Throughput vs Latency · The Serving Trade-Off246. Kv Cache Quantization · Shrinking the Cache

Beyond Transformers

247. Mixture of Experts · Sparse Activation248. Switch Transformers · One Expert Per Token249. Expert Routing · Load Balancing the Gates250. State Space Models · The Ssm Idea251. Mamba · Selective State Spaces252. Rwkv · Rnn with Transformer-Level Quality253. Hybrid Architectures · Attention + Ssm254. Mixture of Depths · Dynamic Compute Per Token255. Diffusion Language Models · Parallel Decoding256. The Architecture Frontier · What is Next

Multimodal Deep

257. Vision Transformers · Patches as Tokens258. Clip Deep · Contrastive Image-Text259. Vision-Language Models · Llava · The Projector260. Diffusion Math · Forward and Reverse Process261. Latent Diffusion · Compressing to Latent Space262. Diffusion Transformers · Dit263. Flow Matching · The Rectified-Flow Path264. Classifier-Free Guidance · Steering Strength265. Image Tokenizers · Vq-Vae · Vqgan266. Audio Models · Codecs · Tts · Asr267. Video Diffusion · Temporal Consistency268. Any-To-Any Models · Unified Multimodal

Evaluation Deep

269. Benchmark Design · Validity and Coverage270. Contamination · When Test Data Leaks Into Training271. LLM-As-Judge · Models Grading Models272. Pairwise Comparison · Elo and Arenas273. Rubric Grading · Structured Scoring274. Agentic Benchmarks · Swe-Bench · Webarena275. Capability Evals · What Can the Model Do276. Safety Evals · What Will the Model Refuse277. Statistical Significance · Error Bars on Evals278. Eval Gaming · Overfitting to the Benchmark279. Holistic Evaluation · Helm-Style Coverage280. Human Evaluation · The Gold Standard and its Cost

Interpretability Deep

281. The Residual Stream · Reading the Model Internals282. Induction Heads · In-Context Learning Circuits283. Superposition · More Features Than Neurons284. Sparse Autoencoders · Extracting Features285. Feature Circuits · Tracing Computation286. Activation Patching · Causal Tracing287. The Logit Lens · Decoding Intermediate Layers288. Probing Classifiers · What Layers Know289. Steering Vectors · Editing Behavior at Inference290. Dictionary Learning · The Monosemanticity Goal

Safety & Alignment Deep

291. The Limits of Rlhf · What Preferences Miss292. Scalable Oversight · Supervising Superhuman Models293. Weak-To-Strong Generalization294. Deceptive Alignment · The Inner-Alignment Risk295. Evaluation Awareness · Models Knowing They are…296. Jailbreak Taxonomy · How Guardrails Fail297. Prompt Injection Defense · Trust Boundaries298. Machine Unlearning · Removing Knowledge299. Watermarking · Detecting AI Output300. The Alignment Frontier · Open Problems