AI · 2026 active
AI Engineering Explained Notes and projects on AI engineering, built from first principles, one algorithm at a time.
By Tai Bui.
Curriculum · 20 phases · 498 lessons Tap a phase to expand its lessons. Each one ships when its math, code, and test are all written.
0.
Wiring the Workshop 0 / 7 00 Configure your machine once, then get out of your own way.
Your Machine, Configured Study Mark done Version Control Study Mark done Compute: GPUs, Cloud & Containers Study Mark done The Command Line and Linux Study Mark done Where You Write and Run Code Study Mark done Fuel: APIs and Data Study Mark done Debugging and Profiling AI Code Study Mark done I.
Math Foundations 0 / 22 01 The intuition behind every AI algorithm, through code.
Linear Algebra Intuition Study Mark done Vectors, Matrices & Operations Study Mark done Matrix Transformations Study Mark done Calculus for Machine Learning Study Mark done Chain Rule & Automatic Differentiation Study Mark done Probability and Distributions Study Mark done Bayes' Theorem Study Mark done Optimization Study Mark done Information Theory Study Mark done Dimensionality Reduction Study Mark done Singular Value Decomposition Study Mark done Tensor Operations Study Mark done Numerical Stability Study Mark done Norms and Distances Study Mark done Statistics for Machine Learning Study Mark done Sampling Methods Study Mark done Linear Systems Study Mark done Convex Optimization Study Mark done Complex Numbers for AI Study Mark done The Fourier Transform Study Mark done Graph Theory for Machine Learning Study Mark done Stochastic Processes Study Mark done II.
ML Fundamentals 0 / 18 02 Classical ML — still the backbone of most production AI.
What Is Machine Learning Study Mark done Linear Regression Study Mark done Logistic Regression Study Mark done Decision Trees and Random Forests Study Mark done Support Vector Machines Study Mark done K-Nearest Neighbors and Distances Study Mark done Unsupervised Learning Study Mark done Feature Engineering & Selection Study Mark done Model Evaluation Study Mark done Bias-Variance Tradeoff Study Mark done Ensemble Methods Study Mark done Hyperparameter Tuning Study Mark done ML Pipelines Study Mark done Naive Bayes Study Mark done Time Series Fundamentals Study Mark done Anomaly Detection Study Mark done Handling Imbalanced Data Study Mark done Feature Selection Study Mark done III.
Deep Learning Core 0 / 13 03 Neural networks from first principles. No frameworks until you build one.
The Perceptron Study Mark done Multi-Layer Networks and Forward Pass Study Mark done Backpropagation from Scratch Study Mark done Activation Functions Study Mark done Loss Functions Study Mark done Regularization Study Mark done Weight Initialization and Training Stability Study Mark done Learning Rate Schedules and Warmup Study Mark done Build Your Own Mini Framework Study Mark done Introduction to PyTorch Study Mark done Introduction to JAX Study Mark done Debugging Neural Networks Study Mark done IV.
Computer Vision 0 / 28 04 From pixels to understanding — image, video, 3D, VLMs, and world models.
Image Fundamentals — Pixels, Channels, Color Spaces Study Mark done Convolutions from Scratch Study Mark done CNNs — LeNet to ResNet Study Mark done Image Classification Study Mark done Transfer Learning & Fine-Tuning Study Mark done Object Detection — YOLO from Scratch Study Mark done Semantic Segmentation — U-Net Study Mark done Instance Segmentation — Mask R-CNN Study Mark done Image Generation — GANs Study Mark done Image Generation — Diffusion Models Study Mark done Stable Diffusion — Architecture & Fine-Tuning Study Mark done Video Understanding — Temporal Modeling Study Mark done 3D Vision — Point Clouds & NeRFs Study Mark done Vision Transformers (ViT) Study Mark done Real-Time Vision — Edge Deployment Study Mark done Build a Complete Vision Pipeline — Capstone Study Mark done Self-Supervised Vision — SimCLR, DINO, MAE Study Mark done Open-Vocabulary Vision — CLIP Study Mark done OCR & Document Understanding Study Mark done Image Retrieval & Metric Learning Study Mark done Keypoint Detection & Pose Estimation Study Mark done 3D Gaussian Splatting from Scratch Study Mark done Diffusion Transformers & Rectified Flow Study Mark done SAM 3 & Open-Vocabulary Segmentation Study Mark done Vision-Language Models — The ViT-MLP-LLM Pattern Study Mark done Monocular Depth & Geometry Estimation Study Mark done Multi-Object Tracking & Video Memory Study Mark done World Models & Video Diffusion Study Mark done V.
NLP: Foundations to Advanced 0 / 29 05 Language is the interface to intelligence.
Text Processing — Tokenization, Stemming, Lemmatization Study Mark done Bag of Words, TF-IDF, and Text Representation Study Mark done Word Embeddings — Word2Vec from Scratch Study Mark done GloVe, FastText, and Subword Embeddings Study Mark done Sentiment Analysis Study Mark done Named Entity Recognition Study Mark done POS Tagging and Syntactic Parsing Study Mark done CNNs and RNNs for Text Study Mark done Sequence-to-Sequence Models Study Mark done Attention Mechanism — The Breakthrough Study Mark done Machine Translation Study Mark done Text Summarization Study Mark done Question Answering Systems Study Mark done Information Retrieval and Search Study Mark done Topic Modeling — LDA and BERTopic Study Mark done Text Generation Before Transformers — N-gram Language Models Study Mark done Chatbots — Rule-Based to Neural to LLM Agents Study Mark done Multilingual NLP Study Mark done Subword Tokenization — BPE, WordPiece, Unigram, SentencePiece Study Mark done Structured Outputs & Constrained Decoding Study Mark done Natural Language Inference: Textual Entailment Study Mark done Embedding Models — The 2026 Deep Dive Study Mark done Chunking Strategies for RAG Study Mark done Coreference Resolution Study Mark done Entity Linking & Disambiguation Study Mark done Relation Extraction & Knowledge Graph Construction Study Mark done LLM Evaluation — RAGAS, DeepEval, G-Eval Study Mark done Long-Context Evaluation — NIAH, RULER, LongBench, MRCR Study Mark done Dialogue State Tracking Study Mark done VI.
Speech & Audio 0 / 17 06 Hear, understand, speak.
Audio Fundamentals — Waveforms, Sampling, Fourier Transform Study Mark done Spectrograms, Mel Scale & Audio Features Study Mark done Audio Classification — From k-NN on MFCCs to AST and BEATs Study Mark done Speech Recognition (ASR) — CTC, RNN-T, Attention Study Mark done Whisper — Architecture & Fine-Tuning Study Mark done Speaker Recognition & Verification Study Mark done Text-to-Speech (TTS) — From Tacotron to F5 and Kokoro Study Mark done Voice Cloning & Voice Conversion Study Mark done Music Generation — MusicGen, Stable Audio, Suno, and the Licensing Earthquake Study Mark done Audio-Language Models — Qwen2.5-Omni, Audio Flamingo, GPT-4o Audio Study Mark done Real-Time Audio Processing Study Mark done Build a Voice Assistant Pipeline — The Phase 6 Capstone Study Mark done Neural Audio Codecs — EnCodec, SNAC, Mimi, DAC and the Semantic-Acoustic Split Study Mark done Voice Activity Detection & Turn-Taking — Silero, Cobra, and the Flush Trick Study Mark done Streaming Speech-to-Speech — Moshi, Hibiki, and Full-Duplex Dialogue Study Mark done Voice Anti-Spoofing & Audio Watermarking — ASVspoof 5, AudioSeal, WaveVerify Study Mark done Audio Evaluation — WER, MOS, UTMOS, MMAU, FAD, and the Open Leaderboards Study Mark done VII.
Transformers Deep Dive 0 / 16 07 The architecture that changed everything.
Why Transformers — The Problems with RNNs Study Mark done Self-Attention from Scratch Study Mark done Multi-Head Attention Study Mark done Positional Encoding — Sinusoidal, RoPE, ALiBi Study Mark done The Full Transformer — Encoder + Decoder Study Mark done BERT — Masked Language Modeling Study Mark done GPT — Causal Language Modeling Study Mark done T5, BART — Encoder-Decoder Models Study Mark done Vision Transformers (ViT) Study Mark done Audio Transformers — Whisper Architecture Study Mark done Mixture of Experts (MoE) Study Mark done KV Cache, Flash Attention & Inference Optimization Study Mark done Scaling Laws Study Mark done Build a Transformer from Scratch — The Capstone Study Mark done Attention Variants — Sliding Window, Sparse, Differential Study Mark done Speculative Decoding — Draft, Verify, Repeat Study Mark done VIII.
Generative AI 0 / 15 08 Create images, video, audio, 3D, and more.
Generative Models — Taxonomy & History Study Mark done Autoencoders & Variational Autoencoders (VAE) Study Mark done GANs — Generator vs Discriminator Study Mark done Conditional GANs & Pix2Pix Study Mark done Diffusion Models — DDPM from Scratch Study Mark done Latent Diffusion & Stable Diffusion Study Mark done ControlNet, LoRA & Conditioning Study Mark done Inpainting, Outpainting & Image Editing Study Mark done Video Generation Study Mark done Audio Generation Study Mark done 3D Generation Study Mark done Flow Matching & Rectified Flows Study Mark done Evaluation — FID, CLIP Score, Human Preference Study Mark done Visual Autoregressive Modeling (VAR): Next-Scale Prediction Study Mark done IX.
Reinforcement Learning 0 / 12 09 The foundation of RLHF and game-playing AI.
MDPs, States, Actions & Rewards Study Mark done Dynamic Programming — Policy Iteration & Value Iteration Study Mark done Monte Carlo Methods — Learning from Complete Episodes Study Mark done Temporal Difference — Q-Learning & SARSA Study Mark done Deep Q-Networks (DQN) Study Mark done Policy Gradient — REINFORCE from Scratch Study Mark done Actor-Critic — A2C and A3C Study Mark done Proximal Policy Optimization (PPO) Study Mark done Reward Modeling & RLHF Study Mark done Multi-Agent RL Study Mark done Sim-to-Real Transfer Study Mark done RL for Games — AlphaZero, MuZero, and the LLM-Reasoning Era Study Mark done X.
LLMs from Scratch 0 / 24 10 Build, train, and understand large language models.
Tokenizers: BPE, WordPiece, SentencePiece Study Mark done Building a Tokenizer from Scratch Study Mark done Data Pipelines for Pre-Training Study Mark done Pre-Training a Mini GPT (124M Parameters) Study Mark done Scaling: Distributed Training, FSDP, DeepSpeed Study Mark done Instruction Tuning (SFT) Study Mark done RLHF: Reward Model + PPO Study Mark done DPO: Direct Preference Optimization Study Mark done Constitutional AI and Self-Improvement Study Mark done Evaluation: Benchmarks, Evals, LM Harness Study Mark done Quantization: Making Models Fit Study Mark done Inference Optimization Study Mark done Building a Complete LLM Pipeline Study Mark done Open Models: Architecture Walkthroughs Study Mark done Speculative Decoding and EAGLE-3 Study Mark done Differential Attention (V2) Study Mark done Native Sparse Attention (DeepSeek NSA) Study Mark done Multi-Token Prediction (MTP) Study Mark done DualPipe Parallelism Study Mark done DeepSeek-V3 Architecture Walkthrough Study Mark done Jamba — Hybrid SSM-Transformer Study Mark done Async and Hogwild! Inference Study Mark done Speculative Decoding and EAGLE Study Mark done Gradient Checkpointing and Activation Recomputation Study Mark done XI.
LLM Engineering 0 / 17 11 Put LLMs to work in production.
Prompt Engineering: Techniques & Patterns Study Mark done Few-Shot, Chain-of-Thought, Tree-of-Thought Study Mark done Structured Outputs: JSON, Schema Validation, Constrained Decoding Study Mark done Embeddings & Vector Representations Study Mark done Context Engineering: Windows, Budgets, Memory, and Retrieval Study Mark done RAG (Retrieval-Augmented Generation) Study Mark done Advanced RAG (Chunking, Reranking, Hybrid Search) Study Mark done Fine-Tuning with LoRA & QLoRA Study Mark done Function Calling & Tool Use Study Mark done Evaluation & Testing LLM Applications Study Mark done Caching, Rate Limiting & Cost Optimization Study Mark done Guardrails, Safety & Content Filtering Study Mark done Building a Production LLM Application Study Mark done Model Context Protocol (MCP) Study Mark done Prompt Caching and Context Caching Study Mark done LangGraph — State Machines for Agents Study Mark done Agent Framework Tradeoffs — LangGraph vs CrewAI vs AutoGen vs Agno Study Mark done XII.
Multimodal AI 0 / 25 12 See, hear, read, and reason across modalities — from ViT patches to computer-use agents.
Vision Transformers and the Patch-Token Primitive Study Mark done CLIP and Contrastive Vision-Language Pretraining Study Mark done From CLIP to BLIP-2 — Q-Former as Modality Bridge Study Mark done Flamingo and Gated Cross-Attention for Few-Shot VLMs Study Mark done LLaVA and Visual Instruction Tuning Study Mark done Any-Resolution Vision: Patch-n'-Pack and NaFlex Study Mark done Open-Weight VLM Recipes: What Actually Matters Study Mark done LLaVA-OneVision: Single-Image, Multi-Image, Video in One Model Study Mark done Qwen-VL Family and Dynamic-FPS Video Study Mark done InternVL3: Native Multimodal Pretraining Study Mark done Chameleon and Early-Fusion Token-Only Multimodal Models Study Mark done Emu3: Next-Token Prediction for Image and Video Generation Study Mark done Transfusion: Autoregressive Text + Diffusion Image in One Transformer Study Mark done Show-o and Discrete-Diffusion Unified Models Study Mark done Janus-Pro: Decoupled Encoders for Unified Multimodal Models Study Mark done MIO and Any-to-Any Streaming Multimodal Models Study Mark done Video-Language Models: Temporal Tokens and Grounding Study Mark done Long-Video Understanding at Million-Token Context Study Mark done Audio-Language Models: the Whisper to Audio Flamingo 3 Arc Study Mark done Omni Models: Qwen2.5-Omni and the Thinker-Talker Split Study Mark done Embodied VLAs: RT-2, OpenVLA, π0, GR00T Study Mark done Document and Diagram Understanding Study Mark done ColPali and Vision-Native Document RAG Study Mark done Multimodal RAG and Cross-Modal Retrieval Study Mark done Multimodal Agents and Computer-Use (Capstone) Study Mark done XIII.
Tools & Protocols 0 / 23 13 The interfaces between AI and the real world.
The Tool Interface — Why Agents Need Structured I/O Study Mark done Function Calling Deep Dive — OpenAI, Anthropic, Gemini Study Mark done Parallel Tool Calls and Streaming with Tools Study Mark done Structured Output — JSON Schema, Pydantic, Zod, Constrained Decoding Study Mark done Tool Schema Design — Naming, Descriptions, Parameter Constraints Study Mark done MCP Fundamentals — Primitives, Lifecycle, JSON-RPC Base Study Mark done Building an MCP Server — Python + TypeScript SDKs Study Mark done Building an MCP Client — Discovery, Invocation, Session Management Study Mark done MCP Transports — stdio vs Streamable HTTP vs SSE Migration Study Mark done MCP Resources and Prompts — Context Exposure Beyond Tools Study Mark done MCP Sampling — Server-Requested LLM Completions and Agent Loops Study Mark done Roots and Elicitation — Scoping and Mid-Flight User Input Study Mark done Async Tasks (SEP-1686) — Call-Now, Fetch-Later for Long-Running Work Study Mark done MCP Apps — Interactive UI Resources via `ui://` Study Mark done MCP Security I — Tool Poisoning, Rug Pulls, Cross-Server Shadowing Study Mark done MCP Security II — OAuth 2.1, Resource Indicators, Incremental Scopes Study Mark done MCP Gateways and Registries — Enterprise Control Planes Study Mark done MCP Auth in Production — Enrollment, JWKS Refresh, Audience-Pinned Tokens Study Mark done A2A — Agent-to-Agent Protocol Study Mark done OpenTelemetry GenAI — Tracing Tool Calls End-to-End Study Mark done LLM Routing Layer — LiteLLM, OpenRouter, Portkey Study Mark done Skills and Agent SDKs — Anthropic Skills, AGENTS.md, OpenAI Apps SDK Study Mark done Capstone — Build a Complete Tool Ecosystem Study Mark done XIV.
Agent Engineering 0 / 42 14 Build agents from first principles — loop, memory, planning, frameworks, benchmarks, production, workbench.
The Agent Loop: Observe, Think, Act Study Mark done ReWOO and Plan-and-Execute: Decoupled Planning Study Mark done Reflexion: Verbal Reinforcement Learning Study Mark done Tree of Thoughts and LATS: Deliberate Search Study Mark done Self-Refine and CRITIC: Iterative Output Improvement Study Mark done Tool Use and Function Calling Study Mark done Memory: Virtual Context and MemGPT Study Mark done Memory Blocks and Sleep-Time Compute (Letta) Study Mark done Hybrid Memory: Vector + Graph + KV (Mem0) Study Mark done Skill Libraries and Lifelong Learning (Voyager) Study Mark done Planning with HTN and Evolutionary Search Study Mark done Anthropic's Workflow Patterns: Simple Over Complex Study Mark done LangGraph: Stateful Graphs and Durable Execution Study Mark done AutoGen v0.4: Actor Model and Agent Framework Study Mark done CrewAI: Role-Based Crews and Flows Study Mark done OpenAI Agents SDK: Handoffs, Guardrails, Tracing Study Mark done Claude Agent SDK: Subagents and Session Store Study Mark done Agno and Mastra: Production Runtimes Study Mark done Benchmarks: SWE-bench, GAIA, AgentBench Study Mark done Benchmarks: WebArena and OSWorld Study Mark done Computer Use: Claude, OpenAI CUA, Gemini Study Mark done Voice Agents: Pipecat and LiveKit Study Mark done OpenTelemetry GenAI Semantic Conventions Study Mark done Agent Observability: Langfuse, Phoenix, Opik Study Mark done Multi-Agent Debate and Collaboration Study Mark done Failure Modes: Why Agents Break Study Mark done Prompt Injection and the PVE Defense Study Mark done Orchestration Patterns: Supervisor, Swarm, Hierarchical Study Mark done Production Runtimes: Queue, Event, Cron Study Mark done Eval-Driven Agent Development Study Mark done Agent Workbench Engineering: Why Capable Models Still Fail Study Mark done The Minimal Agent Workbench Study Mark done Agent Instructions as Executable Constraints Study Mark done Repo Memory and Durable State Study Mark done Initialization Scripts for Agents Study Mark done Scope Contracts and Task Boundaries Study Mark done Runtime Feedback Loops Study Mark done Verification Gates Study Mark done Reviewer Agent: Separate Builder from Marker Study Mark done Multi-Session Handoff Study Mark done The Workbench on a Real Repo Study Mark done Capstone: Ship a Reusable Agent Workbench Pack Study Mark done XV.
Autonomous Systems 0 / 22 15 Long-horizon agents, self-improvement, and the 2026 safety stack.
The Shift from Chatbots to Long-Horizon Agents Study Mark done STaR, V-STaR, Quiet-STaR — Self-Taught Reasoning Study Mark done AlphaEvolve — Evolutionary Coding Agents Study Mark done Darwin Godel Machine — Open-Ended Self-Modifying Agents Study Mark done AI Scientist v2 — Workshop-Level Autonomous Research Study Mark done Automated Alignment Research (Anthropic AAR) Study Mark done Recursive Self-Improvement — Capability vs Alignment Study Mark done Bounded Self-Improvement Designs Study Mark done The Autonomous Coding Agent Landscape (2026) Study Mark done Claude Code as an Autonomous Agent: Permission Modes and Auto Mode Study Mark done Browser Agents and Long-Horizon Web Tasks Study Mark done Long-Running Background Agents: Durable Execution Study Mark done Action Budgets, Iteration Caps, and Cost Governors Study Mark done Kill Switches, Circuit Breakers, and Canary Tokens Study Mark done Human-in-the-Loop: Propose-Then-Commit Study Mark done Checkpoints and Rollback Study Mark done Constitutional AI and Rule Overrides Study Mark done Llama Guard and Input/Output Classification Study Mark done Anthropic Responsible Scaling Policy v3.0 Study Mark done OpenAI Preparedness Framework and DeepMind Frontier Safety Framework Study Mark done METR Time Horizons and External Capability Evaluation Study Mark done CAIS, CAISI, and Societal-Scale Risk Study Mark done XVI.
Multi-Agent & Swarms 0 / 25 16 Coordination, emergence, and collective intelligence.
Why Multi-Agent? Study Mark done Heritage of FIPA-ACL and Speech Acts Study Mark done Communication Protocols Study Mark done The Multi-Agent Primitive Model Study Mark done Supervisor / Orchestrator-Worker Pattern Study Mark done Hierarchical Architecture and Its Failure Mode Study Mark done Society of Mind and Multi-Agent Debate Study Mark done Role Specialization — Planner, Critic, Executor, Verifier Study Mark done Parallel / Swarm / Networked Architectures Study Mark done Group Chat and Speaker Selection Study Mark done Handoffs and Routines — Stateless Orchestration Study Mark done A2A — The Agent-to-Agent Protocol Study Mark done Shared Memory and Blackboard Patterns Study Mark done Consensus and Byzantine Fault Tolerance for Agents Study Mark done Voting, Self-Consistency, and Debate Topology Study Mark done Negotiation and Bargaining Study Mark done Generative Agents and Emergent Simulation Study Mark done Theory of Mind and Emergent Coordination Study Mark done Swarm Optimization for LLMs (PSO, ACO) Study Mark done MARL — MADDPG, QMIX, MAPPO Study Mark done Agent Economies, Token Incentives, Reputation Study Mark done Production Scaling — Queues, Checkpoints, Durability Study Mark done Failure Modes — MAST, Groupthink, Monoculture, Cascading Errors Study Mark done Evaluation and Coordination Benchmarks Study Mark done Case Studies and the 2026 State of the Art Study Mark done XVII.
Infrastructure & Production 0 / 28 17 Ship AI to the real world.
Managed LLM Platforms — Bedrock, Vertex AI, Azure OpenAI Study Mark done Inference Platform Economics — Fireworks, Together, Baseten, Modal, Replicate, Anyscale Study Mark done GPU Autoscaling on Kubernetes — Karpenter, KAI Scheduler, Gang Scheduling Study Mark done vLLM Serving Internals: PagedAttention, Continuous Batching, Chunked Prefill Study Mark done EAGLE-3 Speculative Decoding in Production Study Mark done SGLang and RadixAttention for Prefix-Heavy Workloads Study Mark done TensorRT-LLM on Blackwell with FP8 and NVFP4 Study Mark done Inference Metrics — TTFT, TPOT, ITL, Goodput, P99 Study Mark done Production Quantization — AWQ, GPTQ, GGUF K-quants, FP8, MXFP4/NVFP4 Study Mark done Cold Start Mitigation for Serverless LLMs Study Mark done Multi-Region LLM Serving and KV Cache Locality Study Mark done Edge Inference — Apple Neural Engine, Qualcomm Hexagon, WebGPU/WebLLM, Jetson Study Mark done LLM Observability Stack Selection Study Mark done Prompt Caching and Semantic Caching Economics Study Mark done Batch APIs — the 50% Discount as Industry Standard Study Mark done Model Routing as a Cost-Reduction Primitive Study Mark done Disaggregated Prefill/Decode — NVIDIA Dynamo and llm-d Study Mark done vLLM Production Stack with LMCache KV Offloading Study Mark done AI Gateways — LiteLLM, Portkey, Kong AI Gateway, Bifrost Study Mark done Shadow Traffic, Canary Rollout, and Progressive Deployment for LLMs Study Mark done A/B Testing LLM Features — GrowthBook, Statsig, and the Vibes Problem Study Mark done Load Testing LLM APIs — Why k6 and Locust Lie Study Mark done SRE for AI — Multi-Agent Incident Response, Runbooks, Predictive Detection Study Mark done Chaos Engineering for LLM Production Study Mark done Security — Secrets, API Key Rotation, Audit Logs, Guardrails Study Mark done Compliance — SOC 2, HIPAA, GDPR, PCI-DSS, EU AI Act, ISO 42001 Study Mark done FinOps for LLMs — Unit Economics and Multi-Tenant Attribution Study Mark done Self-Hosted Serving Selection — llama.cpp, Ollama, TGI, vLLM, SGLang Study Mark done XVIII.
Ethics, Safety & Alignment 0 / 30 18 Build AI that helps humanity. Not optional.
Instruction-Following as Alignment Signal Study Mark done Reward Hacking and Goodhart's Law Study Mark done The Direct Preference Optimization Family Study Mark done Sycophancy as RLHF Amplification Study Mark done Constitutional AI and RLAIF Study Mark done Mesa-Optimization and Deceptive Alignment Study Mark done Sleeper Agents — Persistent Deception Study Mark done In-Context Scheming in Frontier Models Study Mark done Alignment Faking Study Mark done AI Control — Safety Despite Subversion Study Mark done Scalable Oversight and Weak-to-Strong Generalization Study Mark done Red-Teaming: PAIR and Automated Attacks Study Mark done Many-Shot Jailbreaking Study Mark done ASCII Art and Visual Jailbreaks Study Mark done Indirect Prompt Injection — Production Attack Surface Study Mark done Red-Team Tooling — Garak, Llama Guard, PyRIT Study Mark done WMDP and Dual-Use Capability Evaluation Study Mark done Frontier Safety Frameworks — RSP, PF, FSF Study Mark done Anthropic's Model Welfare Program Study Mark done Bias and Representational Harm in LLMs Study Mark done Fairness Criteria — Group, Individual, Counterfactual Study Mark done Differential Privacy for LLMs Study Mark done Watermarking — SynthID, Stable Signature, C2PA Study Mark done Regulatory Frameworks — EU, US, UK, Korea Study Mark done EchoLeak and the Emergence of CVEs for AI Study Mark done Model, System, and Dataset Cards Study Mark done Data Provenance and Training-Data Governance Study Mark done Alignment Research Ecosystem — MATS, Redwood, Apollo, METR Study Mark done Moderation Systems — OpenAI, Perspective, Llama Guard Study Mark done Dual-Use Risk — Cyber, Bio, Chem, Nuclear Uplift Study Mark done XIX.
Capstone Projects 0 / 85 19 17 end-to-end products + 9 deep-build tracks. 20-40 hours per project; 4-12 lessons per track.
Capstone 01 — Terminal-Native Coding Agent Study Mark done Capstone 02 — RAG over Codebase (Cross-Repo Semantic Search) Study Mark done Capstone 03 — Real-Time Voice Assistant (ASR to LLM to TTS) Study Mark done Capstone 04 — Multimodal Document QA (Vision-First PDF, Tables, Charts) Study Mark done Capstone 05 — Autonomous Research Agent (AI-Scientist Class) Study Mark done Capstone 06 — DevOps Troubleshooting Agent for Kubernetes Study Mark done Capstone 07 — End-to-End Fine-Tuning Pipeline (Data to SFT to DPO to Serve) Study Mark done Capstone 08 — Production RAG Chatbot for a Regulated Vertical Study Mark done Capstone 09 — Code Migration Agent (Repo-Level Language / Runtime Upgrade) Study Mark done Capstone 10 — Multi-Agent Software Engineering Team Study Mark done Capstone 11 — LLM Observability & Eval Dashboard Study Mark done Capstone 12 — Video Understanding Pipeline (Scene, QA, Search) Study Mark done Capstone 13 — MCP Server with Registry and Governance Study Mark done Capstone 14 — Speculative-Decoding Inference Server Study Mark done Capstone 15 — Constitutional Safety Harness + Red-Team Range Study Mark done Capstone 16 — GitHub Issue-to-PR Autonomous Agent Study Mark done Capstone 17 — Personal AI Tutor (Adaptive, Multimodal, with Memory) Study Mark done Agent Harness Loop Contract Study Mark done Tool Registry with Schema Validation Study Mark done JSON-RPC 2.0 Over Newline-Delimited Stdio Study Mark done Function Call Dispatcher Study Mark done Plan-Execute Control Flow Study Mark done Capstone Lesson 25: Verification Gates and the Observation Budget Study Mark done Capstone Lesson 26: Sandbox Runner with Denylist and Path Jail Study Mark done Capstone Lesson 27: Eval Harness with Fixture Tasks Study Mark done Capstone Lesson 28: Observability with OTel GenAI Spans and Prometheus Metrics Study Mark done Capstone Lesson 29: End-to-End Coding Agent on the Harness Study Mark done BPE Tokenizer From Scratch Study Mark done Tokenized Dataset with Sliding Window Study Mark done Token and Positional Embeddings Study Mark done Multi-Head Self-Attention Study Mark done Transformer Block from Scratch Study Mark done GPT Model Assembly Study Mark done Training Loop and Evaluation Study Mark done Loading Pretrained Weights Study Mark done Capstone Lesson 38: Classifier Fine-Tuning by Head Swap Study Mark done Capstone Lesson 39: Instruction Tuning by Supervised Fine-Tuning Study Mark done Capstone Lesson 40: Direct Preference Optimization from Scratch Study Mark done Capstone Lesson 41: Full Evaluation Pipeline Study Mark done Large Corpus Downloader Study Mark done HDF5 Tokenized Corpus Study Mark done Cosine LR with Linear Warmup Study Mark done Gradient Clipping and Mixed Precision Study Mark done Gradient Accumulation Study Mark done Checkpoint Save and Resume Study Mark done Distributed Data Parallel and FSDP from Scratch Study Mark done Language Model Evaluation Harness Study Mark done Hypothesis Generator Study Mark done Literature Retrieval Study Mark done Experiment Runner Study Mark done Result Evaluator Study Mark done Paper Writer Study Mark done Critic Loop Study Mark done Iteration Scheduler Study Mark done End-to-End Research Demo Study Mark done Vision Encoder Patches Study Mark done Vision Transformer Encoder Study Mark done Projection Layer for Modality Alignment Study Mark done Cross-Attention Fusion Study Mark done Vision-Language Pretraining Study Mark done Multimodal Evaluation Study Mark done Chunking Strategies, Compared Study Mark done Hybrid Retrieval with BM25 and Dense Embeddings Study Mark done Cross-Encoder Reranker Study Mark done Query Rewriting: HyDE, Multi-Query, and Decomposition Study Mark done RAG Evaluation: Precision, Recall, MRR, nDCG, Faithfulness, Answer Relevance Study Mark done End-to-End RAG System Study Mark done Task Spec Format Study Mark done Classical Metrics Study Mark done Code Exec Metric Study Mark done Perplexity and Calibration Study Mark done Leaderboard Aggregation Study Mark done End-to-End Eval Runner Study Mark done Collective Ops From Scratch Study Mark done Data Parallel DDP From Scratch Study Mark done ZeRO Optimizer State Sharding Study Mark done Pipeline Parallel and Bubble Analysis Study Mark done Sharded Checkpoint and Atomic Resume Study Mark done End-to-End Distributed Training Study Mark done Capstone 82 — Jailbreak Taxonomy Study Mark done Capstone 83 — Prompt Injection Detector Study Mark done Capstone 84 — Refusal Evaluation Study Mark done Capstone 85 — Content Classifier Integration Study Mark done Capstone 86 — Constitutional Rules Engine Study Mark done Capstone 87 — End-to-End Safety Gate Study Mark done Complete In progress Planned