# LLM Systems & Interpretability

## Curriculum v6 — Public Curriculum

### Course overview

This is a hands-on course for technically curious learners who want a practical understanding of modern AI systems, from the broad AI landscape down to the internal mechanics of large language models and outward again to complete systems such as ChatGPT, Gemini, Claude, and coding agents.

The course emphasizes understanding how the pieces fit together, inspecting real open models, and using small experiments and labs to make abstract ideas concrete.

### Intended outcomes

By the end of the course, a learner should be able to:

- distinguish major families of AI and explain how they differ;
- distinguish an algorithm, a trained model, a neural-network architecture, an inference engine, an agent harness, and a product;
- explain the path from text to tokens to embeddings to generated output;
- read the major parts of a modern Transformer architecture;
- explain attention, residual streams, multilayer perceptrons (MLPs), normalization, positional encoding, KV caches, and Mixture of Experts;
- distinguish pretraining, post-training, fine-tuning, safety training, and inference-time behavior;
- explain where behavior can live: weights, context, inference settings, orchestration, or external governance;
- understand reasoning effort and inference-time computation;
- build a small tool-using coding agent around an open model;
- inspect activations, routing, and learned features in open models;
- perform controlled interventions such as ablation, steering, and forced MoE routing;
- run reproducible instructional experiments and distinguish correlation from causation.

---

# Course Roadmap

**Module 1 begins** by looking broadly at what AI means, what kinds of systems exist, and how they differ. From there we narrow our focus toward large language models, introducing tokenization and the neural-network foundations needed to understand them.

**Module 2** focuses on Transformers, the architecture behind most modern LLMs. We explore the residual stream, attention, multilayer perceptrons, Mixture of Experts, and the evolution from the original Transformer to modern models.

**Module 3** explains how those models come to exist in the first place. We look at how pretraining turns initially untrained weights into a language model, and how post-training then shapes that model into something that can follow instructions, use tools, reason, and behave more like a useful assistant.

**Module 4** moves from the model itself to what happens when we actually use one. We look at inference and then at the larger systems built around LLMs, including context, retrieval, memory, tools, agents, safety controls, and coding assistants.

**Module 5** turns inward again, introducing mechanistic interpretability and techniques for inspecting activations, learned features, steering, ablation, and other ways of investigating what is happening inside a model.

**Module 6, finally, brings us to experimental methods**: how to design careful LLM experiments, control for misleading effects, and bring the techniques from the course together in a reusable model-inspection environment. The course ends with a small research capstone, inviting the learner to choose a question they find interesting and investigate it for themselves.

---

# How This Course Works

The course has two kinds of activity: **Lessons** and **Labs**. Lessons provide the conceptual spine; Labs make ideas concrete through coding, model inspection, mathematics, paper reading, experiments, or tooling. Some Labs use local hardware, while others use AWS or rented GPU compute when the hardware meaningfully enables the exercise.

Mathematics is introduced when it becomes useful rather than front-loaded as a separate prerequisite course. Each lesson identifies **Math to know / refresh**, distinguishing what is needed immediately from material that is useful to revisit.

We recommend reading a lesson once for the overall picture, then returning to difficult sections interactively and doing Labs where they are recommended. Important ideas should sometimes be explained back, predicted, sketched, compared, or applied rather than simply reread.

Before suitable experiments, make a prediction and record why you expect it. Comparing that prediction with the result is part of the learning process.

Optional papers, documentation, videos, mathematical refreshers, and deeper material appear under **More Learning**. The core lesson should remain understandable without requiring every optional resource.

A recurring question throughout the course is: **Where does this behavior live?** The answer may involve the tokenizer, architecture, trained weights, prompt/context, inference settings, agent or orchestration layer, external controls, or product layer.

---

# Module 1 — Foundations: From AI to LLMs

Module 1 begins with the wider AI landscape, then deliberately narrows toward large language models. After the AI Zoo, the remaining lessons introduce the representations and neural-network foundations needed before we study Transformers in detail.

## 1.1 — The AI Zoo: What Counts as AI?

### Core questions

- What does “AI” actually refer to?
- Which systems are algorithms, which are trained models, and which are products built from many components?
- How is an LLM different from an image generator, video model, genetic algorithm, or agent?

### Topics

- symbolic and rule-based AI;
- search and planning;
- optimization methods;
- evolutionary and genetic algorithms;
- classical machine learning;
- neural networks;
- supervised, unsupervised, and self-supervised learning;
- reinforcement learning;
- LLMs;
- image-generation models;
- video-generation models;
- speech and audio models;
- multimodal models;
- agents and orchestration systems;
- how an LLM may interpret a prompt that is then executed by a different model.

### Math to know / refresh

**Needed now:** none beyond basic functions and probability intuition.

**Useful refresh:** optimization as searching for maxima/minima; probability distributions.

### Early labs

- Map familiar AI systems into the categories above.
- Identify which parts of a modern AI product are models versus software around models.
- Optional AWS familiarization lab: billing dashboard, budgets, VPC/EC2 vocabulary, and launch templates without running expensive compute.

---

*From this point onward, the course focuses primarily on large language models.*

## 1.2 — Text Becomes Numbers

### Topics

- vocabulary;
- tokenizer;
- tokens versus words;
- token IDs;
- special tokens;
- embedding matrices;
- embeddings;
- logits;
- softmax;
- next-token prediction;
- autoregressive generation.

### Core question

What precisely happens between typing text and the model choosing another token?

### Math to know / refresh

**Needed now:** vectors, indexing, exponentials, probability normalization.

**Useful refresh:** matrix lookup notation and basic probability.

### Lab ideas

- Inspect a real open-model tokenizer.
- Encode and decode arbitrary text.
- Compare tokenization of ordinary prose, code, numbers, and unusual words.

---

## 1.3 — Neural Networks Without the Mysticism

### Topics

- vectors and matrices;
- weights;
- biases;
- activations;
- layers;
- linear transformations;
- nonlinear activation functions;
- what a “neuron” means in modern networks;
- training versus inference.

### Math to know / refresh

**Needed now:** matrix multiplication, dot products, functions, simple derivatives.

**Useful refresh:** coordinate transformations and linear algebra notation.

### Lab ideas

- Build and inspect a microscopic neural network.
- Manually compute a forward pass for a tiny example.

---

# Module 2 — The Transformer

The Transformer architecture was introduced in 2017 in *Attention Is All You Need* by Vaswani and colleagues. Its attention-based design had an enormous impact on modern AI and became the architectural foundation for most large language models. This module works through the main pieces of the Transformer and then follows how they evolved into today's LLMs.

### Recommended Lab — Reading *Attention Is All You Need*

Use the original paper as a recurring reference rather than trying to understand it all at once. At the start of the module, read the abstract, introduction, architecture figure, and conclusion; identify the problem the authors were trying to solve and note unfamiliar terminology. After 2.3, return to the architecture and attention sections, connect them to what you have learned, and identify which parts modern decoder-only LLMs retain, change, or omit.

## 2.1 — The Residual Stream

### Topics

- embeddings as vectors;
- Transformer blocks;
- residual stream;
- residual connections;
- normalization;
- information accumulation across layers.

### Math to know / refresh

**Needed now:** vector addition, norms, affine transformations.

**Useful refresh:** geometric interpretation of high-dimensional vectors.

### Lab ideas

- Load a small Transformer and record the residual stream after each layer.
- Compare layer-to-layer changes for selected tokens.

---

## 2.2 — Attention

### Topics

- queries, keys, and values;
- attention scores;
- causal masks;
- self-attention;
- multi-head attention;
- what attention does and does not tell us.

### Core equations

Q = XWq

K = XWk

V = XWv

Attention(Q,K,V) = softmax(QK^T / sqrt(d))V

### Math to know / refresh

**Needed now:** matrix multiplication, transpose, dot products, exponentials, softmax, basic probability.

**Useful refresh:** cosine similarity and vector projection.

### Lab ideas

- Compute a tiny attention example by hand.
- Visualize attention matrices for real sentences.

---

## 2.3 — Multilayer Perceptrons (MLPs), Features, and Superposition

### Topics

- feed-forward / multilayer perceptron (MLP) layers;
- expansion dimensions;
- activation functions;
- features;
- polysemantic neurons;
- superposition;
- why one neuron rarely means one concept.

### Math to know / refresh

**Needed now:** matrix transformations, nonlinear functions.

**Useful refresh:** basis vectors, linear combinations, dimensionality.

### Lab ideas

- Compare MLP activations for contrasting prompts.
- Inspect which units respond across different contexts.

---

## 2.4 — From the Original Transformer to Modern LLMs

### Topics

- decoder-only models;
- positional information;
- RoPE;
- RMSNorm;
- grouped-query attention;
- KV caches;
- context windows;
- local/sliding attention;
- quantization;
- model configuration files.

### Math to know / refresh

**Needed now:** vectors, rotations, norms.

**Useful refresh:** complex-number or rotation-matrix intuition for RoPE; numerical precision.

### Lab ideas

- Read a real model configuration file and translate its parameters into plain English.
- Estimate memory requirements at different numerical precisions.

---

## 2.5 — Mixture of Experts

### Topics

- dense models versus MoE;
- routers;
- expert networks;
- top-k routing;
- routing probabilities;
- total versus active parameters;
- expert specialization;
- load balancing;
- gpt-oss as a primary specimen.

### Math to know / refresh

**Needed now:** softmax, ranking/top-k selection, weighted sums.

**Useful refresh:** sparse computation and conditional probability.

### Lab ideas

- Capture which experts are selected for each token and layer.
- Force, prohibit, or boost selected experts.
- Compare routing across prompt categories.

---

# Module 3 — Where Models Come From

## 3.1 — Pretraining

### Topics

- training corpora;
- prediction objectives;
- loss;
- gradient descent;
- backpropagation;
- batches;
- learning rates;
- checkpoints;
- compute requirements;
- what becomes encoded in weights.

### Math to know / refresh

**Needed now:** derivatives, gradients, chain rule, logarithms, cross-entropy.

**Useful refresh:** multivariable calculus and optimization landscapes.

### Lab ideas

- Train an intentionally tiny language model.
- Plot loss over training and inspect checkpoints.

---

## 3.2 — Post-Training, Alignment, and Safety Learned into Weights

### Topics

- base models versus assistant models;
- instruction tuning;
- supervised fine-tuning;
- preference training;
- RLHF and related methods;
- reasoning training;
- tool-use training;
- policy-following and refusal behavior;
- safety behavior learned into model weights;
- limits of trained-in safety.

### Core distinction

A model may know how to do something while post-training changes how it responds to requests for that capability.

### Math to know / refresh

**Needed now:** loss functions, probability, expected reward.

**Useful refresh:** reinforcement-learning terminology and optimization.

### Lab ideas

- Compare base-model and instruction-tuned behavior where suitable open checkpoints exist.
- Classify examples of behavior as likely coming from pretraining, post-training, or system-level controls.

---

# Module 4 — Running Models and Turning Them into Systems

## 4.1 — Inference: Running a Trained Model

### Topics

- checkpoint loading;
- inference engines;
- GPU memory;
- numerical precision;
- BF16, FP16, FP8, 4-bit formats;
- quantization;
- batching;
- sampling;
- temperature;
- top-p;
- deterministic generation;
- KV cache.

### Math to know / refresh

**Needed now:** probability distributions, logarithms, numerical representation.

**Useful refresh:** floating-point precision and information loss.

### Lab ideas

- Run an open model locally or on rented compute.
- Measure speed, memory, and output changes under different sampling settings.

---

## 4.2 — Reasoning and Inference-Time Compute

### Topics

- reasoning tokens;
- inference-time compute;
- low/medium/high reasoning effort;
- test-time scaling;
- trained reasoning behavior;
- same weights versus different models;
- what reasoning controls can and cannot mean.

### Math to know / refresh

**Needed now:** computational complexity intuition.

**Useful refresh:** scaling relationships and expected-value reasoning.

### Lab ideas

- Compare reasoning-effort settings on a model that exposes them.
- Measure latency, token use, and task performance.

---

## 4.3 — From an LLM to ChatGPT/Gemini-Like Products

### Topics

- system/developer/user messages;
- conversation formats;
- context windows;
- files and retrieval;
- memory;
- tool definitions;
- safety layers;
- context management and compaction;
- product behavior versus model behavior.

### Recurring question

Where does each behavior live: weights, prompt, inference engine, orchestration, or UI/product layer?

### Math to know / refresh

Minimal new mathematics.

### Lab ideas

- Trace a hypothetical request through a full product stack.
- Design a minimal conversation protocol around an open model.

---

## 4.4 — Tool Calling, Agents, and External Governance

### Topics

- function calling;
- tool schemas;
- observations;
- agent loops;
- stopping conditions;
- permissions;
- sandboxes;
- network controls;
- approval gates;
- safety classifiers and safeguard models;
- capability versus authority.

### Math to know / refresh

Minimal new mathematics.

### Lab ideas

- Give an open model a calculator or file-read tool.
- Add permissions and observe what the harness, rather than the model, prevents.

---

## 4.5 — Build Our Own Tiny Codex

### Goal

Build an actual coding agent around an open model.

### Components

- model;
- system instructions;
- repository reader;
- file editor;
- shell/test tool;
- agent loop;
- sandbox;
- permissions;
- context management;
- execution limits.

### Math to know / refresh

No major new mathematics; this is primarily a systems lesson.

### Lab ideas

- Repair a deliberately broken codebase.
- Compare unrestricted versus governed tool access.
- Record failures and retry behavior.

---

# Module 5 — Mechanistic Interpretability

## 5.1 — Looking Inside a Model

### Topics

- hooks;
- activation caches;
- residual-stream inspection;
- attention inspection;
- MLP activations;
- logit lens;
- probing;
- correlation versus intervention.

### Math to know / refresh

**Needed now:** vector norms, dot products, projections, cosine similarity.

**Useful refresh:** linear regression/probes and statistical association.

### Lab ideas

- Track an internal representation through successive layers.
- Compare activations for controlled prompt pairs.

---

## 5.2 — Sparse Autoencoders and Learned Features

### Topics

- superposition;
- sparse representations;
- dictionary learning;
- autoencoders;
- sparse autoencoders;
- learned features;
- feature activation;
- feature interpretation;
- Gemma Scope as a primary specimen.

### Math to know / refresh

**Needed now:** linear combinations, reconstruction error, sparsity.

**Useful refresh:** optimization with regularization; basis/dictionary representations.

### Lab ideas

- Inspect existing SAE features.
- Compare feature activations across prompt families.

---

## 5.3 — Steering, Ablation, and Causal Intervention

### Topics

- activation vectors;
- contrastive vectors;
- activation addition;
- ablation;
- forced routing;
- causal interventions;
- controls and confounders.

### Math to know / refresh

**Needed now:** vector subtraction, projection, normalization, effect measurement.

**Useful refresh:** causal inference vocabulary.

### Lab ideas

- Construct a contrastive activation direction.
- Add/subtract it at selected layers.
- Measure behavioral and internal changes.

---

## 5.4 — Emotion, Persona, and Internal State

### Topics

- emotion-related representations;
- persona vectors;
- behavioral correlates;
- causal intervention;
- representation versus subjective experience;
- anthropomorphism;
- careful interpretation of mechanistic findings.

### Math to know / refresh

**Needed now:** similarity metrics, averaging, variance.

**Useful refresh:** dimensionality reduction and statistical testing.

### Lab ideas

- Compare emotional-state prompt sets.
- Test whether a discovered direction generalizes to held-out prompts.
- Intervene and measure whether behavior changes systematically.

---

# Module 6 — Experimental Methods and Integration

## 6.1 — Experimental Design for LLM Labs

### Topics

- hypotheses;
- independent and dependent variables;
- controls;
- confounders;
- train/test separation;
- sample sizes;
- reproducibility;
- falsification;
- exploratory versus confirmatory analysis.

### Math to know / refresh

**Needed now:** mean, variance, distributions, confidence intervals, effect sizes.

**Useful refresh:** hypothesis testing, multiple comparisons, regression.

### Lab ideas

- Turn one vague claim about model behavior into a falsifiable lab plan.
- Record expected measurements before running the experiment, then compare prediction with result.

---

## 6.2 — The LLM Microscope

### Goal

Integrate the course into a reusable learning environment for instrumenting open models, capturing activations and MoE routing, applying interventions, and producing reproducible lab results.

### Likely pipeline

open model → instrumented inference → activation/routing capture → intervention → repeated experiment → analysis → visualization → write-up

### Math to know / refresh

Depends on the chosen lab or advanced topic.

### Lab ideas

- Provision appropriate AWS compute only when needed.
- Run a complete instructional experiment with saved configuration, data, plots, and conclusions.

---

# Research Capstone — Ask Your Own Question

Choose a question about LLMs or AI systems that genuinely interests you and use the concepts, tools, and experimental habits from this course to investigate it. Define the question, choose an appropriate method, and document what you learn.