Oliver's Notes

Machine Learning

These notes cover transformer internals and interpretability methods for representations, circuits, probes, and interventions.

Concept map

How these interpretability notes depend on each other. Columns are sections; a page sits one row below its deepest prerequisite, so every line runs downward from "read this first" to "read this after". Hover or focus a page to light up everything upstream and downstream of it.

  • selected
  • needed first
  • builds on it
overview
Transformer Internals Interpretability Methods Phenomena & Circuits Attention Residual Stream Residual Stream QK & OV Circuits needs Attention, Residual Str… QK & OV Circuits Probes & Validity needs Residual Stream Probes & Validity Logit & Tuned Lens needs Residual Stream Structural Probes needs Probes & Validity, Subs… Subspace Geometry needs Residual Stream Subspace Geometry Composition Probes needs Probes & Validity Bayesian / MDLProbes needs Probes & Validity CausalInterventions needs Probes & Validity, QK &… Interchange & DAS needs Causal Interventions, S… Head Labels needs QK & OV Circuits Induction Heads needs QK & OV Circuits, Head… Binding needs Probes & Validity Binding C-Command Domains needs Binding, Structural Pro… Lookback Mechanism needs Induction Heads, Binding Mental Spaces needs Binding False-Belief Tasks needs Mental Spaces, Represen… Represented vsExpressed needs Logit & Tuned Lens, Pro… Represented vs Expressed

Pages and descriptions come from the site index; the prerequisite links are an editorial judgement, hand-maintained in src/lib/concept-graph.ts. Longest chain here is 5 pages. The full list of pages, with descriptions, is below.