Pick a page in the map — by pointer or by tabbing through it — to see what it
assumes and what builds on it.
Drag to pan. Zoom with the buttons above, or hold ⌘ /Ctrl while scrolling.
Transformer Internals · starting point
Attention A weighted sum read as an adaptive sufficient statistic, Nadaraya–Watson kernel regression, and entropy-regularized retrieval
Transformer Internals · starting point
Residual Stream & Directions The transformer residual stream as a shared workspace; features as directions; superposition as sparse feature packing
Transformer Internals · 2 deep
QK and OV Circuits An attention head has separate routing and residual-write components
Interpretability Methods · 2 deep
Probes and Validity Probe scores, selectivity controls, lexical controls, and the distinction between decodability and causal use
Interpretability Methods · 2 deep
Logit Lens & Tuned Lens Layerwise vocabulary readouts, tuned affine decoders, and the difference between depth and cognitive time
Interpretability Methods · 3 deep
Dependency Trees & Structural Probes Dependency grammar, tree distance and depth, structural-probe geometry, MST extraction, and syntactic controls
Interpretability Methods · 2 deep
Subspace Geometry PCA projection, principal angles, Procrustes and affine alignment, and relative representations by anchors
Interpretability Methods · 3 deep
Compositionality & Semantic Probes Compositional meaning as a relation between head and dependent vectors, from additive to bilinear and nonlinear probes
Interpretability Methods · 3 deep
Causal Interventions Ablation, activation patching, path patching, attribution patching, and self-repair under component removal
Interpretability Methods · 4 deep
Interchange Interventions & DAS Trained low-rank activation swaps, matched random controls, transfer normalization, and expressive-fit cautions
Phenomena & Circuits · 3 deep
Attention Head Labels Positional, induction, syntactic, rare-word, copy-suppression, and name-mover labels as hypotheses rather than stable kinds
Phenomena & Circuits · 4 deep
Induction Heads The prefix-match then copy mechanism behind the [A][B] ... [A] -> [B] transformer circuit
Phenomena & Circuits · 3 deep
Binding Resolving a use against a nonlocal source — agreement, anaphora, traces, variable use, and logical chaining — with minimal pairs and distractor controls
Phenomena & Circuits · 4 deep
C-Command & Binding Domains Principles A/B/C, c-command domains, dependency-depth proxies, and island boundaries as interactive tree geometry
Phenomena & Circuits · 5 deep
The Lookback Mechanism Store an address, carry a pointer, look back to dereference it: how binding IDs and retrieval heads combine into one in-context recall motif
Phenomena & Circuits · 4 deep
Mental Spaces Fauconnier-style reality, belief, and picture frames with connectors and frame-specific readouts
Phenomena & Circuits · 5 deep
False-Belief Tasks Sally-Anne, unexpected contents, and shortcut-blocking controls for belief-versus-reality stimuli
Phenomena & Circuits · 3 deep
Represented vs. Expressed Knowledge Surprisal, internal readouts, and cases where a model carries information that does not surface in the output distribution