Pick a page in the map — by pointer or by tabbing through it — to see what it
assumes and what builds on it.
Drag to pan. Zoom with the buttons above, or hold ⌘ /Ctrl while scrolling.
Probability & Statistics · starting point
Measure Theory & Random Variables Measurable spaces, probability measures, pushforward measures, densities, and importance sampling
Probability & Statistics · starting point
Named Distributions How Bernoulli, Poisson, Gaussian, Cauchy, chi-square, t, F, conjugate priors, and heavy-tail laws are related
Probability & Statistics · 2 deep
Modes of Convergence Almost sure, in probability, in distribution, in L^p — the implication lattice, counterexamples as sample paths, Markov/Chebyshev/Chernoff bounds, and MCT/DCT/Fatou
Probability & Statistics · starting point
Calculus of Variations First variations, Euler-Lagrange residuals, curve relaxation, and the brachistochrone race
Probability & Statistics · 2 deep
Sufficient Statistics Fisher–Neyman factorization, the fiber picture, Rao–Blackwell variance collapse, and a categorical diagram showing factorization, variance, and Fisher info as one commuting square
Probability & Statistics · 3 deep
The Exponential Family The canonical form, naming and the statistical-physics log-partition story, derivatives of A giving the moments of T(X), canonical links (logit, log) behind GLMs, and an interactive picker stepping through six standard members
Probability & Statistics · 4 deep
Fisher Information Likelihood geometry, score functions, Fisher information, exponential families, log-partition, Jeffreys and max-entropy priors, and Bayesian updates
Probability & Statistics · 3 deep
Hypothesis Testing Type-I error, type-II error, power, and decision thresholds through the classic overlapping-distributions diagram
Probability & Statistics · 4 deep
Distance Correlation Distance-based dependence tests, partial distance correlation, and cases Pearson r misses
Information Theory · 2 deep
Entropy & Mutual Information Average surprise, conditional entropy, shared information, binary distributions, and noisy channels
Information Theory · 3 deep
KL Divergence Directed distribution mismatch, support errors, and the difference between forward and reverse KL
Information Theory · 4 deep
Optimal Transport Wasserstein distance and the transport plan: Sinkhorn iteration, and why OT gives gradients where KL divergence does not
Information Theory · 5 deep
Information Geometry Probability distributions as a manifold: the probability simplex, Fisher-Rao geodesics, dual flatness, and e- vs m-projections
Random Processes · 2 deep
Poisson Processes Rare events in time: exponential waits, Poisson counts, equivalent definitions, thinning, splitting, and process diagnostics
Random Processes · 3 deep
Markov Chains Discrete-time and continuous-time chains, stationary distributions, mixing, recurrence, periodicity, and CTMC rates
Random Processes · 2 deep
LTI Systems on Random Inputs Convolution, correlation propagation, spectra, AR(1) as a leaky integrator, and stationarity through linear filters
Random Processes · 3 deep
Power Spectral Density Autocorrelation, spectra, LTI shaping, and periodograms for stationary random processes
Random Processes · 4 deep
Gaussian Processes for Regression Priors over functions, kernels, posterior conditioning, marginal likelihood, 2-D regression, and acquisition
Bayesian Inference · 5 deep
Choosing a Prior Principles of prior selection: use real prior information when you have it; otherwise group invariance, max entropy, or Jeffreys — and how the three routes disagree near boundaries
Bayesian Inference · 4 deep
Conjugate Priors & the Exponential Family Why some prior–likelihood pairs update in closed form, hyperparameters as pseudo-counts, worked Beta/Normal/Gamma examples, and a table of standard pairs
Bayesian Inference · 5 deep
Posterior Summaries & Bayes Risk Squared, absolute, and zero-one loss pick out the posterior mean, median, and mode — three views of the same posterior, only one of which ignores everything but the peak
Bayesian Inference · 5 deep
Hierarchical Bayes Two-level Normal–Normal model, the posterior formula for borrowing strength across groups, empirical-Bayes fitting of the between-group variance, and the connection to ridge regression
Bayesian Inference · 6 deep
Bayesian Regression: Penalties as Priors OLS, ridge, LASSO, and best-subset selection as MAP under four noise/prior pairs — and why the shape of the prior near zero determines whether the estimator shrinks, selects, or both
Bayesian Inference · 5 deep
Bayesian Graphical Models DAG factorization, d-separation, explaining away, Dirichlet-multinomial CPT learning, and structure scoring
Bayesian Inference · 6 deep
Hidden Markov Models HMM sampling, forward-backward filtering and smoothing, log-domain messages, and Viterbi versus marginal MAP paths
Bayesian Inference · 4 deep
Monte Carlo & MCMC Rejection, importance sampling, Metropolis-Hastings, Gibbs, RJMCMC, simulated annealing, and when to use each method on a static target
Bayesian Inference · 7 deep
Kalman & Particle Filters Sequential inference of a hidden state from noisy observations: Kalman filter for linear-Gaussian models, EKF/UKF for local linearization, particle filter for fully nonlinear non-Gaussian SSMs
Bayesian Inference · 5 deep
Variational Bayes for Gaussian Mixtures CAVI for a 2-D Gaussian mixture with Normal–Wishart and Dirichlet priors, showing component ellipses, automatic pruning of unused components, and the ELBO trace
Bayesian Inference · 6 deep
Bayesian Neural Networks Weight posteriors, predictive function ensembles, Laplace approximation, evidence, Occam's hill, and prior mismatch