Mathematics of AI

Curriculum

Lessons 1-38 are complete and have canonical MDX Lesson Pages. Source Transcripts remain preserved for Lessons 1-25, and Exercise Artifact Modules exist for Lessons 2-38. Lesson 39 is the active Lesson Page.

The next active topic is **Taylor Series and Local Approximation**, connecting function values, gradients, and Hessians to increasingly accurate local models.

Complete

Phase I — Linear Algebra Foundations

This phase builds the geometric foundation for understanding data, transformations, projections, dimensionality reduction, covariance, and later neural network layers.

  • - Rotation/reflection, scaling, rotation/reflection
  • Singular values
  • Rank
  • Low-rank approximation
  • Why SVD is more general than eigendecomposition
  • - Maximum variance directions
  • Dimensionality reduction
  • PCA as a change of basis
  • PCA and SVD

Complete

Phase II — Probability, Statistics, and Geometry

This phase connects linear algebra to uncertainty, distributions, inference, and the probabilistic interpretation of machine learning.

Lesson 12

Covariance

  • - Variance
  • Covariance
  • Covariance matrices
  • Correlation versus covariance
  • - Random variables
  • Probability distributions as structured objects
  • Expectation notation
  • Linearity of expectation
  • Probability vectors and transformations
  • - Covariance matrices as geometric objects
  • Eigenvectors as principal directions
  • Eigenvalues as variance magnitudes
  • Ellipses and ellipsoids
  • Whitening intuition
  • - Gaussian distributions in multiple dimensions
  • Mean vectors
  • Covariance matrices
  • Probability density geometry
  • Mahalanobis distance
  • - Likelihood as a score, not a probability distribution over parameters
  • Log-likelihood
  • Scaling likelihoods
  • Maximum likelihood estimates
  • - Loss as a training objective
  • Negative log-likelihood
  • Squared error
  • Gaussian error assumptions
  • Connecting probability models to optimization targets
  • - Priors
  • Likelihoods
  • Posteriors
  • Maximum a posteriori estimates
  • Regularization as a prior
  • - Bayesian update structure
  • Priors and evidence
  • Posterior distributions
  • Difference between belief updates and point estimates
  • - Updating beliefs with data
  • Posterior movement
  • Confidence and uncertainty
  • How new evidence changes the parameter distribution
  • - Prediction under uncertainty
  • Integrating over parameters
  • Posterior predictive distributions
  • Difference between estimating a parameter and predicting a new observation

Complete through Lesson 25

Phase III — Generalization and Regularization

This phase explains why fitting the training data is not enough and introduces the methods used to control overfitting.

  • - Capacity
  • Underfitting
  • Overfitting
  • Training accuracy versus test accuracy
  • Why more parameters do not guarantee better generalization
  • - Penalizing complexity
  • L1 regularization
  • L2 regularization
  • MAP interpretation of regularization
  • Bias-variance tradeoff intuition
  • - Validation loss
  • Stopping before overfitting dominates
  • Early stopping as implicit regularization
  • Training dynamics

Lesson 25

Dropout

  • - Randomly disabling units during training
  • Reducing co-adaptation
  • Ensemble intuition
  • Training behavior versus inference behavior

Complete through Lesson 33

Phase IV — Optimization

This phase turns loss functions into training procedures. The emphasis is on understanding optimization geometrically and computationally.

  • - Gradients as steepest increase
  • Moving opposite the gradient
  • Learning rate
  • Loss landscapes
  • Convergence intuition
  • - Mini-batches
  • Noisy gradients
  • Why stochasticity can help
  • Epochs and iterations

Lesson 28

Momentum

  • - Accumulating velocity
  • Smoothing noisy updates
  • Escaping shallow ravines
  • Geometric intuition

Lesson 29

Adam

  • - Per-parameter learning rates
  • Adam
  • First and second moment estimates
  • Bias correction
  • - Decay schedules
  • Warmup
  • Cosine schedules
  • Why learning rates change during training
  • - Second derivatives
  • Hessian matrices
  • Curvature directions
  • Why curvature affects optimizer behavior
  • - Convex functions
  • Local versus global minima
  • Saddle points
  • Why neural network optimization is hard but often works
  • - Constraints
  • Lagrange multipliers
  • Penalty methods
  • Constrained learning problems

Lesson 38 complete; Lesson 39 active

Phase V — Calculus and Matrix Calculus for ML

This phase formalizes the calculus needed for backpropagation, optimization, and reading papers with vectorized derivatives.

  • - Functions of multiple variables
  • Holding variables constant
  • Local sensitivity

Lesson 35

Gradients

  • - Gradient vectors
  • Directional derivatives
  • Gradients as covectors / local linear approximations

Lesson 36

Jacobians

  • - Derivatives of vector-valued functions
  • Local linear maps
  • Shape discipline
  • - Composed functions
  • Forward pass and backward pass
  • Local gradients
  • Accumulating derivatives
  • - Numerator and denominator layout conventions
  • Shape checking
  • Gradients with respect to vectors and matrices
  • Common derivative identities
  • - First-order approximation
  • Second-order approximation
  • Why gradients and Hessians describe local behavior

Planned

Phase VI — Neural Networks

This phase introduces neural networks as composed differentiable functions trained by optimization.

  • - Weighted sums
  • Bias terms
  • Linear classification

Lesson 41

Activation Functions

  • - Nonlinearity
  • Sigmoid
  • Tanh
  • ReLU and variants
  • Saturation and dead units

Lesson 42

Feedforward Networks

  • - Composing layers
  • Hidden representations
  • Universal approximation intuition

Lesson 43

Backpropagation

  • - Reverse-mode automatic differentiation
  • Chain rule through layers
  • Gradients of weights and biases

Lesson 44

Initialization

  • - Symmetry breaking
  • Variance propagation
  • Xavier and He initialization

Lesson 45

Normalization

  • - Batch normalization
  • Layer normalization
  • Stabilizing training

Lesson 46

Residual Connections

  • - Identity paths
  • Gradient flow
  • Deep networks

Planned

Phase VII — Information Theory

This phase explains many common loss functions and model objectives in terms of information, coding, and distribution comparison.

Lesson 47

Entropy

  • - Uncertainty
  • Expected information
  • Bits and nats

Lesson 48

Cross Entropy

  • - Expected coding cost under the wrong distribution
  • Classification loss
  • Relation to negative log-likelihood

Lesson 49

KL Divergence

  • - Distribution mismatch
  • Asymmetry
  • Bayesian and variational interpretations

Lesson 50

Mutual Information

  • - Shared information
  • Dependence between variables
  • Representation learning intuition

Lesson 51

Maximum Entropy

  • - Least-assumptive distributions
  • Constraints
  • Why Gaussians appear so often

Planned

Phase VIII — Classical Machine Learning Algorithms

This phase revisits common algorithms after the mathematical foundation is in place.

Lesson 52

Linear Regression Revisited

  • - Least squares
  • MLE interpretation
  • Regularized regression

Lesson 53

Logistic Regression

  • - Classification probabilities
  • Sigmoid link function
  • Cross entropy loss

Lesson 54

Naive Bayes

  • - Conditional independence assumptions
  • Generative classification
  • Bayesian updating

Lesson 55

k-Nearest Neighbors

  • - Distance metrics
  • Decision boundaries
  • Curse of dimensionality

Lesson 56

Decision Trees

  • - Splitting criteria
  • Entropy and Gini impurity
  • Overfitting and pruning

Lesson 57

Random Forests and Ensembles

  • - Bagging
  • Variance reduction
  • Ensemble averaging

Lesson 58

Support Vector Machines

  • - Margins
  • Constrained optimization
  • Kernels

Lesson 59

Clustering

  • - k-Means
  • Hierarchical clustering
  • DBSCAN
  • Geometry of unsupervised learning

Planned

Phase IX — Deep Learning Architecture Math

This phase explains the mathematical objects used in modern deep learning architectures.

Lesson 60

Convolutions

  • - Local receptive fields
  • Translation equivariance
  • Filters and feature maps

Lesson 61

Embeddings

  • - Learned vector representations
  • Similarity geometry
  • Token, item, and feature embeddings

Lesson 62

Softmax

  • - Normalizing scores into probabilities
  • Temperature
  • Softmax gradients

Lesson 63

Sequence Modeling

  • - Recurrent structure
  • Hidden states
  • Sequence-to-sequence intuition

Lesson 64

Attention

  • - Queries, keys, and values
  • Similarity-weighted aggregation
  • Scaled dot-product attention

Lesson 65

Positional Encoding

  • - Position as input information
  • Sinusoidal encodings
  • Learned positional embeddings

Lesson 66

Transformers

  • - Multi-head attention
  • Feedforward blocks
  • Residual paths and normalization
  • Why transformers scale well

Planned

Phase X — Advanced Topics for Paper Reading

This phase is not meant to be exhaustive. It introduces enough of each topic to make research papers less opaque.

Lesson 67

Tensor Algebra

  • - Multidimensional arrays
  • Tensor contractions
  • Shape reasoning

Lesson 68

Spectral Graph Theory

  • - Graphs as matrices
  • Adjacency matrices
  • Graph Laplacians
  • Eigenvectors of graphs

Lesson 69

Manifold Learning

  • - Data on lower-dimensional structure
  • Local linearity
  • Embeddings and charts

Lesson 70

Differential Geometry for ML

  • - Curved spaces
  • Tangent spaces
  • Metrics
  • Optimization on manifolds

Lesson 71

Probabilistic Graphical Models

  • - Conditional dependence
  • Directed and undirected models
  • Factorization
  • Inference intuition

Lesson 72

Variational Inference

  • - Approximate posteriors
  • Evidence lower bound
  • KL minimization

Lesson 73

Monte Carlo Methods

  • - Sampling
  • Estimating expectations
  • MCMC intuition

Planned

Phase XI — Reading Research Papers

This phase focuses on using the accumulated math to understand, reproduce, and critique papers.

Lesson 74

Reading Mathematical Notation

  • - Parsing definitions
  • Tracking symbol tables
  • Recognizing overloaded notation

Lesson 75

Deriving Algorithms from Papers

  • - Moving from objective to update rule
  • Identifying assumptions
  • Reconstructing omitted algebra

Lesson 76

Understanding Proofs

  • - What a proof is trying to establish
  • Common proof strategies
  • Knowing when proof details matter

Lesson 77

Reproducing Results

  • - Minimal implementations
  • Experimental setup
  • Metrics and baselines

Lesson 78

Critical Evaluation

  • - What is being claimed?
  • What is actually shown?
  • What assumptions are hidden?
  • What would change your mind?