Skip to content

Active Curriculum Index

This index covers the active First Principles curriculum: 22 topics across five phases.

Mathematical Foundations

The mathematical prerequisites are maintained in the sister repository applied-mathematics-foundation.

Foundation Main concepts Used heavily by
Linear Algebra Projections, eigendecomposition, SVD, norms Linear Regression, KNN, SVM, PCA, neural networks
Calculus and Optimization Gradients, Hessians, Taylor approximation, convexity Gradient Descent, Logistic Regression, SVM, neural networks
Probability and Statistics Conditional probability, expectation, MLE/MAP Linear/Logistic Regression, Naive Bayes, GMM
Information Theory Entropy, cross-entropy, KL divergence Trees, classification losses, t-SNE
Numerical Computing Conditioning, stability, vectorization Every computational topic

Topic Matrix

Maturity follows the Notebook Standards §9: Planned → Draft → Complete → Verified (Complete plus every §10 validation gate passing). This table is the single source of per-topic maturity; other dashboards link here.

# Topic Core idea Main math Prerequisites Phase Maturity
01 Linear Regression ŷ = Wx + b for continuous values Projection, least squares LA, Calculus, Probability 1 🏅 Verified
02 Gradient Descent θ ← θ − α∇L(θ) Gradients, smoothness, convexity Calculus, 01 1 🏅 Verified
03 Regularization Penalize complexity: L1, L2, MAP Norms, constrained optimization 01, 02 1 🏅 Verified
04 Logistic Regression σ(z) = 1/(1+e⁻ᶻ) → classification Cross-entropy, MLE, convexity Probability, 02 1 🏅 Verified
05 Decision Tree Recursive splits on impurity Entropy, Gini, greedy search Information Theory 2 🏅 Verified
06 Ensemble Methods Bagging (RF) and boosting (GB) Bootstrap, variance reduction 05 2 🏅 Verified
07 KNN Majority vote of K nearest points Norms, metric geometry Linear Algebra 2 🏅 Verified
08 Naive Bayes P(y|x) via Bayes + independence Bayes theorem, conditional independence Probability 2 🏅 Verified
09 SVM Maximum-margin hyperplane Convex optimization, KKT LA, Optimization 2 🏅 Verified
10 PCA Max-variance projection Covariance, eigenvalues, SVD LA, Statistics 1 🏅 Verified
11 Clustering K-Means, DBSCAN, GMM Distances, density, EM LA, Probability 2 🏅 Verified
12 Dimensionality Reduction LDA, t-SNE Scatter matrices, KL divergence 10, Information Theory 2 🏅 Verified
13 Neural Networks Stacked linear + nonlinear layers Chain rule, matrix calculus 02, 04 3 🏅 Verified
14 CNN Convolution for spatial features Cross-correlation, pooling 13 3 🏅 Verified
15 RNN/LSTM Recurrence for sequential data BPTT, gating mechanisms 13 3 🏅 Verified
16 Transformer Self-attention: softmax(QKᵀ/√dₖ)V Attention, positional encoding 13 4 🏅 Verified
17 Autoencoder Encode → latent → decode Representation, reconstruction 13 3 🏅 Verified
18 Reinforcement Learning MDPs, Q-Learning, Policy Gradient Bellman equations, Markov property Calculus, Probability, 13 5 🏅 Verified
19 Generative Models VAE, GAN, Diffusion Models ELBO, Wasserstein distance, SDEs Probability, Calculus, 13, 17 5 🏅 Verified
20 Graph Neural Networks Spectral & Spatial Graph Convolutions Graph Laplacian, Message passing LA, Calculus, 13 5 🏅 Verified
21 LLM Engineering Tokenization, PEFT (LoRA), DPO Subword algorithms, rank decomposition 13, 16 5 🏅 Verified
22 Self-Supervised Learning Contrastive loss, MAE, SimCLR InfoNCE, mutual information bound Probability, 13, 17 5 🏅 Verified

Prerequisite Graph

graph TD
    subgraph Foundations["Foundations"]
        LA["Linear Algebra"]
        CO["Calculus & Optimization"]
        PS["Probability & Statistics"]
        IT["Information Theory"]
        NC["Numerical Computing"]
    end

    CO --> NC
    PS --> IT

    subgraph Phase1["Phase 1 — Core Mathematical ML"]
        T01["01 Linear Regression"]
        T02["02 Gradient Descent"]
        T03["03 Regularization"]
        T04["04 Logistic Regression"]
        T10["10 PCA"]
    end

    LA --> T01
    CO --> T01
    PS --> T01
    LA --> T02
    CO --> T02
    T01 --> T02
    T01 --> T03
    T02 --> T03

    subgraph Phase2["Phase 2 — Classical Machine Learning"]
        T05["05 Decision Tree"]
        T06["06 Ensemble Methods"]
        T07["07 KNN"]
        T08["08 Naive Bayes"]
        T09["09 SVM"]
        T11["11 Clustering"]
        T12["12 Dimensionality Reduction"]
    end

    T01 --> T04
    T02 --> T04
    PS --> T04
    IT --> T05
    PS --> T05
    T05 --> T06
    T02 --> T06
    LA --> T07
    PS --> T08
    IT --> T08
    LA --> T09
    CO --> T09
    T02 --> T09

    subgraph Phase3["Phase 3 — Deep Learning"]
        T13["13 Neural Networks"]
        T14["14 CNN"]
        T15["15 RNN / LSTM"]
        T17["17 Autoencoder"]
    end

    LA --> T10
    PS --> T10
    LA --> T11
    PS --> T11
    T10 --> T12
    IT --> T12

    subgraph Phase4["Phase 4 — Transformers"]
        T16["16 Transformer"]
    end

    LA --> T13
    CO --> T13
    T02 --> T13
    T04 --> T13
    NC --> T13
    T13 --> T14
    T13 --> T15
    T13 --> T16
    T15 --> T16
    T13 --> T17
    T10 --> T17
    IT --> T17

    subgraph Phase5["Phase 5 — Modern AI & Advanced Architectures"]
        T18["18 Reinforcement Learning"]
        T19["19 Generative Models"]
        T20["20 Graph Neural Networks"]
        T21["21 LLM Engineering"]
        T22["22 Self-Supervised Learning"]
    end

    T13 --> T18
    PS --> T18
    T13 --> T19
    T17 --> T19
    PS --> T19
    LA --> T20
    T13 --> T20
    T16 --> T21
    T13 --> T21
    T13 --> T22
    T17 --> T22

Topics 5–9 and 10–12 can be studied in parallel once their prerequisites are met.

Math-to-Algorithm Mapping

graph LR
    subgraph MathConcepts["Mathematical Concepts"]
        PROJ(("Projection"))
        SVD(("SVD"))
        EIGEN(("Eigenvalues"))
        ENTROPY(("Entropy"))
        XENT(("Cross-Entropy"))
        CHAIN(("Chain Rule"))
        NORMS(("Norms"))
        MLE(("MLE"))
        KLD(("KL Divergence"))
        CONVEX(("Convexity"))
        BAYES(("Bayes' Theorem"))
        MATMUL(("Matrix Multiply"))
        SOFTMAX(("Softmax"))
    end

    subgraph Algorithms["Algorithms"]
        LR["Linear Regression"]
        PCA["PCA"]
        DT["Decision Tree"]
        LOGREG["Logistic Regression"]
        NN["Neural Networks"]
        KNN["KNN"]
        SVM["SVM"]
        CLUST["Clustering"]
        GD["Gradient Descent"]
        GMM["GMM"]
        AE["Autoencoder"]
        NB["Naive Bayes"]
        XFORMER["Transformer"]
    end

    PROJ -->|"least squares"| LR
    SVD -->|"decomposition"| PCA
    EIGEN -->|"covariance spectrum"| PCA
    ENTROPY -->|"split criterion"| DT
    XENT -->|"loss function"| LOGREG
    XENT -->|"loss function"| NN
    CHAIN -->|"backpropagation"| NN
    NORMS -->|"distance metric"| KNN
    NORMS -->|"margin"| SVM
    NORMS -->|"centroid distance"| CLUST
    MLE -->|"parameter estimation"| LOGREG
    MLE -->|"EM algorithm"| GMM
    KLD -->|"reconstruction loss"| AE
    CONVEX -->|"convergence"| GD
    CONVEX -->|"dual problem"| SVM
    BAYES -->|"posterior"| NB
    MATMUL -->|"forward pass"| NN
    MATMUL -->|"attention QKᵀ"| XFORMER
    SOFTMAX -->|"output layer"| LOGREG
    SOFTMAX -->|"attention weights"| XFORMER

Algorithm Families

Family Members Common thread
Linear Models LinReg, Ridge, Lasso, LogReg Weighted sum of features + loss
Tree Models Decision Tree, Random Forest, Gradient Boosting Recursive partitioning
Distance-Based KNN, K-Means, DBSCAN, Hierarchical Proximity in feature space
Probabilistic Naive Bayes, GMM, Logistic Regression Explicit probability modeling
Subspace Methods PCA, LDA, t-SNE, Autoencoder Dimensionality reduction
Neural Architectures MLP, CNN, RNN, Transformer, Autoencoder Composable differentiable layers

5 Mathematical Pillars of ML

Pillar Role Examples
Linear Algebra Data as matrices/vectors; the spine of every neural network Matrix multiply in MLPs, SVD in PCA, attention QKᵀ
Calculus Derivatives to minimize error — foundation of backprop Chain rule, Jacobian, gradient descent
Probability & Statistics Reasoning under uncertainty; parameter estimation Bayes' theorem, MLE/MAP, cross-entropy
Optimization Finding optimal parameters — broader than applied calculus SGD, Adam, KKT (SVM), EM algorithm
Information Theory Measuring information and distribution divergence Entropy, KL divergence, mutual information, ELBO

Cross-Topic Navigation