Welcome to the Learning Mechanics DeCal!

Deep learning is in a peculiar situation at this moment in time. Empirically, the capabilities of AI leveraging deep learning (eg. LLMs) have surpassed the expectations of even the most radically optimistic researchers. Publicly available LLMs are disproving math conjectures left and right, AI is being injected into every stage of real drug development life cycles, entire business models have been destroyed by improving AI capabilities, and at this point, AI winning a Nobel Prize is old news. Yet, we lack a comprehensive scientific framework for understanding how exactly these models work and the development of frontier models is closer to alchemy than science.

Not long ago, engineers built working steam engines before anyone understood the science behind them. The effort to understand and build better engines ended up creating an entirely new branch of science: statistical mechanics. Deep learning may be our generation’s steam engine. Learning mechanics is the emerging discipline that aims to understand it from first principles, treating deep learning the way physics treats the natural world: seeking compact mathematical principles, tight connections between theory and experiment, and simple, intuitive explanations for complex phenomena.

This course is for those who want to be part of building that new science.

Readings draw heavily from the perspective paper There Will Be a Scientific Theory of Deep Learning (Simon et al., 2026) and the primary literature it synthesizes.

All lecture notes and homework are also linked inline below, next to the week they belong to.

Course Calendar

Schedule is subject to change.

Wk 1Aug 26
First Week — No Class
Wk 2Sep 2

Lecture 1 Introduction I: Learning Mechanics

What’s the evidence for an emerging scientific theory of deep learning?

Wk 3Sep 9

Lecture 2 Introduction II: Neural Networks

What exactly are neural networks? Why are they hard to study? How will we study them anyways?

Wk 4Sep 16

Lecture 3 Toy Model I: Deep Linear Networks

What can we learn about deep learning from a highly mathematically tractable toy model in deep linear networks?

Wk 5Sep 23

Lecture 4 Toy Model I: Deep Linear Networks (continued)

How can we analytically solve for the training dynamics of deep linear networks?

Wk 6Sep 30

Lecture 5 Toy Model II: Kernel Regression and the NTK

Is there a limit in which neural networks become analytically solvable?

Wk 7Oct 7

Lecture 6 Toy Model II: Kernel Regression and the NTK (continued)

How can we develop a mathematical framework to study kernel regression? Can we predict how kernel regression will perform on real data?

Wk 8Oct 14

Lecture 7 The Lazy (NTK) and Rich (μP) Regimes

In the lazy (NTK) regime, neural networks don’t learn any structure. Is there a regime where they do?

Wk 9Oct 21

Lecture 8 The Lazy (NTK) and Rich (μP) Regimes (continued)

How can we disentangle hyperparameters to maximize feature learning?

Reading: Yang et al. (2024) Lecture Notes Homework
Wk 10Oct 28

Lecture 10 Case Study I: Grokking

How can we apply the tools of learning mechanics to understand grokking?

Wk 11Nov 4

Lecture 11 Case Study II: Representational Geometry

What kind of features are learned by language models? How might we characterize where such features come from and how they’re learned?

Reading: Karkada et al. (2026) Lecture Notes Homework
Wk 12Nov 11
Buffer Week
Wk 13Nov 18

Lecture 13 Final Project Hypothesis Presentations

Wk 14Nov 25
Thanksgiving Break — No Class
Wk 15Dec 2

Lecture 14 Final Project Office Hours

Wk 16Dec 9
RRR Week — No Class


This site uses Just the Docs, a documentation theme for Jekyll.