Welcome to the Learning Mechanics DeCal!

Learning Mechanics is the emerging discipline that treats deep learning the way physics treats the natural world: seeking compact mathematical principles1, tight connections between theory and experiment, and simple, intuitive explanations for complex phenomena. Pieces of a scientific theory for deep learning are beginning to fit together, and in this course, we will examine what has been assembled so far, what remains contested, and where the field is heading.

Deep learning is among the most powerful technologies humans have ever built, and understanding it promises to be one of the defining intellectual challenges of the early 21st century. As of 2026, the engineering success of deep learning has dramatically outpaced our scientific understanding of it. Closing that gap may amount to founding a genuinely new field of science—one whose implications for our understanding of intelligence, data, and learning extend well beyond the neural networks that motivated it.

Readings draw heavily from the whitepaper There Will Be a Scientific Theory of Deep Learning (Simon et al., 2026) and the primary literature it synthesizes.

Course Calendar

Schedule is subject to change.

Wk 1
First Week — No Class
Wk 2

Lecture 1 Introduction I: Learning Mechanics

What’s the evidence for an emerging scientific theory of deep learning?

Wk 3

Lecture 2 Introduction II: Neural Networks

What exactly are neural networks? Why are they hard to study? How will we study them anyways?

Reading: Nielsen (2019) Lecture Notes Homework: optional math review
Wk 4

Lecture 3 Toy Model I: Deep Linear Networks

What can we learn about deep learning from a highly mathematically tractable toy model in deep linear networks?

Wk 5

Lecture 4 Toy Model I: Deep Linear Networks (continued)

How can we analytically solve for the training dynamics of deep linear networks?

Lecture Notes Homework
Wk 6

Lecture 5 Toy Model II: Kernel Regression and the NTK

Is there a limit in which neural networks become analytically solvable?

Wk 7

Lecture 6 Toy Model II: Kernel Regression and the NTK (continued)

How can we develop a mathematical framework to study kernel regression? Can we predict how kernel regression will perform on real data?

Wk 8

Lecture 7 The Lazy (NTK) and Rich (μP) Regimes

In the lazy (NTK) regime, neural networks don’t learn any structure. Is there a regime where they do?

Wk 9

Lecture 8 The Lazy (NTK) and Rich (μP) Regimes (continued)

How can we disentangle hyperparameters to maximize feature learning?

Reading: Yang et al. (2024) Lecture Notes Homework
Wk 10
Thanksgiving Break — No Class
Wk 11

Lecture 10 Case Study I: Grokking

How can we apply the tools of learning mechanics to understand grokking?

Wk 12

Lecture 11 Case Study II: Representational Geometry

What kind of features are learned by language models? How might we characterize where such features come from and how they’re learned?

Reading: Karkada et al. (2026) Lecture Notes Homework
Wk 13
Buffer Week
Wk 14

Lecture 13 Final Project Hypothesis Presentations

Wk 15

Lecture 14 Final Project Office Hours

Wk 16
RRR Week — No Class

  1. From Wikipedia: “In physics and other sciences, theoretical work is said to be from first principles, or ab initio, if it starts directly at the level of established science and does not make assumptions such as empirical model and parameter fitting. “First principles thinking” consists of decomposing things down to the fundamental axioms in the given arena, before reasoning up by asking which ones are relevant to the question at hand, then cross referencing conclusions based on chosen axioms and making sure conclusions do not violate any fundamental laws.” 


This site uses Just the Docs, a documentation theme for Jekyll.