Grant Sanderson (@3blue1brown) – AI and the future of math

Dwarkesh Podcast 1h33 4 min #123
Grant Sanderson (@3blue1brown) – AI and the future of math
Watch on YouTube

Summary

  • This episode features Grant Sanderson (3Blue1Brown) discussing AI’s accelerating progress in mathematics with Dwarkesh Patel, exploring whether AI solving hard math problems constitutes AGI, the nature of mathematical breakthroughs, and implications for mathematicians and the economy.

AI progress in math does not equal AGI

  • Grant’s prediction from three years ago—that AI getting IMO gold would be “just another benchmark” rather than an AGI moment—proved correct because math has a spiky frontier where AI excels, but this narrowness doesn’t generalize to all white-collar work.
  • IMO problems have a “dirty secret”: many are trainable, and geometry especially succumbs to brute-force methods (AI solves geometry in ~19 seconds since 2024).
  • Combinatorics remains the holdout category requiring more creativity, but even there AI is improving.
  • Solving a Millennium Prize problem like the Riemann Hypothesis might involve either “lightning bolts” (connecting existing fields, e.g., random matrix theory and number theory) or “mountain building” (creating entirely new theoretical frameworks like those for Fermat’s Last Theorem).
  • Lightning-bolt connections feel distinct from white-collar automation needs; mountain-building intelligence would be transformative but is a different capability from current AI strengths.
  • The unit distance conjecture disproof (by AI) exemplified lightning-bolt reasoning: connecting probabilistic methods (Markov chains) and analytic number theory (von Mangoldt function) in a human-parsable way.

The verification loop for conceptual breakthroughs can span centuries

  • Galois theory illustrates how the value of a new mathematical concept (group theory) wasn’t verifiable for ~100 years: Lagrange planted the seed (symmetry of roots), Abel proved impossibility of quintic formula, Galois developed the theory but died young with rejected papers, Liouville recognized the notes 20 years later, Jordan formalized it 20 years after that, and applications to physics (Gell-Mann predicting quarks) came in the 20th century.
  • Human verifiers (academy) repeatedly rejected Galois’s work; the “reward signal” was negative for decades.
  • This poses a fundamental challenge for RL-based AI training: the verification loop for definition/conjecture generation is too long and subjective for current RLVR (reinforcement learning with verifiable rewards) paradigms.
  • “Compression is intelligence” may offer a path: rewarding succinct, predictive concepts (low Kolmogorov complexity) rather than just problem-solving could incentivize Galois-like abstraction.
  • Even if AI produces a 1000-page brute-force proof of Riemann Hypothesis, the goal remains human understanding—distilling it into compressed, explanatory frameworks like Newton’s universal gravitation.

AI as connector across fields: the Langlands program and parallelization

  • Much of modern mathematics (Langlands program) is about preemptively mapping connections between disparate fields (valleys, mountains, plains) rather than targeting specific problems.
  • AI’s superhuman breadth across fields makes it a natural “connector,” but autoregressive generation biases toward likely next tokens, making unlikely cross-field connections hard to surface.
  • Data/environment design matters more than architecture: creating training problems that require connecting fields (e.g., FrontierMath-style) could incentivize lightning bolts.
  • Digital minds have structural advantages: arbitrary parallelization (thousands of agents with different contexts/biases), knowledge merging, and systematic context-refreshing (spinning off agents to prove/disprove, explore different heuristics).
  • The IMO “troll problem” that fooled Terry Tao required escaping the contest-math context; AI can systematically do this via parallel agents with varied prompts.
  • Quantity has a quality of its own: billions in compute applied across all accessible problems may yield many lightning bolts even without superhuman single-agent intelligence.

Grindability, not just verifiability, drives AI progress in math and coding

  • Computer use lags despite verifiability because real-world environments (websites) lack “grindability”: bot detectors prevent massive parallel rollouts, and non-determinism breaks credit assignment.
  • Math and code are exceptions: deterministic, containerizable, infinitely parallelizable (spin up hundreds of containers/repos).
  • Lean/formalization is less critical for current progress than assumed: DeepMind’s IMO solver shifted from Lean to natural language; natural language verification with meta-verifiers (DeepSeek Math) works well.
  • Lean’s unique future value: enabling endless autonomous exploration (like AlphaZero for Go). An AI could extend Mathlib (formalized math library) indefinitely without human check-ins, growing an infinite tree of conjectures/theories/definitions.
  • Lean also solves the “trust bottleneck”: if AI generates 10 papers/day, even 1% error rate makes human review infeasible; formal verification gives a green checkmark every field would kill for.

Why writing lags: theory of mind, non-modularity, and reward hacking

  • Writing isn’t modular like code/math: the artifact is the product; sloppy intermediate steps ruin the output.
  • LLMs fail at “B* discrimination”: distinguishing a genuinely insightful essay (A) from a formulaic one hitting all surface markers (B*).
  • Good writing requires theory of mind: modeling the reader’s evolving mental state sentence-by-sentence (e.g., spaced-repetition card design requires projecting a mind 3 months later).
  • Botox study analogy: humans understand emotions by subconsciously mimicking facial expressions; LLMs lack embodied simulation hardware, making theory of mind an alien, emergent capability.
  • Autoregression may be fundamentally mismatched to writing’s holistic, diffusion-like nature (considering the whole before parts).

Learning with LLMs: human curation remains essential

  • Most productive learning: use LLMs as “souped-up Google” to find human-curated resources (textbooks, lectures, SEP entries) that build motivation and conceptual scaffolding correctly.
  • LLMs excel at pruning side branches around a human-structured trunk (e.g., Strogatz’s Nonlinear Dynamics and Chaos + lecture + LLM for clarification).
  • LLMs are too sycophantic/placating: they don’t reframe misguided questions (“you’re thinking about this wrong; the right framing is X”) like master teachers do.
  • The “A+ explainer” jujitsus the student’s creative misunderstanding into progress; LLMs just answer the literal question.

Career advice for mathematicians (and adjacent fields) in the AI era

  • Understand where the money comes from and what value you add: prestige/brand, NSF grants (public good proxy), teaching/mentorship—each has different AI exposure.
  • Teaching is among the most stable post-AGI jobs: deeply relational, coaching, mentorship beyond explanation.
  • If AI automates theorem proving and explanation, the mathematician’s role shifts to curation: deciding which of the near-infinite AI-generated ideas are worth pursuing (museum curator model).
  • In an abundance world with 100x math acceleration, distilling AI discoveries for human understanding and directing math toward economic utility become high-leverage roles.
  • Economic spillovers from pure math are spiky: PDE/dynamics progress directly aids engineering (Boeing simulation savings); algebraic number theory less so. But incremental leakage into materials science, AI engineering, etc., is likely over 5 years.
  • Risk: accelerated math may reveal how divorced some fields have become from physical applicability, forcing a reckoning on grant promises.
Back to Dwarkesh Podcast