Mathematics in the Age of AI: Terence Tao, Goodhart's Law, and the Crisis of Proof Indigestion

Explore Terence Tao's warning on AI in mathematics, Goodhart's law, the 5 stages of proof, and why machine-solved problems risk "proof indigestion."

Mathematics in the Age of AI: Terence Tao, Goodhart's Law, and the Crisis of Proof Indigestion

Key Takeaways (Quick Summary)

  • The Million-Dollar Paradox: If an artificial intelligence solves a centuries-old Millennium Prize problem with a single prompt, but no living human can decipher how or why it works, has the problem actually been solved?
  • The Second Foundational Crisis: Unlike the crisis of the early 20th century triggered by Russell's Paradox (which threatened the logical foundations of truth), the AI era threatens the unwritten human framework of scientific values, credit, and communal understanding.
  • Goodhart's Law in Science: When the metric of "open problems solved" becomes the explicit target, it ceases to be a reliable measure of scientific progress. AI games benchmark numbers while eliminating the decades of tool-building and mentorship that historically accompanied human breakthroughs.
  • The 5 Stages of Proof: Mathematics is not merely proof generation and proof verification; it relies fundamentally on exposition, digestion, and canonicalization. Flooding the literature with opaque machine proofs creates a dangerous intellectual bottleneck known as "proof indigestion."
  • The Leiden Declaration & The Blackboard Test: To preserve scientific integrity, academia must reward conceptual synthesis and clarity over brute problem-solving speed. If an author cannot explain their proof at a physical blackboard, it should not be certified for publication.

Imagine waking up tomorrow to the headline that a frontier artificial intelligence model just cracked the Riemann Hypothesis or solved the Birch and Swinnerton-Dyer conjecture in a single inference run. The solution is fed into an interactive theorem prover like Lean 4, and the machine confirms: the logic is 100% sound.

Yet, when you open the paper, you encounter 60,000 lines of dense, alien code. No living mathematician can comprehend why the proof holds together, what mathematical structures were synthesized, or how to apply the underlying concepts to any other discipline.

Can we legitimately say the problem is solved?

This provocative question lies at the heart of mathematics in the age of AI, a dilemma recently articulated by Fields Medalist Terence Tao in his landmark reflections on the future of research. Tao—widely regarded as one of the greatest living mathematicians—has sounded an urgent alarm: humanity must adapt to an ecosystem where automated tools advance high-level mathematics faster than human minds can metabolize them.

Featured Snippet Bait: Mathematics in the age of AI faces an unprecedented paradigm shift. While machine learning systems and formal verification tools like Lean can generate and verify proofs at scale, they risk divorcing problem-solving from human comprehension. As Terence Tao warns, without human exposition, digestion, and canonicalization, science faces "proof indigestion"—a flood of logically correct but intellectually impenetrable theorems.


1. The Echo of History: From Russell's Paradox to the Second Crisis

To understand the magnitude of what we are facing, we need to travel back more than a century.

In the late 19th and early 20th centuries, mathematics stood on surprisingly fragile ground. Trailblazing mathematicians made breathtaking discoveries, but their proofs relied heavily on physical intuition, geometric visualization, and heuristic arguments. Basic mathematical primitives—such as what constitutes a "set," a "number," or a rigorous "logical proof"—lacked formalization. Mathematicians focused on doing mathematics, leaving definitional pedantry to philosophers.

Then came 1901, and with it, Russell's Paradox.

By considering the set of all sets that do not contain themselves ($R = {x \mid x \notin x}$), Bertrand Russell exposed a catastrophic structural contradiction at the very core of naive set theory. If $R$ is not a member of itself, its definition dictates it must belong to itself; if it does belong to itself, it contradicts its defining criterion.

The mathematical community was paralyzed. If naive set theory allowed a fatal contradiction, the entire edifice of arithmetic and analysis was suspect.

Thirty years later, in 1931, Kurt Gödel proved his famous Incompleteness Theorems, showing that no consistent formal axiomatic system can prove its own consistency or capture all mathematical truths. Yet, from the ashes of this First Foundational Crisis, modern mathematics emerged stronger than ever.

Mathematicians rebuilt the entire discipline on the rigorous foundation of Zermelo-Fraenkel set theory with the Axiom of Choice (ZFC). That arduous rebuilding project yielded modern mathematical logic and directly catalyzed the birth of theoretical computer science through the pioneering work of Alan Turing and Alonzo Church.

graph LR
    subgraph FirstCrisis ["The First Crisis (1900–1930s)"]
        A["Intuitive, Unformalized Foundations"] -->|"Russell's Paradox (1901)"| B["Crisis of Logical Consistency"]
        B -->|"ZFC Axiomatization & Gödel"| C["Birth of Computer Science & Rigorous Foundations"]
    end

    subgraph SecondCrisis ["The Second Crisis (2020s–Present)"]
        D["Human-Centric Scientific Values"] -->|"Superhuman AI Proof Generation"| E["Crisis of Meaning, Credit & Values"]
        E -->|"Goodhart's Law & Proof Indigestion"| F["The Leiden Framework & The Blackboard Test"]
    end

    style FirstCrisis fill:#282828,stroke:#504945,stroke-width:1px,color:#ebdbb2
    style SecondCrisis fill:#282828,stroke:#504945,stroke-width:1px,color:#ebdbb2
    style A fill:#3c3836,stroke:#7c6f64,color:#ebdbb2
    style B fill:#3c1f1e,stroke:#fb4934,color:#fbf1c7
    style C fill:#2e3b2b,stroke:#b8bb26,color:#fbf1c7
    style D fill:#3c3836,stroke:#7c6f64,color:#ebdbb2
    style E fill:#3c1f1e,stroke:#fb4934,color:#fbf1c7
    style F fill:#26383c,stroke:#83a598,color:#fbf1c7

Today, we are crossing the threshold into the Second Great Crisis of Mathematics.

In terms of magnitude, this crisis is just as sweeping as the first. But its nature is completely inverted. The danger today does not threaten the formal truth or consistency of mathematics; rather, it threatens the unwritten human framework of scientific purpose, incentives, and rewards.

What constitutes an authentic contribution? Who—or what—earns credit for a breakthrough? What is the purpose of proof if humans are removed from the loop of understanding?


2. The Working Hypothesis: The First Proof Benchmark

Before analyzing this crisis, we must discard an outdated, dismissive question: "Can AI really do high-level mathematics?"

Skeptics often point to large language models hallucinating basic arithmetic. But that misses the exponential trajectory of modern AI research. In his analysis, Terence Tao urges researchers to adopt a pragmatic working hypothesis:

Assume that AI tools will soon, with high probability, superior quality, and near-zero marginal cost, advance high-level research mathematics.

This is no longer science fiction. We already have empirical proof: the First Proof benchmark project.

In this landmark experiment, elite mathematicians gathered to author 10 entirely novel, fiercely difficult mathematical problems. These questions were constructed from scratch and had zero digital footprint on the internet, making memorization impossible.

When frontier AI systems tackled this battery of problems, they successfully cracked 7 out of the 10 problems—at a fraction of the time and financial cost required by a team of human researchers.

On the surface, this feels like an unqualified victory for human ingenuity. But as economists and anthropologists have warned us for decades, when an ecosystem experiences a sudden, asymmetrical shock to its metrics, unforeseen pathology follows.


3. Goodhart's Law: How AI Breaks the Incentive Engine of Science

To understand why automated problem solving poses an existential hazard to research culture, we have to examine Goodhart's Law.

Coined by British economist Charles Goodhart and elegantly generalized by anthropologist Marilyn Strathern, the principle states:

"When a measure becomes a target, it ceases to be a good measure."

flowchart TD
    A["Human Era: Organic Synergy"] --> B["Deep Theory Building"]
    A --> C["Educating Students"]
    A --> D["Solving Open Problems (Metric)"]
    A --> E["Philosophical Understanding"]

    F["AI Era: Metric Hijacking"] -->|Goodhart's Law| G["Target: Solely Solve Open Problems"]
    G -->|Automated Brute Search| H["Leaderboard Score: 100%"]
    H -.->|Severed Connection| I["Lost: Mentorship, Conceptual Tools & Insight"]

    style A fill:#2e3b2b,stroke:#b8bb26,stroke-width:2px,color:#fbf1c7
    style B fill:#3c3836,stroke:#7c6f64,color:#ebdbb2
    style C fill:#3c3836,stroke:#7c6f64,color:#ebdbb2
    style D fill:#3c3836,stroke:#7c6f64,color:#ebdbb2
    style E fill:#3c3836,stroke:#7c6f64,color:#ebdbb2
    style F fill:#3c1f1e,stroke:#fb4934,stroke-width:2px,color:#fbf1c7
    style G fill:#3a3220,stroke:#fabd2f,color:#fbf1c7
    style H fill:#3b291a,stroke:#fe8019,stroke-width:2px,color:#fbf1c7
    style I fill:#282828,stroke:#928374,stroke-dasharray: 5 5,color:#a89984

The mathematical community operates under several unwritten, overarching goals:

  1. Formulating novel, unifying theories
  2. Discovering deep, conceptual truths about the universe
  3. Mentoring, inspiring, and training the next generation of scholars
  4. Solving famously difficult open problems

Historically, these objectives lived in perfect harmony. Advancing one goal inevitably advanced all the others. Because of this natural alignment, academic institutions settled on a simple, legible proxy metric to evaluate mathematicians: the number and importance of open problems they solved.

In a world populated solely by human brains, this proxy metric worked brilliantly.

If a mathematician devoted a decade of their life to resolving an Erdős conjecture, they could not snap their fingers and manufacture a solution. To conquer the problem, they had to invent novel technical machinery, write extensive expository monographs, present preliminary findings at conferences, debate peer reviewers, and mentor graduate students who assisted with the lemmas.

The single metric—solving the problem—acted as an engine that dragged all the other unwritten values forward.

The Double Rupture Caused by AI

Generative AI shatters that harmony in two distinct ways:

  1. Epistemic Ungroundedness: Machine learning models lack human contextual grounding. They do not perceive mathematics as an interconnected tapestry of physical intuition and philosophical meaning. Instead, an AI optimizes strictly for output tokens that satisfy a formal checker or maximize an objective function. It aims for a solution that is formally valid on paper, irrespective of whether it yields conceptual enlightenment.
  2. Commercial Benchmarking Incentives: The artificial intelligence industry is governed by financial incentives that reward quantifiable benchmark dominance. Progress is measured by pass rates on public leaderboards, funding rounds, and compute scaling laws.

When an ungrounded model attacks the metric of "open problems solved," it achieves the target while gutting the underlying substance. The benchmark records a triumph, the problem is labeled "closed," but the human insight, educational legacy, and theoretical synthesis completely evaporate.

| Dimension | The Human Mathematical Era | The AI Generation Era | | :--- | :--- | :--- | | Primary Incentive | Holistic understanding, theory building & mentorship | Maximizing benchmark scores & leaderboard rankings | | Solving an Open Problem | Requires years of tool-building, exposition, and seminars | Instantaneous output generation via automated search | | Scientific Byproducts | New mathematical fields, student theses, shared intuition | Massive formal code bases, opaque inference traces | | Knowledge Transfer | Deep cognitive integration across the global community | Concentrated in machine-readable verification files |


4. Deconstructing Proof: The 5 Stages of Problem Solving

Problem-solving is not an atomic, indivisible act. In his analysis, Terence Tao breaks mathematical discovery down into five distinct evolutionary stages:

flowchart TD
    subgraph Phase1 ["1. Proof Generation"]
        A["Raw Idea & Candidate Argument"]
    end

    subgraph Phase2 ["2. Proof Verification"]
        B["Formal Checkers: Lean 4, Isabelle, Coq"]
    end

    subgraph Phase3 ["3. Proof Exposition"]
        C["Human-Readable Translation & Pedagogical Friction"]
    end

    subgraph Phase4 ["4. Proof Digestion"]
        D["Community Debate, Assimilation & Extracting Invariants"]
    end

    subgraph Phase5 ["5. Proof Canonicalization"]
        E["Distillation, Textbook Synthesis & Interdisciplinary Bridges"]
    end

    A -->|AI Automated| B
    B -->|Interactive Theorem Provers| C
    C -->|Cognitive Struggle| D
    D -->|Centuries of Cultural Work| E

    style Phase1 fill:#26383c,stroke:#83a598,stroke-width:1px,color:#ebdbb2
    style Phase2 fill:#283935,stroke:#8ec07c,stroke-width:1px,color:#ebdbb2
    style Phase3 fill:#3a3220,stroke:#fabd2f,stroke-width:1px,color:#ebdbb2
    style Phase4 fill:#3b291a,stroke:#fe8019,stroke-width:1px,color:#ebdbb2
    style Phase5 fill:#362635,stroke:#d3869b,stroke-width:1px,color:#ebdbb2
    style A fill:#3c3836,stroke:#83a598,color:#fbf1c7
    style B fill:#3c3836,stroke:#8ec07c,color:#fbf1c7
    style C fill:#3c3836,stroke:#fabd2f,color:#fbf1c7
    style D fill:#3c3836,stroke:#fe8019,color:#fbf1c7
    style E fill:#3c3836,stroke:#d3869b,color:#fbf1c7

Stage 1: Proof Generation

This is the genesis of an argument—brainstorming raw conjectures, identifying candidate strategies, and sketching pathways through uncharted mathematical terrain. Historically, this represented the ultimate cognitive bottleneck for humans. Today, frontier AI systems commoditize this stage, executing rapid combinatorial searches across problem spaces.

Stage 2: Proof Verification

Once a raw proof is proposed, it must be verified. Is the logic sound, or is it riddled with subtle fallacies? Historically, verification required years of grueling peer review by exhausted subject-matter experts.

Today, interactive theorem provers (ITPs) like Lean 4, Isabelle, and Coq have transformed verification. Once an argument is formalized in Lean, a computer kernel checks every deduction down to foundational axioms. If the kernel compiles the file, the proof is mathematically watertight.

Stage 3: Proof Exposition

Here is where the fault lines appear. Proof exposition is the art of translating formal symbolic logic into natural language that humans can read and appreciate.

Generative text models excel at writing fluent, polished prose. But as Tao observes, this effortless fluency is precisely what ruins mathematical learning.

When humans write mathematics, their drafts contain natural pedagogical friction. The phrasing slows down around subtle steps; authors write caveats, use analogies, and leave signposts warning the reader: "Pay attention here; this lemma looks trivial, but it contains the entire conceptual trick."

AI-generated text strips this friction away. It produces frictionless, uniform prose where a profound conceptual leap is formatted identically to a routine algebraic expansion. The result reads smoothly, but it severely impairs comprehension.

Stage 4: Proof Digestion

Proof digestion is a slow, deeply social cognitive metabolism. A mathematician reads an expository paper, discusses it over coffee with colleagues, debates its assumptions in seminar rooms, and ponders why the argument works.

Through digestion, the mathematical community extracts the invariant core of an idea, discovering how techniques developed in algebraic geometry might unlock unsolved mysteries in number theory.

Stage 5: Proof Canonicalization

This is arguably the most valuable intellectual activity in human civilization. In this final stage, the scientific community takes a messy, 400-page pioneering proof, strips away unnecessary detours, discovers more elegant formulations, and integrates the core insights into standard university textbooks.

Consider this irony: the only reason modern AI systems can write mathematics today is that they were trained on centuries of human-canonicalized literature. Machine learning models feast on the polished, distilled textbooks that human scholars bled to produce over generations.


5. "Proof Indigestion": The Paradox of Abundance

For millennia, the mathematical community operated in an economy of proof scarcity. Rigorous proofs were precious, rare artifacts forged through decades of human labor.

We are hurtling toward an economy of proof abundance.

In this impending reality, automated agents and theorem provers can generate tens of thousands of airtight, machine-verified proofs weekly. Every single one of these proofs is logically sound. Not a single step violates the rules of formal logic.

Yet, almost none of them can be digested by a human mind.

graph TD
    A["Era of Proof Scarcity (Past)"] -->|"Few Proofs, High Understanding"| B["Healthy Scientific Metabolism"]
    C["Era of Proof Abundance (Future)"] -->|"Infinite AI Proofs, Zero Understanding"| D["Severe Proof Indigestion"]
    
    D --> E["Academic Journals Overwhelmed"]
    D --> F["Students Unable to Extract Concepts"]
    D --> G["Black-Box Mathematics Disconnected from Reality"]

    style A fill:#2e3b2b,stroke:#b8bb26,stroke-width:1px,color:#fbf1c7
    style B fill:#283935,stroke:#8ec07c,stroke-width:1px,color:#fbf1c7
    style C fill:#3b291a,stroke:#fe8019,stroke-width:1px,color:#fbf1c7
    style D fill:#3c1f1e,stroke:#fb4934,stroke-width:2px,color:#fbf1c7
    style E fill:#3c3836,stroke:#fb4934,color:#ebdbb2
    style F fill:#3c3836,stroke:#fb4934,color:#ebdbb2
    style G fill:#3c3836,stroke:#fb4934,color:#ebdbb2

This structural pathology is what Terence Tao labels Proof Indigestion.

Academic journals will not possess the human bandwidth to peer-review the conceptual significance of these synthetic submissions. Doctoral students will find themselves marooned in a sea of verified facts with no intuitive compass to guide them. Mathematics risks deteriorating into an inscrutable oracle: we will know that certain statements are true, but we will have lost the ability to explain why.


6. The Leiden Declaration and the Blackboard Test: Reclaiming Scientific Value

How can humanity navigate this deluge without sacrificing the soul of scientific discovery?

A pivotal step forward came with the formulation of the Leiden Declaration on AI and Mathematics. Developed by leading figures across the international mathematical community, the declaration charts a concrete path toward safeguarding scientific inquiry:

flowchart LR
    A["The Leiden Framework"] --> B["1. Absolute Transparency & AI Auditing"]
    A --> C["2. Shifting Academic Prestige to Digestion"]
    A --> D["3. The Blackboard Test for Publication"]
    A --> E["4. Restructuring Mathematical Pedagogy"]

    style A fill:#3a3220,stroke:#fabd2f,stroke-width:2px,color:#fbf1c7
    style B fill:#3c3836,stroke:#83a598,color:#ebdbb2
    style C fill:#3c3836,stroke:#d3869b,color:#ebdbb2
    style D fill:#2e3b2b,stroke:#b8bb26,stroke-width:2px,color:#fbf1c7
    style E fill:#3c3836,stroke:#8ec07c,color:#ebdbb2

1. Mandatory Transparency and Disclosure

Authors must explicitly document where, how, and to what degree artificial intelligence tools were employed in their research. Submitting machine-generated proofs under the guise of unassisted human labor is not merely academic dishonesty; it poisons the scientific well with uncurated, indigestible artifacts.

2. Rewriting the Academic Reward System

The scientific community must dismantle the romantic cult of raw problem-solving speed. Academia must stop lionizing solely the individual who first planted a flag on an open problem.

Instead, universities, grant agencies, and prize committees must allocate prestige and tenure to those who engage in the unglamorous labor of digestion, exposition, and canonicalization. Writing a lucid expository guide that demystifies a Byzantine proof must be recognized as equal in scholarly value to discovering the original theorem.

3. The "Blackboard Test"

Perhaps the most potent countermeasure discussed is the Blackboard Test.

Under this standard, a manuscript cannot be accepted for publication in a premier journal—even if an interactive theorem prover on a supercomputer certifies its formal correctness—unless the human authors can stand in front of a physical chalkboard and deliver a transparent, conceptually coherent lecture to their peers.

The authors must explain the historical context, delineate the central strategic intuition, identify the crucial friction points, and defend why the proof hangs together. If the author cannot defend the ideas without reading an AI prompt transcript, the paper does not belong in the scientific canon.

4. Reimagining Mathematical Education

The crisis in research mirrors the crisis in the classroom. If software can instantaneously solve calculus and linear algebra problem sets, evaluating students purely on their ability to output correct final answers is obsolete.

Mathematics pedagogy must pivot from answer-oriented mechanics to conceptual articulation, dialectic defense, and structural reasoning. Students should be evaluated on their capacity to explain the architectural choices behind a solution, identify hidden assumptions, and communicate mathematical concepts to others.


The Soul of Mathematics: Understanding Over Verification

Mathematics has never been a sport about generating true statements for their own sake. If the goal of mathematics were merely to accumulate verified tautologies, we could configure a simple script to enumerate every formal consequence of ZFC axioms until the heat death of the universe.

Mathematics is a uniquely human quest to find harmony, pattern, and meaning in the abstract fabric of reality.

As artificial intelligence systems grow exponentially more capable, the challenge facing humanity is not whether we can build machines that out-calculate us. The true challenge is whether we have the institutional courage and wisdom to preserve human comprehension at the center of the scientific enterprise.

Without understanding, mathematics ceases to be a beacon of enlightenment—and becomes nothing more than a catalog of incomprehensible miracles.


FAQ (Frequently Asked Questions)

:::details What is Terence Tao's core warning about AI in mathematics? Fields Medalist Terence Tao warns that rapid advancements in AI will fundamentally alter the incentives and values of the mathematical community. By automating proof generation and formal verification, AI risks creating "proof indigestion"—an unmanageable flood of logically sound, machine-verified proofs that humans cannot understand, contextualize, or teach. :::

:::details How does Goodhart's Law apply to artificial intelligence solving math problems? Goodhart's Law states that when a measure becomes a target, it ceases to be a good measure. In mathematics, solving open problems was historically a reliable proxy for scientific progress because human researchers had to build theories, teach students, and write clear explanations along the way. AI games this proxy metric by delivering raw solutions, decoupling problem-solving from communal understanding. :::

:::details What is the "Blackboard Test" proposed for mathematical papers in the AI era? The Blackboard Test is an evaluation threshold requiring authors to stand before an audience of peer mathematicians at a physical chalkboard and clearly explain the conceptual intuition, structure, and history behind their proof. If the authors cannot articulate the reasoning themselves, the paper cannot be published, even if verified by formal proof software like Lean. :::