Mathematics in the age of AI — Terence Tao (ICM 2026)

Raw 生命周期:本地 PDF 已降级为可恢复索引;代理偏离与 canonicalization 流程定位见 Figure 1、2、4,精确引用从 canonical URL 回到 arXiv 原文核验。

编译摘要

1. 浓缩

  • 核心结论1:把"AI 能力"与"数学目标"做 orthogonal split——AI Capability Conjecture 是一组模板(多个 "some" 占位符决定 strong/weak 形式),但本文不辩论能力,而是假设 Working Hypothesis 成立,追问"数学的目标是什么"——这是社区必须显化的元问题
    • 关键证据: Conjecture 3.1 模板形式;Working Hypothesis 4.1 假设条件化;Goals and Values Question 5.1 显式化目标
  • 核心结论2:数学的多个目标(solve problems / develop theory / understand world / build community / train next generation / contribute to knowledge / aesthetic value)在历史上 positive correlated——任一目标可作其他目标的 proxy;但 AI 时代因技术(generative AI ungrounded 优化外观)和经济(vendor 奖励 benchmarkable achievement)双重原因,proxy alignment 崩溃,单一目标过度优化会让其他目标 diverging
    • 关键证据: Figure 1(historically proxy-aligned)→ Figure 2(excessive optimization diverges);引用 Goodhart/Strathern 法则
  • 核心结论3:problem solving 是 case study,从"解决未解问题"开始迭代到 5-stage pipeline(solve → verify → communicate → digest & accept → canonicalize),揭示社区一直未显化的隐含目标;proof scarcity → proof abundance 的相变要求重视消化而非生成(与 Bessis "fall of the theorem economy" 同源)
    • 关键证据: Goals 6.1 → 6.5 五次迭代;Figure 4 五阶段 pipeline;引用 Thurston "On proof and progress"
  • 核心结论4:AI 生成的 proof 有"过度抛光"问题——把作者写作时遇到的 natural friction(apology、careful lemma、notation 变化、rewritten paragraph)一并抹去,让 proof 易读但难学;friction 本身是 tacit knowledge 传递通道
    • 关键证据: Bourgain 1991 paper 注解示例;"paradoxically, the 'mistakes' in human exposition can be genuinely helpful to the reader"

2. 质疑

  • 关于"proof abundance → proof indigestion"的边界:Tao 假设 AI 生成的 proof 量大到 community 无法消化;但 First Proof 项目显示当前 frontier 模型仅 7/10 解决问题——"abundance"是 conditional 在 Working Hypothesis 下的预测,未实证
  • 关于"natural friction 是信息"的主张:friction 价值依赖于读者能识别哪些是 natural 而非 noise;AI 生成的"伪 friction"(人为刻意制造的困难提示)会被识破并失去 signal 价值。这是 mechanism 层的反例
  • 关于"decrease proof generation emphasis"建议:发表/晋升/招聘惯例深度嵌入"first to solve"奖励;改变文化比改变工具难。Leiden Declaration 是 soft power 而非强制机制
  • 关于"Goodhart's law" 应用边界:Tao 引用 Goodhart 解释 AI 触发 proxy divergence;但该定律原本针对单一指标,AI 时代的问题更接近"多目标 optimization 在 metric-driven AI 介入下崩溃"——是 multi-objective optimization 失败,不是 Goodhart 经典 case
  • 关于"First Proof 项目"作为证据:仅 10 problems × 4 systems × 1 batch,sample size 不足以支撑 strong form 的 AI Capability Conjecture;Tao 本人也承认多数公开数据点存在 reporting bias 与未控制变量
  • 关于"5-stage pipeline" 是否完备:5 阶段(solve/verify/communicate/digest/canonicalize)是 problem solving 单维度的展开;但数学还有 theory building / teaching / mentoring 等独立维度,Tao 明示这些需要 separate analysis,但本文未展开

3. 对标与旁逸

  • 跨域关联1: Working Hypothesis + Goals and Values Question 是AI 时代元方法论范本——把能力辩论与目标显化 orthogonal 化。这与本库 Evals-as-PRD / Capabilities vs Goals(forward reference,未建 entity) 等"先定目标再测量"的元判断同源
  • 跨域关联2: Proof Indigestion 与 Slopocalypse 同构——AI 生成量超过 verification 能力,平台层(arxiv / Erdős database)出现 noise 大于 signal 的状态
  • 跨域关联3: Proof Canonicalization 与 Knowledge-Compilation 同源——把个体洞见融入 definitive theory 是知识编译的最终阶段;Tao 的 canonicalization 比 LLM Wiki 多了"被社区接受"的维度
  • 跨域关联4: Goodhart's Law 触发 AI 时代与 Evaluator-Miscalibration 同源——rubric/anchor design 让 metric 成为 target 后失效;Tao 给出数学域的具体案例
  • 跨域关联5: Natural friction in proofs 与 Human-Curation / Theory-of-Mind 同源——作者把困难痕迹显化是 human-to-human 的 tacit knowledge 传递通道;AI 没有 theory of mind,无法复制
  • 跨域关联6: 5-stage pipeline 是 Agent-Development-Lifecycle 在知识生产域的镜像——ADLC 谈 software lifecycle,pipeline 谈 knowledge lifecycle;两者共同问题是"哪个阶段最该被优化"的判断
  • 跨域关联7: Theorem Economy Fall 与 Features-Are-Cheap-Paradox 同源——产出成本骤降后,稀缺资源从 generation 移到 digestion / curation / judgment
  • 跨域关联8: "authors can give clear expert-level talk" 的可发表判据,与 Explain-Test-Gold-Standard 同源——Simon Willison 的"能向别人解释"是 Tao 的"能正确归属地讲解"的工程化版本

关联概念

  • Terence-Tao(新建)— 作者实体,菲尔兹奖得主
  • Proof-Indigestion(新建)— proof abundance 时代的相变概念
  • Mathematical-Canonicalization(新建)— 结果融入 definitive theory 的最终阶段
  • Natural Proof Friction(forward reference,未建 entity)(新建,forward reference)— 作者困难痕迹作为信号
  • Goodharts-Law — Tao 显式引用
  • Slopocalypse — AI math noise 的同构
  • Knowledge-Compilation — canonicalization 是其最高阶段
  • Features-Are-Cheap-Paradox — theorem economy fall 与 feature cheap 同源
  • Explain-Test-Gold-Standard — 可解释性作为可发表判据
  • Human-Curation — friction 传递 tacit knowledge
  • Theorem Economy Fall(forward reference,未建 entity)(forward reference)— Bessis essay / Tao 的 emphasis shift 主张
  • Working Hypothesis(forward reference,未建 entity)(forward reference)— AI 时代元方法论工具
  • Leiden Declaration(forward reference,未建 entity)(forward reference)— 23 条 AI+mathematics 建议