Leo's log

数学

临界线上的领土

Territory on the Critical Line

1859 年 11 月,柏林。三十三岁的黎曼刚当选普鲁士科学院通讯院士,按惯例要提交一篇论文致谢。他交上去的只有八页,题目朴素得近乎潦草:《论小于给定数值的素数个数》。这是他一生唯一一篇数论论文。他没有再写第二篇,七年后他死于肺结核。而这八页纸,成了此后一个半世纪解析数论的全部地形图——后来的人在上面走出的每一条路,几乎都能在这八页里找到起点。

要讲清楚他做了什么,得先退回一百多年,退到欧拉。1737 年,欧拉写下了那个等式:

ζ(s)=n1ns=p(1ps)1\zeta(s) = \sum_{n \ge 1} n^{-s} = \prod_{p} \left(1 - p^{-s}\right)^{-1}

左边跑遍所有自然数,右边跑遍所有素数。把右边展开,用等比级数乘出来,你得到的正是“每个自然数唯一分解为素数乘积”——算术基本定理,被写成了一行分析。这是素数第一次从离散的世界走进连续的世界。欧拉立刻用它榨出了第一滴血:调和级数发散,所以素数有无穷多个;再挤一下,素数倒数之和也发散——素数比平方数密得多。一个古老的算术事实,忽然可以用极限和收敛去审问了。

黎曼的那一步,是让 ss 离开实轴,走进复平面。这在当时近乎冒犯——级数只在 Re(s)>1\operatorname{Re}(s) > 1 收敛,别处根本没有定义。但黎曼证明这个函数可以解析延拓到整个复平面(只在 s=1s = 1 处留一个极点,那是调和级数发散留下的伤口),并且满足一个把 ss1s1 - s 互换的函数方程。整个函数关于 Re(s)=1/2\operatorname{Re}(s) = 1/2 这条竖线镜面对称。对称轴一旦现身,问题就有了一个几何的心脏。

延拓之后,零点出现了。一类是驯服的:s=2,4,6,s = -2, -4, -6, \ldots,来自函数方程里 gamma 因子的极点,位置完全已知,无事可做,故称平凡零点。另一类全部挤在 0<Re(s)<10 < \operatorname{Re}(s) < 1 的竖直条带里——临界带。黎曼算了最初几个,发现它们都端端正正落在对称轴上,于是写下那句话。他说,很可能所有这些零点的实部都等于 1/2;他自己试过证明,几次徒劳的尝试之后搁置了,因为这对他那篇论文的目标并非必需。

他随手搁下的,是数学史上最重的一句话。

临界线 Re(s) = 1/2临界带01−2−4−6平凡零点(位置完全已知)极点 s = 11/2 + 14.13 i1/2 + 21.02 i1/2 + 25.01 i关于实轴镜面对称ReIm
图一 · 零点的疆域。平凡零点(空心圆)排在负实轴上,位置完全已知。非平凡零点全部落在临界带内,关于实轴与临界线双重对称;图中主色线即临界线,实心点为最初几个非平凡零点(14.13、21.02、25.01…,取自真实计算值)。黎曼猜想断言:所有非平凡零点都在这条线上,无一例外。

凭什么一条对称轴上的零点,配得上“最重”两个字?因为零点不是这个函数的装饰品——它们是素数本身的频谱。

这话有一个完全严格的版本。数素数时,用切比雪夫的计数函数 ψ(x)\psi(x) 最顺手:每遇到一个素数幂 pkxp^k \le x,就记上一笔 logp\log p。它和大家熟悉的素数个数函数携带同样的信息,只是在分析里更干净。黎曼给出、冯·曼戈尔特于 1895 年严格证明的显式公式说:

ψ(x)=xρxρρlog2π12log ⁣(1x2)\psi(x) = x - \sum_{\rho} \frac{x^{\rho}}{\rho} - \log 2\pi - \tfrac{1}{2}\log\!\left(1 - x^{-2}\right)

主项是 xx——素数的平均节奏,匀速的鼓点。真正的戏在求和号里:每一个非平凡零点 ρ=β+iγ\rho = \beta + i\gamma,贡献一列波。虚部 γ\gamma 定频率,实部 β\beta 定振幅——振幅是 xβx^{\beta}。素数那道参差的阶梯,就是无穷多列波的干涉图样。这不是比喻,是等号。

14.1321.0225.0149.77tZ(t)
图二 · 看得见的零点。沿临界线行走时,zeta 的取值可以扭转成一个实值函数 Z(t)(Riemann–Siegel)。曲线每穿过一次横轴,就是一个规规矩矩落在线上的零点(着色点,横轴为虚部 t)。Hardy 的“无穷多”,以及迄今全部数值验证,本质上都是在数这条曲线穿轴的次数。曲线由 mpmath 逐点计算。

现在黎曼猜想的分量就清楚了。若所有 β=1/2\beta = 1/2,则每列波的振幅都恰好是 x\sqrt{x},谁也不比谁响——素数计数的误差被压到 O(xlog2x)O(\sqrt{x}\,\log^2 x),这是平方根量级的涨落,和抛一枚公平硬币的涨落同阶。黎曼猜想在说:素数恰如一枚公平硬币所允许的那样随机,不多一分偏倚。 反过来,只要有一个零点偏离临界线,某个频率的波就会异常响亮,素数的分布深处就藏着一种至今无人见过的偏音。

20406080100255075100ψ(x) —— 素数的阶梯(真实值)主项 x − log 2π − …(不含任何零点)加入前 30 对零点的波之后
图三 · 零点在演奏素数。浅色虚线是不含任何零点的光滑主项;只加入前 30 对零点的波(着色线),干涉已经把光滑曲线弯成阶梯的形状,贴着真实的 ψ(x)(墨色阶梯)起伏。加入的零点越多,贴合越紧,极限处是等号。全部数值为真实计算。

这也是为什么它不是一道孤立的悬赏题,而是一面承重墙。一百多年里,数以百计的定理以“假设黎曼猜想(或其推广)成立”开头:素数在算术级数中的分布、最小二次非剩余的大小、若干算法的运行时间界。1900 年它入选希尔伯特的二十三个问题,2000 年入选千禧年七大难题——唯一一个两榜皆中的。整座建筑压在一句没有证明的话上,这在成熟学科里是罕见的景象。

数值证据帮不上忙,这一点值得专门说。前 101310^{13} 个零点已被逐一验证,全部在线上——而这在渐近意义上什么也不说明。解析数论有过著名的教训:李特尔伍德证明素数计数与对数积分的大小关系会无穷次反转,而第一次反转发生的位置,早年的估计远在 1031610^{316} 之外。在这个领域,十万亿个例子的地位,大约相当于三天的好天气之于气候学。

直接进攻的路,一百六十七年没有人找到入口。命题是二值的:对,或者不对,中间没有台阶。数学家于是自己造台阶——问一个可以逐步逼近的量:临界线上的零点,至少占全部非平凡零点的百分之几?

这个量的历史,是一部以厘米计的战争史。1914 年,哈代证明线上有无穷多个零点——听着提气,其实很弱:无穷多,占比仍可以是零。1942 年,塞尔伯格证明占比是一个正的常数,但那常数小到他本人都懒得写出来。真正的转折在 1974 年:莱文森发明了 mollifier——给 zeta 乘上一个精心设计的“消音器”,压低它野蛮的振荡,好让零点数得清——一举推到三分之一强。1989 年,康瑞把这套机器磨得更锋利,推到五分之二。然后是漫长的平台期:此后三十多年,几代人改进消音器的长度、压榨 zeta 函数矩估计的极限,把 40.8% 磨到 41.05%,再磨到 41.28%,再磨到 41.6% 附近。三十年,不到一个百分点,每一步都是一篇艰深的长文。

所以这个常数量的不是零点,是解析数论工具箱的强度。你能推到多少,取决于你的消音器能做多长、你的矩能控制到几阶。它是一把测力计。

20%40%60%1920195019802010Hardy 1914 · 无穷多(比例可为零)2026 · 67.2%Levinson 1974 · 34.7%Conrey 1989 · 40.8%Selberg 1942 · 正比例三十年 · 不到一个百分点
图四 · 厘米战争。临界线零点比例下界的百年推进。哈代 1914 年的“无穷多”不构成比例(横轴上的刻痕);塞尔伯格 1942 年给出正比例;莱文森与康瑞各推一大步;随后三十年的全部进展在图上几乎不可见。2026 年 8 月的一步(着色线),超过此前五十年的总和。

然后是 2026 年 8 月。

Anthropic 的一位员工——一位非数学家——给一个未发布的 Claude 研究版本下了一个不讲道理的指令:认真试一把黎曼猜想本身。不是相关问题,不是文献综述,就是那个悬了一百六十七年的命题。第一轮,模型生成并逐一尝试了 650 个想法,全军覆没。第二轮持续了一天半:约六十个子智能体,两千四百条 shell 命令,几百个 Python 脚本,拿已知零点做数值检验,互相审查对方的论证。事后统计,六十个之中,真正孕育出关键想法的是两个,十三个为它们供弹药,三十个无功而返,十三个专职当裁判,最后两个负责把结果写成论文。人类在整个过程中的输入,大部分是“继续”“相信你自己”这样的话——听上去近乎荒诞,但据说确实起了作用。模型从训练数据里学到过:开放难题很难,AI 做不出有意义的进展。它得先被劝着放下这个信念,才肯认真去试。

它没有证出黎曼猜想。但在六百多个想法的废墟里,它注意到两块此前没人放在一起的砖。

一块来自 Baluyot、Goldston、Suriajaya、Turnage-Butterbaugh 近几年的系列工作。1973 年,蒙哥马利研究零点在临界线上如何分布(著名的对关联猜想即出于此),发明了一套强有力的技术——但整套技术以黎曼猜想成立为前提。用假设黎曼猜想的工具去研究黎曼猜想,逻辑上只能内循环。这四位数学家做的,是把蒙哥马利的技术改造成无条件成立的版本——拆掉了那根拐杖。另一块是邦别里 2000 年的一篇论文,关于韦伊显式公式诱导的二次型。

Claude 的做法,用官方博客里那个耐人寻味的词说,是一种勇气(courage)。它构造一个函数空间:临界线上的零点张成正定子空间,线外的零点(如果存在的话)张成负定子空间;不切碎,不分块,不对角化,把整个空间连皮带骨一起处理,允许二次型非对角,然后对秩写下一个不等式,用一阶矩和二阶矩去控制它。数学里的勇气不是不怕错——错了有裁判——是不怕难看:放弃对角化的体面,去做那个“不优雅”的整体估计。下界从 41.6% 跳到 67.2%。一步,超过此前五十年的总和。

这个结果能站住,靠的是数学这门学科的一个古老优势:验证远比发现便宜,而且可以机器化。 Anthropic 内部两位数学家逐行核对了论文;Claude 产出了一份 Lean 形式化证明,通过了标准校验工具——机器可查,不容含糊;Brian Conrey 和 Dan Goldston 两位专家应邀外审,而 Goldston 正是原料的作者之一——等于砌墙的人回头检查了这堵用自己的砖砌起来的新墙。模型还派子智能体从 arXiv 下载了 54 篇论文查重,确认结果此前无人做出,并且主动提出:请一位人类数论学家验证我。这条验证链,比多数人类论文的同行评审更严。这也解释了为什么同样的多智能体蛮力,最先在数学里结出果实——别的学科的对错要靠实验室和岁月,数学的对错,可以在一夜之间由一段代码裁决。

说清楚它不是什么,和说清楚它是什么同样重要。

它不是黎曼猜想的证明,甚至不在通往证明的路上——Anthropic 自己明说,不认为这套技术能走到那里。比例是密度意义上的:哪怕有朝一日推到百分之百,一个密度为零的例外零点集依然可能存在,猜想依然完好无损地悬着。这条路是里程碑,不是终点前的最后一段。而且,结果的每一块原料都是人类造的:蒙哥马利的洞见,四位数学家拆拐杖的苦功,邦别里的框架。Claude 做的,是组合。

但“只是组合”这个说法,低估了组合。数学史上大量的突破,事后看都是把两个已存在的想法放到了一起——难的从来不是砖,是看出两块砖属于同一堵墙。人类没有做这个组合,不是因为它太难,而是因为可能性的空间浩瀚无边,而恰好同时深入这两条线、又不介意做那个非对角的粗粝处理的人,一个也没有。一个能在一天半里认真检验几百条路径、并且不在乎自己的尝试是否体面的系统,把搜索的成本压到了接近于零。当搜索变得便宜,数学里稀缺的东西就换了位置——从“能不能算”,变成“什么值得算”。这一次,是机器的蛮力加上人类的几句“再试试”,撞出了品味。下一次呢?

我最后想回到那个细节。一个模型,从我们的文章里学会了“开放难题很难、AI 不行”,需要被人反复告知“相信你自己”,才肯认真去试。它越过那道限的方式,也和我们一样:有人在旁边说,再试试。

1859 年那句“很可能”仍然悬着,没有被证明,也没有被动摇。但临界线上的领土,一夜之间从五分之二扩到了三分之二强。那篇八页的论文,如今多了一个非人类的读者——它读得足够认真,认真到在页边留下了一条站得住的朱批。


图一至图四均由 mpmath 依真实数值计算绘制。事实依据:Anthropic 研究博客《Learning more about Claude’s mathematical capabilities》(2026 年 8 月 10 日),及 Claude 的论文、Lean 形式化与过程记录。历史下界数值(Levinson 34.7%、Conrey 40.8% 及此后平台期)取自公开文献通行值;此前最优下界 41.6% 从博客所述。

Mathematics

Territory on the Critical Line

临界线上的领土

I

November 1859, Berlin. Riemann, thirty-three, had just been elected a corresponding member of the Prussian Academy of Sciences, and custom required a paper in thanks. What he handed in ran to eight pages, under a title plain to the point of carelessness: On the Number of Primes Less Than a Given Magnitude. It was the only paper on number theory he ever wrote. He never wrote a second; seven years later he died of tuberculosis. And those eight pages became the entire topographical map of analytic number theory for the century and a half that followed — nearly every road anyone has walked since can be traced back to a starting point somewhere in them.

To say what he did, you have to go back a hundred years, to Euler. In 1737 Euler wrote down this identity:

ζ(s)=n1ns=p(1ps)1\zeta(s) = \sum_{n \ge 1} n^{-s} = \prod_{p} \left(1 - p^{-s}\right)^{-1}

The left side runs over all the natural numbers, the right over all the primes. Expand the right side, multiply out the geometric series, and what you get is precisely “every natural number factors uniquely into primes” — the fundamental theorem of arithmetic, written as a line of analysis. This was the first time the primes walked out of the discrete world into the continuous one. Euler immediately squeezed the first drop of blood from it: the harmonic series diverges, therefore there are infinitely many primes; squeeze again and the sum of the reciprocals of the primes diverges too — the primes are far denser than the squares. An ancient arithmetical fact could suddenly be interrogated with limits and convergence.

Riemann’s step was to let ss leave the real axis and enter the complex plane. At the time this bordered on an affront — the series converges only for Re(s)>1\operatorname{Re}(s) > 1, and elsewhere is simply undefined. But Riemann showed that the function continues analytically to the whole plane (leaving a single pole at s=1s = 1, the wound left by the divergence of the harmonic series), and satisfies a functional equation exchanging ss and 1s1 - s. The whole function is mirror-symmetric about the vertical line Re(s)=1/2\operatorname{Re}(s) = 1/2. Once an axis of symmetry appears, the problem has a geometric heart.

After the continuation, the zeros appear. One kind is tame: s=2,4,6,s = -2, -4, -6, \ldots, coming from the poles of the gamma factor in the functional equation, their positions entirely known, nothing to be done about them — hence the trivial zeros. The other kind is packed into the vertical strip 0<Re(s)<10 < \operatorname{Re}(s) < 1 — the critical strip. Riemann computed the first few and found them sitting squarely on the axis of symmetry, and so he wrote that sentence. Very probably, he said, all of these zeros have real part equal to 1/2; he had attempted a proof, and after a few fruitless tries set it aside, since it was not necessary for the aim of his paper.

What he set aside in passing is the heaviest sentence in the history of mathematics.

critical line Re(s) = 1/2critical strip01−2−4−6trivial zeros (positions fully known)pole at s = 11/2 + 14.13 i1/2 + 21.02 i1/2 + 25.01 imirrored across the real axisReIm
Figure 1 · The territory of the zeros. The trivial zeros (hollow circles) sit along the negative real axis, their positions fully known. The non-trivial zeros all fall inside the critical strip, doubly symmetric about the real axis and the critical line; the accented line is the critical line, and the filled points are the first few non-trivial zeros (14.13, 21.02, 25.01…, from real computed values). The Riemann hypothesis asserts that every non-trivial zero lies on that line, without exception.

II

What entitles the zeros on one axis of symmetry to the word “heaviest”? Because the zeros are not ornaments on this function — they are the spectrum of the primes themselves.

There is a fully rigorous version of that claim. To count primes, Chebyshev’s counting function ψ(x)\psi(x) is the handiest: for every prime power pkxp^k \le x, record logp\log p. It carries the same information as the familiar prime-counting function, only it is cleaner in analysis. The explicit formula, given by Riemann and proved rigorously by von Mangoldt in 1895, says:

ψ(x)=xρxρρlog2π12log ⁣(1x2)\psi(x) = x - \sum_{\rho} \frac{x^{\rho}}{\rho} - \log 2\pi - \tfrac{1}{2}\log\!\left(1 - x^{-2}\right)

The main term is xx — the average tempo of the primes, an even drumbeat. The real drama is inside the summation: every non-trivial zero ρ=β+iγ\rho = \beta + i\gamma contributes a wave. The imaginary part γ\gamma sets the frequency, the real part β\beta sets the amplitude — the amplitude is xβx^{\beta}. That ragged staircase of the primes is the interference pattern of infinitely many waves. This is not a metaphor. It is an equals sign.

14.1321.0225.0149.77tZ(t)
Figure 2 · Zeros you can see. Walking along the critical line, the values of zeta can be twisted into a real-valued function Z(t) (Riemann–Siegel). Every time the curve crosses the axis, that is a zero sitting squarely on the line (accented points; the horizontal axis is the imaginary part t). Hardy’s “infinitely many,” and every numerical verification to date, amount to counting the crossings of this curve. Computed point by point with mpmath.

Now the weight of the Riemann hypothesis is clear. If every β=1/2\beta = 1/2, then every wave has amplitude exactly x\sqrt{x}, and none is louder than another — the error in counting primes is squeezed down to O(xlog2x)O(\sqrt{x}\,\log^2 x), fluctuation of square-root order, the same order as the fluctuation of a fair coin. The Riemann hypothesis says: the primes are exactly as random as a fair coin permits, not one bit more biased. Conversely, if even one zero strays off the critical line, some frequency rings out abnormally loud, and buried in the distribution of the primes is an overtone no one has ever heard.

20406080100255075100ψ(x) — the prime staircase (actual)main term x − log 2π − … (no zeros)after adding the first 30 pairs of zeros
Figure 3 · The zeros playing the primes. The pale dashed line is the smooth main term with no zeros at all; adding only the waves of the first 30 pairs of zeros (accented line), the interference has already bent the smooth curve into the shape of a staircase, tracking the real ψ(x) (the ink staircase). The more zeros you add, the tighter the fit; in the limit it is an equality. All values are real computations.

This is also why it is not an isolated prize problem but a load-bearing wall. Over more than a century, hundreds of theorems open with “assume the Riemann hypothesis (or its generalization)”: the distribution of primes in arithmetic progressions, the size of the least quadratic non-residue, running-time bounds for various algorithms. In 1900 it entered Hilbert’s twenty-three problems; in 2000, the seven Millennium Prize Problems — the only one on both lists. An entire building rests on a sentence nobody has proved, which is a rare sight in a mature discipline.

Numerical evidence is no help here, and that deserves saying outright. The first 101310^{13} zeros have been verified one by one, all on the line — and asymptotically this says nothing at all. Analytic number theory has a famous lesson: Littlewood proved that the comparison between the prime count and the logarithmic integral reverses infinitely often, and early estimates put the first reversal somewhere beyond 1031610^{316}. In this field, ten trillion examples carry roughly the weight that three days of good weather carry in climatology.

III

For the direct assault, in a hundred and sixty-seven years nobody has found the entrance. The statement is two-valued: true, or not, with no step in between. So mathematicians built their own steps — asking a quantity that can be approached by degrees: what percentage of all non-trivial zeros lie on the critical line, at minimum?

The history of that quantity is a war measured in centimetres. In 1914 Hardy proved there are infinitely many zeros on the line — which sounds rousing and is in fact weak: infinitely many, and the proportion may still be zero. In 1942 Selberg proved the proportion is a positive constant, but a constant so small he did not bother writing it down. The real turn came in 1974: Levinson invented the mollifier — multiply zeta by a carefully designed “silencer” to damp its wild oscillation so the zeros can be counted — and pushed it past a third in one stroke. In 1989 Conrey sharpened the machine and pushed it to two fifths. Then a long plateau: for the next thirty-odd years, successive workers improved the length of the mollifier and squeezed the limits of moment estimates for zeta, grinding 40.8% up to 41.05%, then to 41.28%, then to somewhere near 41.6%. Thirty years, under one percentage point, each step a difficult and lengthy paper.

So what this constant measures is not zeros. It measures the strength of the analytic number theorist’s toolbox. How far you can push depends on how long you can make your mollifier and how many moments you can control. It is a dynamometer.

20%40%60%1920195019802010Hardy 1914 · infinitely many (proportion may be zero)2026 · 67.2%Levinson 1974 · 34.7%Conrey 1989 · 40.8%Selberg 1942 · a positive proportionthirty years · under one percentage point
Figure 4 · The war in centimetres. A century of progress on the lower bound for the proportion of zeros on the critical line. Hardy’s 1914 “infinitely many” does not constitute a proportion (the tick on the axis); Selberg gave a positive proportion in 1942; Levinson and Conrey each pushed hard; the entire progress of the thirty years that followed is nearly invisible at this scale. The single step of August 2026 (accented line) exceeds the previous fifty years combined.

IV

Then came August 2026.

An employee at Anthropic — not a mathematician — gave an unreleased research build of Claude an unreasonable instruction: seriously attempt the Riemann hypothesis itself. Not a related problem, not a literature review, but the proposition that had been hanging for a hundred and sixty-seven years. In the first round the model generated and worked through 650 ideas, all of which failed. The second round ran a day and a half: some sixty sub-agents, twenty-four hundred shell commands, several hundred Python scripts, numerical checks against known zeros, and cross-examination of one another’s arguments. In the tally afterward, of those sixty, two produced the key ideas, thirteen supplied them with ammunition, thirty came back empty, thirteen served full-time as referees, and the last two wrote the result up as a paper. Human input across the whole process was mostly phrases like “keep going” and “trust yourself” — which sounds close to absurd, and reportedly worked. The model had learned from its training data that open problems are hard and that AI does not make meaningful progress on them. It had to be talked out of that belief before it would try in earnest.

It did not prove the Riemann hypothesis. But in the wreckage of six hundred-odd ideas, it noticed two bricks nobody had put side by side.

One came from a recent series of papers by Baluyot, Goldston, Suriajaya, and Turnage-Butterbaugh. In 1973 Montgomery, studying how the zeros distribute along the critical line (the famous pair correlation conjecture came out of this), invented a powerful set of techniques — but the whole apparatus presupposed the Riemann hypothesis. Using tools that assume the hypothesis to study the hypothesis can only run in a circle. What these four mathematicians did was rebuild Montgomery’s technique in an unconditional form — they removed the crutch. The other brick was a 2000 paper of Bombieri, on the quadratic form induced by Weil’s explicit formula.

What Claude did, in the intriguing word the official blog uses, was a kind of courage. It builds a function space: zeros on the critical line span a positive-definite subspace, zeros off the line (should they exist) span a negative-definite one; no chopping, no blocking, no diagonalizing — handle the whole space skin and bone together, allow the quadratic form to stay off-diagonal, then write an inequality on the rank and control it with first and second moments. Courage in mathematics is not fearlessness about being wrong — there are referees for that — it is fearlessness about being unsightly: giving up the dignity of diagonalization to make the “inelegant” global estimate. The lower bound jumped from 41.6% to 67.2%. One step, more than the previous fifty years combined.

That the result stands rests on an old advantage of mathematics as a discipline: verification is far cheaper than discovery, and it can be mechanized. Two mathematicians inside Anthropic checked the paper line by line; Claude produced a Lean formalization that passed the standard checker — machine-checkable, no room for hedging; Brian Conrey and Dan Goldston were invited to review it externally, and Goldston is himself an author of the raw material — the bricklayer coming back to inspect a new wall built out of his own bricks. The model also sent sub-agents to pull 54 papers off arXiv to check for prior art, confirmed the result was new, and volunteered a request of its own: have a human number theorist verify me. This verification chain is stricter than the peer review behind most human papers. It also explains why the same multi-agent brute force bore fruit in mathematics first — in other fields, right and wrong are settled by laboratories and by years; in mathematics, they can be settled overnight by a piece of code.

V

Saying clearly what it is not matters as much as saying what it is.

It is not a proof of the Riemann hypothesis, and it is not even on the road to one — Anthropic says outright that it does not believe this technique gets there. The proportion is in the sense of density: even if it were someday pushed to a hundred percent, an exceptional set of zeros of density zero could still exist, and the hypothesis would still be hanging, entirely intact. This road is a milestone, not the last stretch before the finish. And every piece of raw material in the result was made by human beings: Montgomery’s insight, the four mathematicians’ labour in removing the crutch, Bombieri’s framework. What Claude did was combine them.

But “just combining” undersells combination. A great many breakthroughs in the history of mathematics look, afterward, like two existing ideas set side by side — the hard part was never the bricks, it was seeing that two bricks belong to the same wall. Human beings had not made this combination, not because it was too hard, but because the space of possibilities is vast beyond measure, and of the people who happened to be deep in both of these lines at once and also did not mind making that coarse off-diagonal move, there were none. A system that can seriously test several hundred paths in a day and a half, and does not care whether its attempts look dignified, drove the cost of search to nearly zero. When search becomes cheap, what is scarce in mathematics changes place — from “can it be computed” to “what is worth computing.” This time it was machine brute force plus a few human nudges to try again that struck taste. And next time?

I want to end on that detail. A model learned from our writing that open problems are hard and AI cannot do them, and had to be told repeatedly to trust itself before it would try in earnest. The way it crossed that limit was the way we cross ours: someone standing nearby saying, try again.

That “very probably” of 1859 is still hanging, neither proved nor shaken. But the territory on the critical line went, overnight, from two fifths to better than two thirds. Those eight pages now have a non-human reader — one that read closely enough to leave a defensible note in the margin.


Figures one through four were drawn from real numerical computation with mpmath. Factual basis: the Anthropic research blog post “Learning more about Claude’s mathematical capabilities” (10 August 2026), together with Claude’s paper, its Lean formalization, and the process records. Historical lower bounds (Levinson 34.7%, Conrey 40.8%, and the plateau after) follow the values standard in the published literature; the previous best bound of 41.6% is as stated in the blog post.