Jack J. Garzella

Julia Stadlmann and bounded gaps between primes

In this post, I will chronicle the story of some recent advances towards the twin primes conjecture. I believe this story is important because I think it gives us a glimpse into the future of what math research might look like in the age of AI.

So after telling you the real story, I'll give a counterfactual story of what I imagine it might have been like in a world without AI. In both stories, Julia Stadlmann is the main character, because her ideas are key ingredients to all of the recent advances.

To set the context for our story, I want to introduce one bit of jargon. There is a condition in the literature called DHL[k,2]\mathrm{DHL}[ k,2]. It does not matter what this means at all, for us DHL[k,2]\mathrm{DHL}[ k,2] is a black box. However, what you should know is that if DHL[k,2]\mathrm{DHL}[ k,2] holds, then there are infinitely many primes such that the gap between them is less than H(k)H(k), where H(k)H(k) is given by the OEIS sequence A008407.

For example, the Polymath project proved that DHL[50,2]\mathrm{DHL}[ 50,2], showing that we have infinitely many bounded gaps of length 246. If someone ever can prove DHL[2,2]\mathrm{DHL}[ 2,2], then the twin primes conjecture is true.

The other key thing you need to know is that one key step in proving DHL[k,2]\mathrm{DHL}[ k,2] is evaluating a certain very complicated integral on a certain domain SS. I'll call this the key integral in the discussion below. All of the proofs end up doing this using computers in some way or other. According people with expertise in the area, this integral is very finnicky; being off by even a little bit can lead to wildly wrong results. Polymath gets around this difficulty by choosing SS to be a simplex, and using a closed formula for the integral.

Disclaimer: I am not an expert in analytic number theory or in bounded gaps between primes. The perspective of this post is as a mathematician who is interested in the effect of AI on the pace of research. The information in this post all comes from my reading of informal expositions of the material (such as introductions) and discussing the results with AI chatbots and experts. I take responsibility for any errors or imprecisions, and I welcome corrections from experts.

The Real Life Story

September 1, 2023: Stadlmann releases her paper On primes in arithmetic progressions and bounded gaps between many primes, which will henceforth be known to us as Staldmann's EE paper (EE is for Equidistribution Estimates). This makes an improvement to a key technical ingredient for proofs in the area, and is (as far as I can tell) the first advance in this class of problems in years.

This paper has a title based on it's main application, which is a related but different statement about gaps between many primes (not just two). An outside observer might suspect that these results have relevance to bounded gaps between primes, but it might not be obvious even to an expert. Or maybe it would be immediately obvious to an expert, as a non-expert myself I can't tell. Either way, one might suspect that Stadlmann herself had an idea of where this might go next.

August 31, 2026: Stadlmann releases her paper Bounded gaps between primes, which proves DHL[49,2]\mathrm{DHL}[ 49,2] and thus gives the prime gap bound of 240 (i.e., there are infinitly pairs of primes with gap 240 or less). The proof is human-made, and gives an interesting and novel way to overcome one of the tradeoffs in the Polymath proof. It is the first improvement since Polymath, which also makes it significant. As an outsider, I would guess that this paper likely belongs in a top 5 journal given the notoriety of the result, and that Stadlmann's EE paper was published in Advances in Mathematics.

This paper will henceforth be referred to as Stadlmann's main paper, because it is central to the story. Do not confuse it with the EE paper 😊.

Stadlmann's main paper uses ideas from Stadlmann's EE paper in a crucial way, including adapting one of the equidistribution estimates arguments.

Finally, to compute the key integral, Stadlmann's main paper uses a novel recursive algorithm whose bottleneck step is multiplication of matrices of rational numbers. This is clearly the start of something and not the end, because it admits in the introduction that with more computational resources the result could likely be improved.

September 1, 2026: Shiva Kintali, using AI with custom orchestration, releases an improvement which improves on Stadlmann's main paper slightly, proving DHL[48,2]\mathrm{DHL}[ 48,2] and gives a prime gap bound of 236.

Despite not currently working in academia, Kintali is no crank–he has a PhD from Georgia Tech, and clearly has sufficient expertise to understand proofs in the area. However, has not previously worked on bounded gaps in primes, which would have made it exceedingly hard to write such a paper in the pre-AI age. The paper has an AI disclosure that describes the collaboration with AI, which in my opinion is compatible with the Leiden declaration.

The analytic ideas in the paper aren't doing anything but using Stadlmann's ideas and maybe being a bit more careful, and this is freely admitted on Kintali's website. So even if/when this gets published, it probably isn't in a top 5 journal or whatever.

Kintali's paper claims that it evaluates the key integral exactly, but few details are given in the writeup, despite the fact that this is a key step. I had an AI analyze the code, and seems that the integral is evaluated by using a more or less naive triangulation of the domain of integration and evaluating integrals on simplices a la Polymath.

September 3, 2026: AxiomMath releases a preprint which improves Stadlmann's methods, proving DHL[45,2]\mathrm{DHL}[ 45,2] and thus giving a prime bound gap of 212. This comes with an automatically generated Lean proof. Like Kintali, the main ideas are due to Stadlmann. They say they rely primarily on numerical computations, especially parameter searches for their improvement.

There is no AI disclosure in this paper, so it's really not clear what role fronteir LLM agents played. The only thing that is mentioned is that AxiomMath's internal tool created the Lean proof.

AxiomMath also evaluates the key integral exactly via a similar method to Kintali - they have some way of finding a triangulation of the domain SS, and then they evaluate on simplices, using the machinery they made during their formalization of the Polymath proof.

September 3, 2026: Ingo Althofer writes on his "Math with AI diary" that he has used GPT to "milk" Stadlmann's argument into a proof of DHL[44,2]\mathrm{DHL}[ 44,2]. Althofer claims that the ideas are all Stadlmann's, and he is only the "milkman".

As of this writing the linked preprint is very sparse, has many of the hallmarks of LLM-generation, and outlines an argument for DHL[46,2]\mathrm{DHL}[ 46,2]. To be fair, that preprint does not seem to be intended for publication.

The writeup does not actually provide any indication on how to evaluate the key integral; it says that the values "should be evaluated" in a rigorous way, as if the AI hasn't actually done the computation. It's not clear the key integral was computed or not.

September 3, 2026: Coinciding with the release of Astra, OpenAI announces a proof of DHL[40,2]\mathrm{DHL}[ 40,2], obtaining a prime gap bound of 186. The proof uses a different method than Stadlmann's main paper. This approach, which involves the use of so-called "3-densely divisible moduli", was known to the experts since Polymath, certainly Stadlmann and many of the Polymath authors knew about it.

Despite not using Stadlmann's main paper, OpenAI does use arguments and ideas from Stadlmann's EE paper in a crucial way. Stadlmann's EE paper is the only paper that is post-2017 and cited by the OpenAI paper. Clearly this technical ingredient is very important for future advances.

The switch to "3-densely divisible moduli" requires evaluating a much harder key integral. Evaluating such an integral exactly would be a big challenge, which would likely require interesting novel mathematics research. However, OpenAI skips this entirely, instead making extensive use of numerical computations, and justifying the computation with a rigorous bound. There are a bunch of of potential pitfalls with this approach (for example, software bugs), but in theory it could have been done by any of the previous contributions. Evaluating such an integral numerically can be considered the "brute force" way to solve the problem. Until OpenAI, all previous contributions had decided not to do this, either because it wasn't feasible with available compute, or because it wasn't as interesting as developing the math research to evalute the integral exactly.

I had an AI analyze OpenAI's code, and the repo seems to do a good job of avoiding the pitfalls of floating-point rounding error. It has a separate writeup for the numerical analysis that translates the bounds used in the proof into something that can be run on a computer. It uses rigorous bounds for floating-point that are better than a generic interval arithmetic library, and more custom than something like FPTaylor or Gappa. It takes care to make sure that it runs operations that will be more numerically stable, but again being a bit more custom than something like Herbie. It uses Arb/FLINT for to do ball arithmetic for certain steps, even checking for a particular bug in a particular version.

In theory, this kind of brute force computation could have been done 10 years ago; you'd have needed 1-2 experts in numerical computations and a bunch of compute. But this is a lot easier with AI, and I think this is clearly where the majority of the contribution lives.

The Lean proof that OpenAI provides has nothing about any of these numerical methods. I consider this to be a big omission, given that the whole strategy of this proof is to brute force it, and the brute force isn't verified at all. Some more Lean-minded people might even consider the current AI artifacts to be not a proof.

The date on the paper (August 30) is before Stadlmann's main paper was available publically. However, Stadlmann's main paper is mentioned as independent work in the paper. This suggests that OpenAI and Julia Stadlmann may have been mutually aware of each other's proofs before either got published.

Reading between the lines, one might guess that the following happened: first, OpenAI mathematicians reached out to Stadlmann before releasing their work, and discovered the mutually independent progress. Stadlmann might not have been quite intending to release her paper yet; likely, OpenAI wanted to release their work by the release of Astra which was already set, and both parties agreed that Stadlmann would write up and release her paper before the release, causing Stadlmann to have to finish her paper quickly, hence the result without doing parameter searches that the "milkmen" were able to do in days using AI.

Discovering independent work on the same problem and arranging for a simultaneous or close-to-simultaneous release is fairly common in mathematical practice. Operating on a deadline measured in weeks and not months or years is very uncommon in mathematical practice.

September 3, 2026: Anthropic employee Levent Alpöge discloses via Twitter/X that Claude has obtained a prime gap bound of 188, presumably by showing DHL[41,2]\mathrm{DHL}[ 41,2]. The argument is not publically available; based on Alpöge's comments, this seems to be because the argument has been sent to be looked over by an expert before public release. Also based on Alpöge's comments, it seems like the argument might be pretty similar to the ChatGPT one.

If the world was stuck in 2023

Imagine that the world was stuck in 2023, when AI could not contribute substantially to mathematics research. What might have happened instead?

As I'm not an expert in analytic number theory, these counterfactuals aren't meant to be super technically accurate; vaguely, I'm using the kk in DHL[k,2]\mathrm{DHL}[ k,2] as a rough measure for progress, and construction this based on my observations about how progress happens in my own areas of expertise.

In what follows, we well set 2026 to be Year 0, and count up from Year 0, rather than using counterfactual dates.

Year 0, November: Stadlmann's paper comes out with a proof of DHL[48,2]\mathrm{DHL}[ 48,2], having more time to run computations since there wasn't a deadline.

Year 1: Stadlmann gives a bunch of talks about progress on bounded gaps between primes. Stadlmann starts to become a trendy name in analytic number theory.

Year 1, May: Quanta Magazine writes an article about Stadlmann's work. It gives a profile of her, and tells the story of all recent progress on bounded gaps between primes, much like previous articles.

Year 1, August: A workshop is convened with relevant experts in a variety of topics. Stadlmann is either an organizer or the guest of honor. She gives a lecture series, and many of the participants work on numerical optimizations during the workshop.

Year 2, April: The workshop collaboration releases a paper which proves DHL[47,2]\mathrm{DHL}[ 47,2].

Year 2: Stadlmann starts a tenure-track job. At this point, she is a regular attendee and has various other papers having to do with equidistribution estimates.

Year 2: Stadlmann takes a PhD student, who let's call Thor. Thor's thesis project is to improve on DHL[47,2]\mathrm{DHL}[ 47,2]. Thor spends a lot of time scratching and clawing and optimizing Stadlmann's recursive algorithm, and gets DHL[46,2]\mathrm{DHL}[ 46,2], the smallest possible improvement.

Year 6: Thor graduates, releasing the DHL[46,2]\mathrm{DHL}[ 46,2] proof into the wild. Meanwhile, Stadlmann has been making other interesting advances in the field.

Year 6 Thor, Julia Stadlmann, and one or two of the experts from the conference publish a paper improving the result to DHL[44,2]\mathrm{DHL}[ 44,2]. At this point, the idea from Stadlmann's original paper is well and truly "milked". People move on to doing different things.

Year 8: Julia Stadlmann gets a PhD student who is exceptional, a way better student than Thor, lets call this student Black Widow. Stadlmann gives Black Widow the thesis project of trying to use "3-densely divisible moduli" to improve on bounded gaps between primes.

Year 12: Black Widow graduates, having successfully showed DHL[41,2]\mathrm{DHL}[ 41,2].

Year 14: Another "milking" process, kicked off by Black Widow, completes with the community having proved DHL[40,2]\mathrm{DHL}[ 40,2] or DHL[39,2]\mathrm{DHL}[ 39,2]. The problem becomes dormant again.

Observations

I'd like to share a few assorted thoughts about the whole situation.

As one parting thought, I hope that the mathematical community can recognize that Stadlmann's fingerprints are all over all of these advances, and give her the credit she deserves.

Update: This is a revised version of this post. After some feedback from experts, I now believe that the original version overstated the contributions of some of the papers involved. In particular, the original verison did not mention the "key integral" computation. For transparency purposes, you can read the original verison here.