Jack J. Garzella

Original Article: Julia Stadlmann and bounded gaps between primes

Warning: After hearing the feedback of experts, I now believe that this article overstates the contributions of some of the papers discussed. Please read the updated version for a more accurate portrayal.

In this post, I will chronicle the story of some recent advances towards the twin primes conjecture. I believe this story is important because I think it gives us a glimpse into the future of what math research might look like in the age of AI. So after telling you the real story, I'll give a counterfactual story of what I imagine it might have been like in a world without AI. In both stories, Julia Stadlmann is the main character, because her ideas are key ingredients to all of the recent advances.

To set the context for our story, I want to introduce one bit of jargon. There is a condition in the literature called DHL[k,2]\mathrm{DHL}[ k,2]. It does not matter what this means at all, for us DHL[k,2]\mathrm{DHL}[ k,2] is a black box. However, what you should know is that if DHL[k,2]\mathrm{DHL}[ k,2] holds, then there are infinitely many primes such that the gap between them is less than H(k)H(k), where H(k)H(k) is given by the OEIS sequence A008407.

For example, the Polymath project proved that DHL[50,2]\mathrm{DHL}[ 50,2], showing that we have infinitely many bounded gaps of length 246. If someone ever can prove DHL[2,2]\mathrm{DHL}[ 2,2], then the twin primes conjecture is true.

Disclaimer: I am not an expert in analytic number theory or in bounded gaps between primes. The perspective of this post is as a mathematician who is interested in the effect of AI on the pace of research. The information in this post all comes from my reading of informal expositions of the material (such as introductions) and discussing the results with AI chatbots. I take responsibility for any errors or imprecisions, and I welcome corrections from experts.

The Real Life Story

September 1, 2023: Stadlmann releases her paper On primes in arithmetic progressions and bounded gaps between many primes, which will henceforth be known to us as Staldmann's EE paper (EE is for Equidistribution Estimates). This makes an improvement to a key technical ingredient for proofs in the area, and is (as far as I can tell) the first advance in this class of problems in years.

This paper has a title based on it's main application, which is a related but different statement about gaps between many primes (not just two). An outside observer might suspect that these results have relevance to bounded gaps between primes, but it might not be obvious even to an expert. Or maybe it would be immediately obvious to an expert, as a non-expert myself I can't tell. Either way, one might suspect that Stadlmann herself had an idea of where this might go next.

August 31, 2026: Stadlmann releases her paper Bounded gaps between primes, which proves DHL[49,2]\mathrm{DHL}[ 49,2] and thus gives the prime gap bound of 240 (i.e., there are infinitly pairs of primes with gap 240 or less). The proof is human-made, and gives an interesting and novel way to overcome one of the tradeoffs in the Polymath proof. It is the first improvement since Polymath, which also makes it significant. This paper likely belongs in a top 5 journal given that Stadlmann's EE paper was published in Advances in Mathematics, and this paper's title result is better. Could this be an Annals paper? As a non-expert, I can't quite tell, but the point is that it's a really good paper.

This paper will henceforth be referred to as Stadlmann's main paper, because it is central to the story. Do not confuse it with the EE paper 😊.

Stadlmann's main paper uses ideas from Stadlmann's EE paper in a crucial way, including adapting one of the equidistribution estimates arguments.

Finally, Stadlmann's main paper is clearly the start of something and not the end, because it admits in the introduction that with more computational resources the result could likely be improved.

September 1, 2026: Shiva Kintali, using AI with custom orchestration, releases an improvement which improves on Stadlmann's main paper slightly, proving DHL[48,2]\mathrm{DHL}[ 48,2] and gives a prime gap bound of 236.

Despite not currently working in academia, Kintali is no crank–he has a PhD from Georgia Tech, and clearly has sufficient expertise to understand proofs in the area. However, has not previously worked on bounded gaps in primes, which would have made it exceedingly hard to write such a paper in the pre-AI age. The paper has an AI disclosure that describes the collaboration with AI, which in my opinion is compatible with the Leiden declaration.

The ideas in the paper aren't doing anything but using Stadlmann's ideas and being a bit more careful, and this is freely admitted on Kintali's website. So even if/when this gets published, it doesn't deserve to be in a top 5 journal or anything like that.

September 3, 2026: AxiomMath releases a preprint which improves Stadlmann's methods, proving DHL[45,2]\mathrm{DHL}[ 45,2] and thus giving a prime bound gap of 212. This comes with an automatically generated Lean proof. Like Kintali, the main ideas are due to Stadlmann. They say they rely primarily on numerical computations and optimization.

There is no AI disclosure in this paper, so it's really not clear what role fronteir LLM agents played. The only thing that is mentioned is that AxiomMath's internal tool created the Lean proof.

September 3, 2026: Ingo Althofer writes on his "Math with AI diary" that he has used GPT to "milk" Stadlmann's argument into a proof of DHL[44,2]\mathrm{DHL}[ 44,2], though as of this writing the linked preprint is very sparse, has many the hallmarks of LLM-generation, and outlines an argument for DHL[46,2]\mathrm{DHL}[ 46,2]. Althofer claims that the ideas are all Stadlmann's, and he is only the "milkman".

September 3, 2026: Coinciding with the release of Astra, OpenAI announces a proof of DHL[40,2]\mathrm{DHL}[ 40,2], obtaining a prime gap bound of 186. The proof uses a different idea to Stadlmann's main paper. As far as I can tell, it seems to be weakining a standard assumption in these arguments from something about smooth numbers to pairs (tuples?) of numbers that are "triply densely divisible". To me, this seems like the sort of thing that would be pretty technical and a pain in the butt for a human to do; on the other hand, the paper is only 39 pages so maybe my impression is wrong.

The OpenAI paper does not use Stadlmann's main paper, but it does use arguments and ideas from Stadlmann's EE paper in a crucial way. Stadlmann's EE paper is the only paper that is post-2017 and cited by the OpenAI paper.

The date on the paper (August 30) is before Stadlmann's main paper was available publically. However, Stadlmann's main paper is mentioned as independent work in the paper. This suggests that OpenAI and Julia Stadlmann were mutually aware of each other's proofs before either got published.

Reading between the lines, one imagines that the following scenario may have happened: first, OpenAI mathematicians reached out to Stadlmann before releasing their work, and discovered the mutually independent progress. Stadlmann might not have been quite intending to release her paper yet; likely, OpenAI wanted to release their work by the release of Astra which was already set, and both parties agreed that Stadlmann would write up and release her paper before the release, causing Stadlmann to have to finish her paper quickly, hence the result without doing numerical optimizations that the "milkmen" were able to do in days using AI.

Discovering independent work on the same problem and arranging for a simultaneous or close-to-simultaneous release is fairly common in mathematical practice. Operating on a deadline measured in weeks and not months or years is very uncommon in mathematical practice.

September 3, 2026: Anthropic employee Levent Alpöge discloses via Twitter/X that Claude has obtained a prime gap bound of 188, presumably by showing DHL[41,2]\mathrm{DHL}[ 41,2]. The argument is not publically available; based on Alpöge's comments, this seems to be because the argument has been sent to be looked over by an expert before public release. Also based on Alpöge's comments, it seems like the argument might be pretty similar to the ChatGPT one.

If the world was stuck in 2023

Imagine that the world was stuck in 2023, when AI could not contribute substantially to mathematics research. What might have happened instead?

As I'm not an expert in analytic number theory, these counterfactuals aren't meant to be super technically accurate; vaguely, I'm using the kk in DHL[k,2]\mathrm{DHL}[ k,2] as a rough measure for progress, and construction this based on my observations about how progress happens in my own areas of expertise.

In what follows, we well set 2026 to be Year 0, and count up from Year 0, rather than using counterfactual dates.

Year 0, November: Stadlmann's paper comes out with a proof of DHL[48,2]\mathrm{DHL}[ 48,2], having more time to run computations since there wasn't a deadline.

Year 1: Stadlmann gives a bunch of talks about progress on bounded gaps between primes. Stadlmann starts to become a trendy name in analytic number theory.

Year 1, May: Quanta Magazine writes an article about Stadlmann's work. It gives a profile of her, and tells the story of all recent progress on bounded gaps between primes, much like previous articles.

Year 1, August: A workshop is convened with relevant experts in a variety of topics. Stadlmann is either an organizer or the guest of honor. She gives a lecture series, and many of the participants work on numerical optimizations during the workshop.

Year 2, April: The workshop collaboration releases a paper which proves DHL[47,2]\mathrm{DHL}[ 47,2].

Year 2: Stadlmann starts a tenure-track job. At this point, she is a regular attendee and has various other papers having to do with equidistribution estimates.

Year 2: Stadlmann takes a PhD student, who let's call Thor. Thor's thesis project is to improve on DHL[47,2]\mathrm{DHL}[ 47,2]. Thor spends a lot of time scratching and clawing and working on numerical optimizations, and gets DHL[46,2]\mathrm{DHL}[ 46,2], the smallest possible improvement.

Year 6: Thor graduates, releasing the DHL[46,2]\mathrm{DHL}[ 46,2] proof into the wild. Meanwhile, Stadlmann has been making other interesting advances in the field.

Year 6 Thor, Julia Stadlmann, and one or two of the experts from the conference publish a paper improving the result to DHL[44,2]\mathrm{DHL}[ 44,2]. At this point, the idea from Stadlmann's original paper is well and truly "milked". People move on to doing different things.

Year 8: Julia Stadlmann gets a PhD student who is exceptional, a way better student than Thor, lets call this student Black Widow. Stadlmann gives Black Widow a thesis project whose idea is to find some truly new method for lowering the bound, much like her own thesis. We'll cast aside the idea of whether the idea AI came up with would be the same one that the humans would come up with, and for the sake of argument say that this is the same idea as OpenAI.

Year 12: Black Widow graduates, having successfully showed DHL[41,2]\mathrm{DHL}[ 41,2].

Year 14: Another "milking" process, kicked off by Black Widow, completes with the community having proved DHL[40,2]\mathrm{DHL}[ 40,2] or DHL[39,2]\mathrm{DHL}[ 39,2]. The problem becomes dormant again.

Observations

I'd like to share a few assorted thoughts about the whole situation.

As one parting thought, I hope that the mathematical community can recognize that Stadlmann's fingerprints are all over all of these advances, and give her the credit she deserves.