In this post, I will chronicle the story of some recent advances towards the twin primes conjecture. I believe this story is important because I think it gives us a glimpse into the future of what math research might look like in the age of AI.
So after telling you the real story, I'll give a counterfactual story of what I imagine it might have been like in a world without AI. In both stories, Julia Stadlmann is the main character, because her ideas are key ingredients to all of the recent advances.
To set the context for our story, I want to introduce one bit of jargon. There is a condition in the literature called . It does not matter what this means at all, for us is a black box. However, what you should know is that if holds, then there are infinitely many primes such that the gap between them is less than , where is given by the OEIS sequence A008407.
For example, the Polymath project proved that , showing that we have infinitely many bounded gaps of length 246. If someone ever can prove , then the twin primes conjecture is true.
The other key thing you need to know is that one key step in proving is evaluating a certain very complicated integral on a certain domain . I'll call this the key integral in the discussion below. All of the proofs end up doing this using computers in some way or other. According people with expertise in the area, this integral is very finnicky; being off by even a little bit can lead to wildly wrong results. Polymath gets around this difficulty by choosing to be a simplex, and using a closed formula for the integral.
Disclaimer: I am not an expert in analytic number theory or in bounded gaps between primes. The perspective of this post is as a mathematician who is interested in the effect of AI on the pace of research. The information in this post all comes from my reading of informal expositions of the material (such as introductions) and discussing the results with AI chatbots and experts. I take responsibility for any errors or imprecisions, and I welcome corrections from experts.
September 1, 2023: Stadlmann releases her paper On primes in arithmetic progressions and bounded gaps between many primes, which will henceforth be known to us as Staldmann's EE paper (EE is for Equidistribution Estimates). This makes an improvement to a key technical ingredient for proofs in the area, and is (as far as I can tell) the first advance in this class of problems in years.
This paper has a title based on it's main application, which is a related but different statement about gaps between many primes (not just two). An outside observer might suspect that these results have relevance to bounded gaps between primes, but it might not be obvious even to an expert. Or maybe it would be immediately obvious to an expert, as a non-expert myself I can't tell. Either way, one might suspect that Stadlmann herself had an idea of where this might go next.
August 31, 2026: Stadlmann releases her paper Bounded gaps between primes, which proves and thus gives the prime gap bound of 240 (i.e., there are infinitly pairs of primes with gap 240 or less). The proof is human-made, and gives an interesting and novel way to overcome one of the tradeoffs in the Polymath proof. It is the first improvement since Polymath, which also makes it significant. As an outsider, I would guess that this paper likely belongs in a top 5 journal given the notoriety of the result, and that Stadlmann's EE paper was published in Advances in Mathematics.
This paper will henceforth be referred to as Stadlmann's main paper, because it is central to the story. Do not confuse it with the EE paper 😊.
Stadlmann's main paper uses ideas from Stadlmann's EE paper in a crucial way, including adapting one of the equidistribution estimates arguments.
Finally, to compute the key integral, Stadlmann's main paper uses a novel recursive algorithm whose bottleneck step is multiplication of matrices of rational numbers. This is clearly the start of something and not the end, because it admits in the introduction that with more computational resources the result could likely be improved.
September 1, 2026: Shiva Kintali, using AI with custom orchestration, releases an improvement which improves on Stadlmann's main paper slightly, proving and gives a prime gap bound of 236.
Despite not currently working in academia, Kintali is no crank–he has a PhD from Georgia Tech, and clearly has sufficient expertise to understand proofs in the area. However, has not previously worked on bounded gaps in primes, which would have made it exceedingly hard to write such a paper in the pre-AI age. The paper has an AI disclosure that describes the collaboration with AI, which in my opinion is compatible with the Leiden declaration.
The analytic ideas in the paper aren't doing anything but using Stadlmann's ideas and maybe being a bit more careful, and this is freely admitted on Kintali's website. So even if/when this gets published, it probably isn't in a top 5 journal or whatever.
Kintali's paper claims that it evaluates the key integral exactly, but few details are given in the writeup, despite the fact that this is a key step. I had an AI analyze the code, and seems that the integral is evaluated by using a more or less naive triangulation of the domain of integration and evaluating integrals on simplices a la Polymath.
September 3, 2026: AxiomMath releases a preprint which improves Stadlmann's methods, proving and thus giving a prime bound gap of 212. This comes with an automatically generated Lean proof. Like Kintali, the main ideas are due to Stadlmann. They say they rely primarily on numerical computations, especially parameter searches for their improvement.
There is no AI disclosure in this paper, so it's really not clear what role fronteir LLM agents played. The only thing that is mentioned is that AxiomMath's internal tool created the Lean proof.
AxiomMath also evaluates the key integral exactly via a similar method to Kintali - they have some way of finding a triangulation of the domain , and then they evaluate on simplices, using the machinery they made during their formalization of the Polymath proof.
September 3, 2026: Ingo Althofer writes on his "Math with AI diary" that he has used GPT to "milk" Stadlmann's argument into a proof of . Althofer claims that the ideas are all Stadlmann's, and he is only the "milkman".
As of this writing the linked preprint is very sparse, has many of the hallmarks of LLM-generation, and outlines an argument for . To be fair, that preprint does not seem to be intended for publication.
The writeup does not actually provide any indication on how to evaluate the key integral; it says that the values "should be evaluated" in a rigorous way, as if the AI hasn't actually done the computation. It's not clear the key integral was computed or not.
September 3, 2026: Coinciding with the release of Astra, OpenAI announces a proof of , obtaining a prime gap bound of 186. The proof uses a different method than Stadlmann's main paper. This approach, which involves the use of so-called "3-densely divisible moduli", was known to the experts since Polymath, certainly Stadlmann and many of the Polymath authors knew about it.
Despite not using Stadlmann's main paper, OpenAI does use arguments and ideas from Stadlmann's EE paper in a crucial way. Stadlmann's EE paper is the only paper that is post-2017 and cited by the OpenAI paper. Clearly this technical ingredient is very important for future advances.
The switch to "3-densely divisible moduli" requires evaluating a much harder key integral. Evaluating such an integral exactly would be a big challenge, which would likely require interesting novel mathematics research. However, OpenAI skips this entirely, instead making extensive use of numerical computations, and justifying the computation with a rigorous bound. There are a bunch of of potential pitfalls with this approach (for example, software bugs), but in theory it could have been done by any of the previous contributions. Evaluating such an integral numerically can be considered the "brute force" way to solve the problem. Until OpenAI, all previous contributions had decided not to do this, either because it wasn't feasible with available compute, or because it wasn't as interesting as developing the math research to evalute the integral exactly.
I had an AI analyze OpenAI's code, and the repo seems to do a good job of avoiding the pitfalls of floating-point rounding error. It has a separate writeup for the numerical analysis that translates the bounds used in the proof into something that can be run on a computer. It uses rigorous bounds for floating-point that are better than a generic interval arithmetic library, and more custom than something like FPTaylor or Gappa. It takes care to make sure that it runs operations that will be more numerically stable, but again being a bit more custom than something like Herbie. It uses Arb/FLINT for to do ball arithmetic for certain steps, even checking for a particular bug in a particular version.
In theory, this kind of brute force computation could have been done 10 years ago; you'd have needed 1-2 experts in numerical computations and a bunch of compute. But this is a lot easier with AI, and I think this is clearly where the majority of the contribution lives.
The Lean proof that OpenAI provides has nothing about any of these numerical methods. I consider this to be a big omission, given that the whole strategy of this proof is to brute force it, and the brute force isn't verified at all. Some more Lean-minded people might even consider the current AI artifacts to be not a proof.
The date on the paper (August 30) is before Stadlmann's main paper was available publically. However, Stadlmann's main paper is mentioned as independent work in the paper. This suggests that OpenAI and Julia Stadlmann may have been mutually aware of each other's proofs before either got published.
Reading between the lines, one might guess that the following happened: first, OpenAI mathematicians reached out to Stadlmann before releasing their work, and discovered the mutually independent progress. Stadlmann might not have been quite intending to release her paper yet; likely, OpenAI wanted to release their work by the release of Astra which was already set, and both parties agreed that Stadlmann would write up and release her paper before the release, causing Stadlmann to have to finish her paper quickly, hence the result without doing parameter searches that the "milkmen" were able to do in days using AI.
Discovering independent work on the same problem and arranging for a simultaneous or close-to-simultaneous release is fairly common in mathematical practice. Operating on a deadline measured in weeks and not months or years is very uncommon in mathematical practice.
September 3, 2026: Anthropic employee Levent Alpöge discloses via Twitter/X that Claude has obtained a prime gap bound of 188, presumably by showing . The argument is not publically available; based on Alpöge's comments, this seems to be because the argument has been sent to be looked over by an expert before public release. Also based on Alpöge's comments, it seems like the argument might be pretty similar to the ChatGPT one.
Imagine that the world was stuck in 2023, when AI could not contribute substantially to mathematics research. What might have happened instead?
As I'm not an expert in analytic number theory, these counterfactuals aren't meant to be super technically accurate; vaguely, I'm using the in as a rough measure for progress, and construction this based on my observations about how progress happens in my own areas of expertise.
In what follows, we well set 2026 to be Year 0, and count up from Year 0, rather than using counterfactual dates.
Year 0, November: Stadlmann's paper comes out with a proof of , having more time to run computations since there wasn't a deadline.
Year 1: Stadlmann gives a bunch of talks about progress on bounded gaps between primes. Stadlmann starts to become a trendy name in analytic number theory.
Year 1, May: Quanta Magazine writes an article about Stadlmann's work. It gives a profile of her, and tells the story of all recent progress on bounded gaps between primes, much like previous articles.
Year 1, August: A workshop is convened with relevant experts in a variety of topics. Stadlmann is either an organizer or the guest of honor. She gives a lecture series, and many of the participants work on numerical optimizations during the workshop.
Year 2, April: The workshop collaboration releases a paper which proves .
Year 2: Stadlmann starts a tenure-track job. At this point, she is a regular attendee and has various other papers having to do with equidistribution estimates.
Year 2: Stadlmann takes a PhD student, who let's call Thor. Thor's thesis project is to improve on . Thor spends a lot of time scratching and clawing and optimizing Stadlmann's recursive algorithm, and gets , the smallest possible improvement.
Year 6: Thor graduates, releasing the proof into the wild. Meanwhile, Stadlmann has been making other interesting advances in the field.
Year 6 Thor, Julia Stadlmann, and one or two of the experts from the conference publish a paper improving the result to . At this point, the idea from Stadlmann's original paper is well and truly "milked". People move on to doing different things.
Year 8: Julia Stadlmann gets a PhD student who is exceptional, a way better student than Thor, lets call this student Black Widow. Stadlmann gives Black Widow the thesis project of trying to use "3-densely divisible moduli" to improve on bounded gaps between primes.
Year 12: Black Widow graduates, having successfully showed .
Year 14: Another "milking" process, kicked off by Black Widow, completes with the community having proved or . The problem becomes dormant again.
I'd like to share a few assorted thoughts about the whole situation.
First, it kinda sucks for Stadlmann, instead of getting to go on tour and get hailed as an expert, she has the next three projects stolen by people whose biggest contribution is that they have a lot more compute (read: money) than a postdoc.
In hindsight, it's no surprise Claude and ChatGPT seem to have found roughly the same argument (even though we aren't completely sure). Daniel Litt talked about this recently, where he argues that doing math this way could lead to lower quality mathematics.
You can tell that the "If we were stuck in 2023" mathematical methodology has a certian flow to it. Someone comes up with a good idea, and then the community milks all the low-hanging fruit over 2-5 years. There are some mathematicians who more or less only do milking, and many more start out this way. Once all the low hanging fruit has been milked, the problem goes dormant until a really good student or senior mathematician puts a lot of effort into improving it (usually by doing the same amount of work as in a PhD thesis), after which the community has more new ideas that can be milked.
Building on the last point, it is also pretty clear that even the pace of these incremental improvements (much less the GPT/Claude proofs) absolutely breaks the previous system in math. Milking is now essentially instant (if you have the cash).
After seeing this, I think as a pre-tenure mathematician, it's important to be really careful what you say in public. Clearly, there are people watching the arXiv feed who will throw tokens and compute at whatever they find interesting. Anything related to something too popular is liable to be tokenmaxxed. Will it become necessary to give your paper a more obscure title in the hopes that people will miss it? Should you even be posting stuff on the arXiv at all before it's been accepted for publication anymore? I don't have answers, but we really have to rethink how we operate.
Participation in pure math as a whole has been broadened: as I said before Kintali obviously not a crank, but he also isn't working in academia. Academia doesn't really know how to deal with this type of contributor. In 2023's world, such a person would have likely been unable to have the resources to contribute to research and do his day job, but now he can just go ahead and let the bots run.
As one parting thought, I hope that the mathematical community can recognize that Stadlmann's fingerprints are all over all of these advances, and give her the credit she deserves.
Update: This is a revised version of this post. After some feedback from experts, I now believe that the original version overstated the contributions of some of the papers involved. In particular, the original verison did not mention the "key integral" computation. For transparency purposes, you can read the original verison here.