OpenAI Published a Batch of Math Results, and So Did I

An AI did hours of tedious math for me, but the real effort went into making this essay about the experience sound like a human wrote it.

Cover Image for OpenAI Published a Batch of Math Results, and So Did I

In September 2026, OpenAI posted on GitHub formal Lean proofs for finite-time blow-up in the Navier-Stokes and Euler equations. On October 6, they released 722 manuscripts produced by an internal model. This is the domain of top human minds and massive compute clusters. I also published a batch of math results. Mine were computed by Claude on my laptop in a few hours. I only sent a few prompts, and even after writing this essay, I am still not quite sure what the math problem itself is. But this essay, the one you are reading now, took many rounds and several versions of back and forth with AI. The contrast is stark: having an AI do math felt natural, but getting it to write a few thousand words about it turned into a tug-of-war.

In 2026, using AI to set new records in mathematics is already a routine event, something I knew well. The target Claude picked this time was on Erich Friedman's website. Friedman is a mathematician who taught for 26 years at Stetson University in the US. Since retiring in 2018, he has focused on recreational math. He maintains a site called the Packing Center, which lists over a hundred geometric packing problems. These problems are about fitting circles, squares or polygons into a container, or placing points in space as compactly or evenly as possible. For each problem, a page lists the best known arrangement for each size, next to the name of the person who found it. If someone finds a better arrangement, the name on the page changes. Friedman updates these pages himself, by hand, based on submissions he receives by email.

Claude found a problem about placing points in three-dimensional space. The table on the page stopped at 30 points, but the rules allowed for up to 50. The 20 cells from 31 to 50 had never been filled. There were two reasons to pick this problem. First, nobody had noticed the page. Second, the problem is not technically difficult and has little mathematical value, so nobody wanted to spend time on it. Claude filling in these 20 cells was like sweeping out a neglected corner.

The computation side went unusually smoothly. In that world, every step is either clearly right or clearly wrong. Starting on September 28, Claude used four cores on my laptop to reproduce the table's existing best results in just five and a half minutes. The first real run took about five hours, but the computer restarted during the night and the results were lost. The rerun took about seven and a half hours. By the morning of September 29, all 20 results were done.

Every number in those 20 results was checked exactly using fractions. Mathematical verification is rigorous and requires no human intervention. To be safe, I asked Claude if there were any issues with the results. It spun up an independent AI, one that could not see the code, to review them. This review agent only needed the final coordinates to check the distances. It caught a formatting error: the last digit had been rounded up in every case. After the fix, I submitted the results from my own email on September 29, under the name Limin Ge. The code and results were later open-sourced on GitHub. Now, the 20 entries are waiting to be added.

The whole thing should have ended there. The numbers are precisely reproducible. An independent check can find errors. Machines operate with ease in domains governed by clear rules. But when I tried to write this process into an article, the trouble began.

On September 29, I had Claude write the first draft. It then did its own pass to remove the AI tone, but it still read as machine-made to me.

I turned to Gemini. I called four of its models, gave them five different angles, and had them each write a draft, for a total of 20. I had hoped that different models and different prompts would produce some variety in expression. After reading all 20, they felt the same to me, and every one of them had flaws. Numbers have right and wrong answers and can be verified by another machine. But a piece of writing that expresses a view or an experience has no single correct answer. When faced with a task that has no standard answer, a model will fall back on the average of its training data, toward the safest possible phrasing. It generates paragraphs that are grammatically correct but tedious to read. In the world of words, the machine pursues only one thing: mediocrity without mistakes.

Factual errors turned out to be the easiest problem to fix in these drafts. I had another AI check the 20 drafts, and it found that only one was free of invented or incorrect facts. Among the fabrications, the models had turned a Monday into "the weekend," added a dramatic "woke up to find the data gone," and even put words I never said inside quotation marks. Math results have an independent review process because coordinates have precise values. Fabrications in text can also be caught by a program, but they reveal how the model works: it does not know what a fact is. It only knows that in human stories, data loss often happens in the morning, and the protagonist tends to make a few remarks at that point.

To revise the drafts, I built a tool for myself. Editing in a traditional chat window often has cascading effects; asking a model to rewrite one paragraph can change the tone of the entire piece. My tool limits the changes to a local area: I strike through sentences I do not like on a web page, write a note, and the model rewrites only the paragraphs that were struck. By October 8, I had gone through four rounds of striking out text (12, 8, 9, and 4 places respectively), yielding seven different versions.

Even with this tool for local rewriting, the process was full of resistance. Sometimes I would write in a note the exact sentence I wanted to use, and the model would change it to something else anyway. A word I struck in one round would reappear, as if nothing had happened, a round later. The model seemed to have a stubborn attachment to certain words and sentence structures. Once freed from the precise constraints of mathematics, it began to wander the probability grid of language, replacing my expressions with what it deemed more appropriate. In the end, I had to set a rule: any words I write must be used exactly as written. It was the only way I could preserve my own patience.

This is a strange contrast. I, along with thousands of other people, sit in front of a screen every day, wrestling with the writing habits of a language model. People think that putting AI to work in scientific research is about expanding the frontiers of human knowledge. The reality is that the machines can already handle the verifiable parts of the process on their own. The most draining work is what cannot be quantified. The time I spent was not on checking facts, which another AI did better than me, but on arguing over the phrasing of a sentence or the feel of a word, fighting against the statistically determined mediocrity of the model. When machines have filled dozens of empty cells in a table that nobody was paying attention to, what people are left doing for themselves is using AI to make AI's work look less like AI's output.


Get new articles by email

One email when a new article lands. Nothing else.