Essay · September 2026

Stop Asking Whether It Thinks. Ask What It Can Reach.

What I learned from the week a machine proved a theorem about moving water, why I think we have been using the wrong word for this technology, and the place I would rather point it.

A librarian, and two piles of paper

Start in a library, because that is where this actually starts.

A librarian at the University of Chicago named Don Swanson spent years worrying about a problem that sounds like a riddle and turns out to be a fact about the world. Science publishes far more than any person can read. So it is ordinary — not rare — for one group of researchers to establish that A affects B, and for a second group, in a different field, reading different journals, going to different conferences, to establish that B affects C, and for nobody, ever, to write down the sentence connecting A to C. The knowledge is public. It is sitting on the shelves. It is undiscovered anyway, because no one person read both shelves. He gave the condition a name: undiscovered public knowledge.

Then he went and did it. He read one literature on the blood of people with Raynaud’s syndrome, a condition where the small vessels in the fingers clamp shut painfully in the cold. He read a second literature, which did not cite the first, on what dietary fish oil does to blood. Neither set of authors had read the other. Put side by side, they licensed a hypothesis nobody had stated: that fish oil might help. Swanson ran no experiment. He wrote the sentence, and published it.

I want to be accurate about what happened next, because the honest version is the useful one. A small double-blind trial three years later found a benefit in primary Raynaud’s and none in the secondary form. His second case, magnesium and migraine, drew two trials in the same journal in the same year that disagreed with each other. The field he founded has spent the decades since arguing about how to even evaluate itself. So this is not a story about a librarian who was always right. It is a story about where a hypothesis came from. It came from reading, and from nothing else.

Hold that shape in your head. Everything below is the same shape at a different scale.

The week in September

On the eighth of September, OpenAI published a paper claiming that an internal system of theirs had proved that a three-dimensional fluid, starting from rest, under a smooth push, with its energy staying finite the whole time, can tear itself into a singularity in finite time. A hundred and sixty-six pages. A machine-checked formalization alongside it. The company said the run began on the first of the month, that the winning group ran on the order of ten thousand agents at once, and that the proof arrived about eighty-eight hours later.

About twelve hours before that announcement, Tristan Buckmaster of NYU posted a statement. He and Levent Alpöge, a mathematician who works at Anthropic, had been working privately for about a year, and had just released results of their own on closely related equations, obtained in mid-August and machine-verified a week after that. His statement says he was told, in calls that weekend, that an internal OpenAI model had proved the forced Navier–Stokes case, and that when he heard the word forced it was, in his phrase, a bright red flag — because the route through a smooth force was the one he and Alpöge had quietly chosen, and almost nobody else was on it. He says he asked whether the model had been trained on their sessions in OpenAI’s coding tool, where they had been putting their drafts for the length of the project, and that he did not get an answer on training. He describes proposals about authorship that he declined, and an exchange that turned sharp.

OpenAI’s account is that its researchers and its agents did not see any of that work through any means until it was public, and that no specific user data was accessed to solve the problem. Its original post added that, while unlikely, it could not rule out that de-identified data derived from their use of the products had helped improve the models. Two days later the company updated the page to say that an investigation had confirmed the prompts could not have influenced the system in any way, including through training. The researcher at the center of the sharpest exchange has apologized for his words and disputes the characterization of what he asked for.

I am not going to tell you who is right, and I am not writing this to. I will point out the thing I found most striking, which is that Buckmaster himself declines to. His statement says it outright: he has not seen OpenAI’s proof, he does not know what their model did, he does not know whether their data was used, and he is not accusing anyone of anything. He is stating what he was told and when. And he adds that if an OpenAI model really did close that gap, it is a remarkable thing and should be said loudly, with the history intact.

The history, in this case, is not in dispute by anyone. Both sides credit the same two people. Diego Córdoba and Luis Martínez-Zoroa spent years building the program that made this line of attack possible; they are in OpenAI’s own bibliography, and Buckmaster writes in his statement that he believes Martínez-Zoroa deserves a Fields Medal. Nobody’s result here came from nowhere. It came from a road two humans built.

The sentence that made me want to write this

It is in Buckmaster’s statement, and it is about his own work, not OpenAI’s. Describing the first proof his collaborator’s model produced, he writes that it was the most horrendous thing he has ever read. They verified it in Lean, the proof assistant, on the twenty-second of August. And then, in his words, they worked around the clock to understand this proof and turn it into something readable.

Read that again slowly. The machine found it. Two human beings then spent weeks understanding it. He is candid about how that went: he calls one of the resulting write-ups AI slop, apologizes for it, and says the community and the problems deserve better care than he had time to give.

That gap — between a thing being found and a thing being understood — is the whole subject of this essay. It is not a gap that embarrasses the technology. It is a description of what the technology is.

What the instrument actually does

Here is the claim I want to make, and it is deliberately smaller than the one you have been hearing.

A large language model is an instrument for reaching regions of an enormous space of already-existing information that we could not practically reach before. Not a mind. Not a colleague. A reader with an inhuman span, which can hold open more of the record than any of us and notice which two pages are reaching for each other. Swanson did that by hand with two literatures. This does it across a space no person could walk.

This is not a metaphor I invented, and the literature is better than the metaphor. In a study in Nature, researchers trained word embeddings on millions of materials-science abstracts and recovered the structure of the periodic table without being taught any chemistry — and, more to the point, surfaced thermoelectric materials years before those materials were published as discoveries. The authors’ own sentence is the cleanest statement of the idea I have found: latent knowledge regarding future discoveries is, to a large extent, embedded in past publications.

Look at how the strongest systems are actually built and you find the same admission. FunSearch, in Nature, pairs a language model with an automatic evaluator, and its authors are explicit that the pairing exists precisely because models confabulate: the model proposes, the evaluator throws away everything that does not survive checking, and what is left is new. AlphaProof works in Lean for the same reason. Two philosophers of mathematics, writing in a peer-reviewed journal last year, called these systems embodiments of brute-force search, and I do not think that is an insult. It is a specification.

And there is a beautiful piece of evidence for the boundary of the thing. When Nature had working mathematicians put AlphaProof through its paces, the pattern that came back was that it did well on problems built from concepts already defined in Lean’s shared mathematical library, and much less well where they were not. Kevin Buzzard, who is leading the effort to formalize Fermat’s Last Theorem, could not get use out of it at all, because his development is full of bespoke definitions that the library — and so the training — does not contain. His summary was blunt: no AI system is anywhere near useful to him right now.

That is the thesis stated as a measurement. The instrument navigates the space that is already represented. Which is exactly why it is powerful, and exactly why calling it a general intelligence gets the engineering wrong.

The time everyone got this wrong in public

In the autumn of last year a researcher posted that a model had solved a batch of open problems from a well-known database of Erdős’s unsolved questions, and declared that science acceleration via AI had officially begun. It went everywhere. Then the person who maintains the database explained what the word open meant on his site. It meant open to him — problems whose solutions he personally had not yet found in the literature. The model had not solved them. The model had gone and found the papers that already solved them, some of them obscure, and it had done that very well. The head of a rival lab called the whole episode embarrassing.

I think about that story constantly, because nothing in it was fake. The capability was real and, honestly, wonderful. A machine read a corpus nobody had finished reading and returned the connection. The only thing that was wrong was the word we reached for. We said solved. It had searched.

And here is what bothers me about that mistake: searching was the more useful thing. Swanson’s whole career says so.

Where I could be wrong, stated at full strength

If I only told you the part above, I would be doing the thing I am complaining about, in the other direction. So here is the strongest evidence against my own framing.

In July, the same Levent Alpöge published an explicit counterexample to the Jacobian conjecture, a problem that had stood since the nineteen-thirties, crediting an Anthropic model with finding it. A counterexample is not like a proof. You do not have to trust whoever produced it. You substitute the numbers and look. It either is or is not a counterexample, and this one is. And you cannot retrieve from a corpus an object that is not in the corpus.

I do not know how to fit that cleanly into a story about searching a space of existing information, and I am not going to pretend otherwise. The most honest thing I can say is that the space these systems navigate is not only the space of things people have written. It seems to include the space of things constructible from what people have written, which is unimaginably larger, and which we have no map of. That is more than a library. It is still not a mind.

The part that makes it knowledge

Now the detail from September that I find genuinely moving, and that almost nobody wrote about.

OpenAI’s formal proof is checked against problem statements it did not write. The statements were taken, at a pinned version, from an independent repository of formalized open conjectures maintained by a competing lab. So the company that produced the proof did not get to define what counted as proving it. Someone else held the definition, in machine-checkable form, and the proof had to satisfy that.

I would like that to be the most-copied idea of this whole episode. It is the answer to the question everyone keeps asking in the wrong key. You do not need to know whether the machine understood anything. You need the statement of the problem to live somewhere the machine’s owner does not control, and you need the check to be mechanical.

Which is why nothing is settled yet, and why that is fine. The Clay Mathematics Institute has recognized no one; its rules require publication, years of elapsed time, and general acceptance by the field, and its president says the evaluation will be deliberately unhurried. There is a real disagreement among mathematicians about significance, too: the official problem statement offers four alternatives, and while the two that concern breakdown do permit a smooth applied force — so the claim is valid on the face of the text — the unforced question, the one most people mean when they say Navier–Stokes, is still open. OpenAI has said it does not intend to claim the prize.

Terence Tao, who has thought about these equations for most of his career, gave the objection its best form. Getting the answer this way, he told CNN, is a little like watching a movie by jumping from the first ten minutes to the last ten. Technically the plot lines all resolve. Most of the value of the experience is gone. In a lecture this year he made the quieter and more damaging point: the public evidence about what these systems can do is subject to severe reporting bias, because successes are announced and failures are not.

Why the framing is not just semantics

Five days before the mathematics announcement, at a briefing for a different model, OpenAI’s president told reporters it was not unreasonable to feel that we are now in the AGI era. Those two events got welded together in the retelling, and they should not be: different week, different model, different claim.

But the welding is the whole problem, and it is why I care about a word.

If the frame is artificial general intelligence, the only available question is whether the machine is smarter than us, and that question has no engineering answer. It cannot be measured, it cannot be designed against, and it turns every result into a referendum on human worth. If the frame is an instrument for reaching a space, then the questions become answerable and, better, actionable. Which region can it reach? What is represented in that region and what is missing? Who holds the statement of the problem? What is the check, and can the machine’s owner edit it? How much did the search cost, and who can afford it — a fair question here, since one mathematician observed that very few mathematicians will ever have resources at that scale.

Those are questions a graduate student can work on. The other one is a question you can only have opinions about.

Where I would rather point it

Now the science fiction, except that every piece of it has already been built and published, which is why I think it is not fiction so much as an unfinished assignment.

Picture a person’s medical record as what it actually is: a time series. Visits, codes, prescriptions, values, decades long, written by dozens of people who never met each other and were each solving that day’s problem. It is a library, and nobody has read the whole of it — not the patient, and not, in any real sense, any one of their doctors.

Last year a group published a generative model in Nature trained on the records of about four hundred thousand people, which predicts rates for more than a thousand diseases from a person’s history and can generate plausible trajectories two decades forward. They then ran it, unchanged, against the national registry records of nearly two million people in another country. A separate group trained on Danish registry sequences and validated on American veterans’ data to flag pancreatic cancer risk years before diagnosis. This is Swanson’s shape again: A and B in one part of the record, B and C in another, and the sentence connecting them in nobody’s chart.

And now the guardrail, which belongs in the same paragraph as the hope rather than in a footnote after it.

That Nature model’s accuracy fell when it crossed the border, and fell further the farther ahead it looked. Its own authors report that it learned artifacts of how the data was collected — diseases that only ever appear in hospital records were predicted far more often in anyone who had any other hospital record — and they warn against reading it causally. That is what an honest paper looks like. The dishonest version is also on the record: a widely deployed proprietary sepsis alert, evaluated independently at a university health system, scored far below what its vendor had reported and missed about two-thirds of the sepsis cases while firing on nearly a fifth of all hospitalizations. And in the most cited case of all, a risk algorithm used on millions of patients turned out to be predicting cost rather than illness, which quietly meant Black patients had to be sicker to get the same score.

So the instrument does not get to diagnose. It gets to notice. A clinician decides. And between noticing and deciding sits the unglamorous machinery I have come to think is the actual frontier: prospective validation, subgroup breakdowns, and reporting standards that exist and have names. My own small corner of this is a controlled experiment on retrieval-augmented mammography report generation, accepted this month to the Pacific Symposium on Biocomputing for the proceedings and an oral presentation, with my advisor Dr. Wenjing Yang. I will publish it when the camera-ready is in. What that work taught me is that the hard parts are almost never the model. They are the questions about what a comparison entitles you to claim.

The thing I want you to keep

A journalist told a small story on a podcast this month that has stayed with me more than the theorem did. He had been covering a puzzle about dice — sets of dice fair enough that players can each roll one to decide who goes first, with no ties and no advantage — which a loose group of enthusiasts and mathematicians had worked on for more than a decade. After filing the story he went back to a harder version of it, opened a chatbot, and got a candidate answer out of it in minutes using nothing but plain English. He is a computer scientist by training and says his own mathematics is dusty. What he contributed was not mathematics. It was encouragement: the model would stall, ask whether it should try looking somewhere else instead, and he would say yes, go on.

I should be careful with that story, because I checked it and it is smaller than it sounds: solutions for that size of set already existed publicly, including on the wiki the puzzle’s own community keeps, and no one has independently verified his. But the part that matters survives checking. A person with no standing in that field reached into a space he could not otherwise have entered, and the only thing he supplied was direction and permission to continue.

That is the future I actually believe in, and it is smaller and better than the one being sold. Not a machine that thinks for us. A machine that goes where we cannot go and comes back with something, while the deciding what it means, and the checking whether it is true, and the caring who it is for, stay exactly where they have always been.

Every correction in this essay — the misquoted line, the dice claim that was bigger than the record, the two announcements welded into one — came from a person opening a source and reading it. That is not a defeat for the technology. It is the other half of it, and it is the half that is ours.

Finding is not a lesser act than creating. Swanson settled that forty years ago in a library, and the machines have only made the point louder. But nothing that is found becomes knowledge until somebody understands it. Buckmaster and Alpöge lost weeks of sleep to that, and it was the most human thing in the whole story.

Don’t forget which half is yours.


Sources

Everything above is checkable, so here is what I read. Where a claim in the news cycle did not survive checking, I left it out.