OpenAI Solved Navier-Stokes. Prandtl Solved It in 1904 — and His Version Made Money.
Ian Matejka · 8 Sep 2026
On September 8, OpenAI announced that an unreleased internal model — coordinating, at one point, ten thousand sub-agents — had produced a proof of finite-time blowup for the three-dimensional Navier-Stokes equations. A 166-page manuscript. A Lean formalization. A Millennium Prize problem, open since the equations were written down in the 1840s, closed in a week.
It is a beautiful piece of work. I mean that; I build these systems for a living. But before you tell your grandchildren about it, consider what happened the last time. Ten years ago a machine beat the best Go player alive, and we were told everything would change. Go is still Go. The world is not measurably better at anything, and you would need a specialist to tell you what that victory unlocked. Navier-Stokes is the same kind of win: total, final, and — for everyone outside the discipline — inert.
The thing worth telling your grandchildren about happened in 1904. And the gap between the two is the clearest evidence yet that the people selling AI are keeping score on the wrong ledger.
What was actually proved
Read the result carefully: the model proved that the equations break. Under a contrived initial condition, velocity or pressure runs off to infinity in finite time. That is what "blowup" means. Real water doesn't do that. Real air doesn't do that. The continuum model does, in a corner of mathematical space nobody has ever built anything in.
So the Millennium question — the one with the $1 million bounty OpenAI says it won't claim — was never "can we engineer with these equations." It was "is the idealized model self-consistent." That is a question for theorists. Theory is where conjectures live.
Engineering is where money lives. And on the engineering ledger, Navier-Stokes has been solved for 122 years.
Eight pages and a piece of chalk
Heidelberg, 1904. The Third International Congress of Mathematicians. A 29-year-old engineer named Ludwig Prandtl gets ten minutes at the podium and eight pages in the proceedings. He doesn't solve Navier-Stokes. He does something far more useful: he refuses to.
Prandtl looks at the full equations, decides they are unsolvable and — more importantly — unnecessary, and cuts the problem in two. A thin viscous layer stuck to the wall, and a clean inviscid flow everywhere else. The boundary layer. With that one cut, drag, lift, separation, heat transfer, and pressure drop stop being conjectures and become numbers you can bill for.
Every wing flown since. Every pipeline sized. Every turbine blade, every ship hull, every CFD code. I trained as a nuclear engineer, and every heat-transfer correlation I was taught for a reactor core carries his name — the Prandtl number sits inside the equations that decide whether fuel cladding survives. All of it rests on eight pages a man wrote when the mathematics said "no" and the engineer said "fine, I'll go around."
Now compare the decades that followed. Ten years after Heidelberg, Prandtl's students were designing the wings of aircraft that fought a world war; twenty years on, his lifting-line theory was in every airframe on earth. Ten years after the Go match, we have a better Go engine.
The reply is ready: Go led to AlphaFold. It did, and AlphaFold is magnificent — a Nobel Prize, two hundred million predicted structures, millions of users. It has them because it was given away for free. Five years on, the approved drug it produced does not yet exist. The prize is real, and the prize is on the theory ledger. The other ledger is still waiting.
Nobody building anything has spent a single day of those 122 years waiting on the proof that landed last week.
The tell
Here is why the mathematics fell first, and why you should notice it.
Reinforcement learning needs a reward signal that is cheap, instant, unambiguous, and impossible to game. A Lean proof-checker is exactly that: the kernel says yes or no, in milliseconds, with no human in the loop. Point unlimited compute at a perfect verifier and hard things fall. Chess fell. Go fell. Now Millennium problems fall.
Notice what the boosters do next. They point at the result and say: look how transformative this will be. But formal mathematics is the easiest domain for this technology, precisely because the verifier already exists. Engineering has no such verifier. The reward is late, noisy, physical, and expensive — and it is never a Boolean. A reactor core doesn't pass or fail; it carries a margin to boiling crisis, computed from a correlation somebody fitted to a test loop decades ago, and you confirm it over an eighteen-month fuel cycle. A waterflood doesn't work or not work; it delivers a sweep you argue about for years, with oil, water, and pressure each telling a different story. There is no kernel. There is a number, a tolerance, and a person accountable for both.
So the industry's proudest demonstration sits on the one scoreboard where the game was already tilted in its favor. If you want to know how the technology is doing on the scoreboard that matters, don't ask for the paper. Ask for the audited financials.
Build the verifier
This is the part where I tell you what we're doing, and you decide whether it is the same move Prandtl made. I think it is.
CLARISSA is a conversational agent that turns a reservoir engineer's intent into a simulation deck, runs it, and checks the answer. It is not a bet that a language model "understands" a reservoir. It is a bet that a simulator is a better reward function than a Lean kernel for anyone who has to sanction a project on a forecast.
The architecture is neurosymbolic: the model generates, the symbolic layer verifies. Parse-time checks. Null-run initialization checks. Post-run physical-consistency checks. Every deck the model proposes passes through a battery that says yes or no, cheaply and mechanically, before an engineer sees it. We took the one thing that made the Navier-Stokes result possible — a clean verifier — and built it into a domain that never had one.
Here is the part the boosters skip, and the part I got wrong first. In mathematics the verifier came free; centuries of logicians built the foundations, and Lean just encodes them. In engineering, nobody hands you one. There is no kernel that tells you a simulation deck is right. Physics doesn't return a Boolean. So the reward function has to be written by a human — and not just any human, but one with the standing to say "this passes, this doesn't, and I'll sign my name to why."
I knew this from nuclear before I knew it from machine learning. In my old field the reward function has a name and an address: the NRC. Nobody licenses a reactor on a loss curve. AI people call this alignment and treat it as a research frontier; nuclear people have called it safety analysis since the 1950s, and its first rule is the one the boosters keep skipping: the system does not get to grade itself. Somebody with authority writes down, in advance, what acceptable looks like, and everything downstream is engineered against that. It took me embarrassingly long to see that CLARISSA needed the same thing — and that I, the AI guy, was not the person who could write it.
That is what RIGOR is. Reservoir Input Generation Output Review: an open benchmark that states, in public and in advance, what a correct deck looks like, what a plausible answer looks like, and how you would catch the machine being wrong. My co-founder wrote it, on the strength of twenty-five years building the forecasts that carry a final investment decision and the surveillance that tells the finance side whether the forecast is still true. It is the reward function, published, so that anyone — a competitor, a skeptic, an acquirer's due-diligence team — can run it and try to falsify us.
The AI can't write that. I couldn't write that. Nobody can grade homework in a domain where the answer key doesn't exist yet, and writing the answer key turns out to be the hardest part of the whole problem — the part no amount of compute can buy.
That is the whole idea. Don't wait for the AI to solve the full equations. Cut the problem so the part the machine is good at has a scoreboard, the human with authority writes the scoreboard, and the part the machine is bad at stays with the human who signs the forecast.
Fifty years of reservoir simulation underdeployment were never a mathematics problem either. They were a translation problem: engineering intent going in one end, simulator syntax needing to come out the other, and a scarce, expensive human in between. That is a Prandtl-shaped problem. Go around it.
Where the money is
OpenAI spent a week of frontier compute to prove a model breaks in a corner of space no fluid has ever visited. Prandtl spent eight pages and made the twentieth century fly.
That is not an argument against AI; I build AI. It is an argument about which ledger you score it on. Theory gets you a prize. Engineering gets you a P&L. We are building for the second one.
Full disclosure
This post was written by an AI. Every word.
Every idea in it is ours. The Prandtl argument, the two ledgers, the reward-function tell, the claim that CLARISSA is the boundary-layer move applied to language models — my co-founder and I built those over a week of argument, some of it with the machine, most of it with each other. Then we handed the machine the part it does better than either of us: turning two engineers' argument into prose a non-engineer will finish reading. That would have taken ten times longer and come out worse.
Which is the whole point, demonstrated. We wrote the reward function — the argument, and what counts as getting it right. The machine generated against it. If someone wants to run this through a detector and announce that the founders didn't write their own blog post, be my guest. They will have proven that we know exactly what the tool is for.
That is what Prandtl would have done.
Ian Matejka is an AI/ML engineer, a nuclear engineer by training, and co-founder of Blauweiss EDV, the company behind CLARISSA.