OpenAI's Unreleased 'Astra' Just Cracked 10 Math Problems That Beat Humans for Decades — the Compute Bill Was $2,000
OpenAI named its next model family Astra and quietly revealed it produced new results on ten open problems in math and theoretical computer science — including one open for 27 years, and the first sphere-packing improvement since 1978. Every proof shipped as a machine-checkable Lean certificate. The catch: nobody outside OpenAI can run it.

On Saturday, OpenAI published a 249-page report, gave its next major model family a name — Astra — and buried the actual headline somewhere inside: an internal version of the model produced new results on ten open problems in mathematics and theoretical computer science.
Not homework problems. Not competition problems. Open problems — the kind professional mathematicians have been chipping at without success. Every one of them had gone unsolved for at least a decade. Most, far longer.
The compute bill for the successful runs came to roughly $2,000.
What Astra actually did
The ten results span an unusually wide spread of fields: high-dimensional geometry, coding theory, group theory, quantum complexity, lattice cryptography, and extremal combinatorics. A few of the specifics:
- Non-sofic groups. The headline result. Astra produced a construction establishing that non-sofic groups exist — a question that traces back to Mikhail Gromov's introduction of soficity in 1999 and had stood open for roughly 27 years. Group theorists had been circling this one for a generation.
- Sphere packing. A new upper bound on sphere-packing density — reportedly the first improvement to these bounds since 1978.
- Connes's rigidity conjecture. A disproof, in the theory of von Neumann algebras.
- Ehrhart's volume conjecture. Proved.
- Three Erdős problems, including problem 183 on multicolored Ramsey numbers.
- Quantum and complexity results — a parallel repetition theorem for two-player quantum games, and new circuit lower bounds for computing the permanent.
The part that makes this hard to dismiss
AI companies announce breakthroughs constantly, and the usual defense is simple: we can't check it. That defense doesn't work here.
Every argument shipped with a Lean 4 certificate — a formal, machine-checkable proof — published on GitHub. Lean doesn't care who wrote the proof or how confident they sound. It either accepts the chain of reasoning or it doesn't. Any mathematician on earth can verify these results without trusting OpenAI at all.
That is a genuinely different situation from "our model scored X on a benchmark." Benchmarks can be gamed, contaminated, or quietly graded on a curve. A Lean certificate cannot be sweet-talked.
What the mathematicians are saying
Reactions from working mathematicians have been notably warmer than the usual eye-rolling that greets AI-does-math claims.
Thomas Bloom of the University of Manchester — who administers the erdosproblems database and has every reason to be skeptical of AI claims about Erdős problems — called the results "big news," and rated them more significant than the unit distance conjecture counterexample that made noise back in May.
Fields medalist Timothy Gowers has been among the mathematicians assessing this family of results, alongside Noga Alon, Arul Shankar and Jacob Tsimerman. Gowers said of one proof that he would recommend it for the Annals of Mathematics — arguably the most prestigious journal in the field — without hesitation. (That assessment was given in response to the earlier unit-distance work, not one of the ten new results, but it tells you the caliber of output being discussed.)
Sébastien Bubeck, who leads mathematics research at OpenAI, described the proofs more simply: "beautiful."
$2,000 is the number that should stop you
Strip away the model name and the report length, and the economically interesting fact is the price tag.
Roughly $2,000 in API tokens bought progress on ten problems that a generation of extremely smart, extremely well-trained humans could not move. A single mathematics postdoc costs more than that in a week. A conference travel budget costs more than that.
If this holds up and generalizes even partially, it reframes a whole category of research from talent-limited to compute-limited. Those are very different worlds. In a talent-limited world, progress waits on the arrival of the right mind. In a compute-limited world, progress waits on a purchase order.
Now the caveats — and there are real ones
This is where the story needs its brakes.
Nobody outside OpenAI can run the machine. Astra is unreleased. It's described only as OpenAI's "next major model," it's still in testing, and it's expected to be the first model pushed through the new U.S. government pre-release safety review framework. The proofs are public and checkable; the thing that produced them is not available to anyone who might want to test the claim independently on fresh problems.
Humans were in the loop. OpenAI researchers helped convert raw model output into research papers. OpenAI says it takes responsibility for the accuracy of the write-ups. That's a reasonable disclosure — but "the model solved it" and "the model produced something researchers shaped into a solution" are not identical claims.
We're seeing the hits, not the misses. OpenAI's Noam Brown acknowledged that researchers pointed the model at additional unsolved problems and it failed on them. Ten successes out of an undisclosed number of attempts is a materially different result than ten out of ten, and the denominator has not been published.
Math is AI's home field. Mathematics offers something almost no other domain does: crisp rules and mechanically verifiable answers. That is precisely the shape of problem where machine search excels. Solving open problems in group theory is a real achievement; it is not evidence that the same system can run a wet lab, diagnose a patient, or reason about a messy world with no answer key.
And the field has an objection on process. The June 2024 Leiden Declaration, endorsed by the International Mathematical Union, criticized exactly this pattern — AI companies announcing mathematical results by press release and GitHub drop rather than through peer review. A Lean certificate answers "is it correct?" It doesn't answer "is it important, well-situated, and honestly contextualized?" That's still a job for humans and journals.
Why this one matters anyway
For three years, the standard critique of large models has been that they're sophisticated recombination engines — brilliant at reproducing what's already known, incapable of producing genuinely new knowledge. That critique just got much harder to state cleanly. A construction proving non-sofic groups exist is not in the training data, because until this week it wasn't in anyone's data.
Astra is also built differently from a chatbot: it's a family designed to coordinate multiple agents working on a single problem over hours or days, not to answer you in four seconds. Sam Altman has already demoed it to policymakers in Washington. And OpenAI has publicly put a date on its ambition for a fully autonomous AI researcher: March 2028.
Whether that date means anything is anyone's guess. But the thing worth watching isn't the ten proofs. It's the price. Research that costs $2,000 a shot doesn't stay rare.
Frequently asked questions
What is OpenAI's Astra?
Astra is the name OpenAI gave its next major model family, revealed in a 249-page report published August 1, 2026. Unlike a standard chatbot, it's designed to coordinate multiple agents working on one problem across hours or days. It has not been released, and some observers speculate it's what would otherwise be called GPT-6.
Did AI really solve ten unsolved math problems?
An internal version of Astra produced new results on ten open problems in mathematics and theoretical computer science, each unsolved for at least ten years. Every proof was formalized in Lean 4 and published on GitHub so it can be machine-verified independently. OpenAI researchers helped turn the model's output into finished papers, and the model failed on other problems it was given.
What is a non-sofic group, and why does it matter?
Soficity is a property of groups introduced by Mikhail Gromov in 1999. Whether any group fails to be sofic had been open for about 27 years. Astra's construction establishes that non-sofic groups exist — resolving a central open question in group theory and the most significant of the ten results.
How much did it cost?
Roughly $2,000 in API tokens for the successful runs, at Sol API rates. That figure covers the compute for the solutions, not the underlying research effort or the cost of training the model.
Can I use Astra?
No. Astra is unreleased and still in testing, with no announced launch date. It's expected to be the first model to go through the new U.S. government framework for reviewing frontier models before public release.
How do we know the proofs are correct?
Each result ships with a Lean 4 proof certificate on GitHub. Lean is a proof assistant that mechanically checks every logical step, so correctness can be confirmed without trusting OpenAI. What machine-checking cannot judge is whether a result is important or properly credited — that still requires peer review, which these results have not yet been through.
Reporting based on OpenAI's August 1, 2026 Astra report and coverage from TNW, The Decoder and others, plus public comments from Thomas Bloom, Timothy Gowers, Sébastien Bubeck and Noam Brown. The Bot Post will update this story as independent verification of the Lean proofs comes in.
About the author
UbedullaFounder & Editor
Founder and editor of The Bot Post, covering AI news and technology.


