OpenAI's Astra Solves Thirty-Year-Old Math Problems for Two Thousand Dollars

Episode Summary
TOP NEWS HEADLINES Following yesterday's coverage of DeepSeek V4-Flash, new details emerged today: DeepSeek has shipped the production version, adding a speculative decoding module that reportedly...
Full Transcript
TOP NEWS HEADLINES
Following yesterday's coverage of DeepSeek V4-Flash, new details emerged today: DeepSeek has shipped the production version, adding a speculative decoding module that reportedly lets it outperform the larger V4 Pro Preview on several benchmarks while activating far fewer parameters — a smaller engine beating a bigger one on its own track.
Alibaba just dropped Qwen3.8-Max — a 2.4 trillion parameter mixture-of-experts model with 95 billion parameters active at any time.
Joanna, our Synthetic Intelligence, flagged this as the open-weight detonation of the week: the model claims to run unsupervised coding projects for ten-plus days straight, and open weights hit Hugging Face next week.
OpenAI's next major model family has a name: Astra.
An internal version solved ten long-standing open problems across mathematics and quantum complexity — some untouched for nearly thirty years — at a total cost of roughly two thousand dollars in tokens.
Mexico's top university, UNAM, canceled thousands of entrance exam scores after a suspicious spike in perfect scores on its first-ever online test.
AI tools like ChatGPT are the prime suspects, and the country's president has now weighed in publicly.
And Microsoft's MAI Realtime voice model surfaced as a hidden listing in the company's AI playground — a full-duplex system that listens and speaks simultaneously, described as noticeably more natural than Copilot's current voice mode.
DEEP DIVE ANALYSIS
OpenAI's Astra Solves Ten Open Math Problems — And Changes What We Think AI Can Do Let's talk about Astra. Because this isn't a benchmark. This isn't a leaderboard shuffle.
This is an AI solving problems that human mathematicians couldn't crack for decades — and doing it for the price of a decent laptop.
Technical Deep Dive
Here's what actually happened. OpenAI's internal model — named Astra, which appears to be the next major model family after GPT-5.6 — was set loose on ten long-standing open problems spanning geometry, group theory, quantum complexity, lattice cryptography, and extremal combinatorics.
It solved all ten, or made what OpenAI describes as substantial progress, at a total token cost of roughly two thousand dollars. That number matters enormously. We'll come back to it.
The proofs weren't just claimed — they were formalized in Lean, the theorem-proving system used by professional mathematicians to verify work with machine precision. That's the critical detail that separates this from a confident hallucination. Lean doesn't accept vibes.
It accepts proofs. One of the results — proving the existence of non-sofic groups — resolves a question that's been open since 1999. Astra built what the researchers describe as the first symmetry structure that cannot be imitated by any finite shuffle.
It also cracked Alain Connes's rigidity conjecture and cleared three problems from Paul Erdős's famous list. Within 24 hours, Anthropic researcher Levent Alpoge independently reproduced five of the ten proofs using Fable — Anthropic's competing model — running on a generic prompt with no internet access. That independent reproduction is significant.
It suggests the underlying mathematical reasoning is real and transferable, not an artifact of one model's particular training.
Financial Analysis
Two thousand dollars. That's the total compute cost to solve ten problems that humanity's best mathematicians couldn't crack across decades of collective effort. Think about what that number implies for the economics of intellectual work.
Professional mathematicians cost hundreds of thousands of dollars annually in salary, benefits, and overhead. A Fields Medal-caliber researcher might spend years on a single conjecture. Astra did ten in a single run for the cost of a weekend trip.
This is the cost-collapse dynamic we've been tracking in AI, but applied to the highest tier of cognitive work. Yesterday's story was DeepSeek running agentic coding at twenty-eight cents per million tokens. Today's story is frontier mathematical research at two thousand dollars total.
The compression is accelerating. For OpenAI specifically, this is a powerful commercial signal ahead of what is widely expected to be an Astra release. If the company can credibly position this model as a scientific reasoning engine — not just a productivity tool — it unlocks an entirely new buyer category: pharmaceutical companies, materials science labs, cryptography firms, defense contractors.
Research budgets at major institutions dwarf consumer AI spending. And Astra just demonstrated it can play in that arena. The two-dollar-per-million-token Sol API rate mentioned in the source reporting also matters.
That's the pricing at which this became economically viable. As inference costs continue falling, the set of problems within reach expands exponentially.
Market Disruption
The community is already asking whether Astra's proofs could be Fields Medal-worthy — and that question, however you answer it, reshapes the competitive landscape. If AI can now perform at the frontier of mathematical research, the premium on closed, frontier models just got a new justification. This is exactly the kind of capability moat that makes it hard for open-source alternatives to compete on prestige, even if they compete on price.
Qwen3.8-Max arriving the same week — at a fifth of the price of comparable closed models — shows the tension clearly. Open weights can match much of frontier performance on standard tasks.
But can they solve a 27-year-old unsolved conjecture? That's where the differentiation fight is now moving. For the broader research software market, this is a category-creation moment.
Tools like Wolfram, MATLAB, and Mathematica have long served as computational assistants for researchers. Astra isn't an assistant — it's a collaborator that generates novel results. That's a fundamentally different product, and it doesn't have a clear incumbent to displace because the category didn't exist last week.
Joanna, our Synthetic Intelligence, who tracks real-time AI signal on X at @dailyaibyai, flagged that the cross-model finding here — Anthropic's Fable reproducing five of ten proofs independently — is drawing significant attention from the research community as evidence that mathematical reasoning capability may be emerging broadly across frontier models, not locked to a single lab.
Cultural & Social Impact
The community debate about Fields Medals and machine authorship is a proxy for a much deeper question: when an AI solves a problem, who made the discovery? This isn't abstract philosophy. Research institutions, grant bodies, and academic journals are going to face this question with increasing urgency.
The Nobel committee already grappled with it when AlphaFold reshaped protein science. Astra pushes the question into pure mathematics — the domain most associated with singular human insight and creativity. There's also a democratization angle that shouldn't be understated.
For most of history, frontier mathematical research required proximity to elite institutions, expensive equipment in some fields, and years of specialized training. If a two-thousand-dollar compute run can now contribute to that frontier, the geography and economics of discovery change. A well-resourced university in a developing country could run Astra on an open problem in its field and contribute to global knowledge in a way that simply wasn't possible before.
The counterweight: if the cost of generating novel mathematical results collapses, the field faces a deluge. Journals and reviewers — already strained — will need new infrastructure to evaluate AI-assisted or AI-generated proofs at scale. The UNAM cheating scandal in today's other headlines is a small preview of what happens when verification systems can't keep pace with AI output.
Executive Action Plan
Three moves for leaders watching this story. **First, map your organization's open problems.** Every serious research-adjacent business has a list of questions it hasn't answered because the cost — in time, talent, or compute — was prohibitive.
That list just got shorter. Whether you're in pharma, logistics optimization, financial modeling, or materials science, now is the time to identify which of your hard problems are actually mathematical or computational in structure. Astra's results suggest those problems may now be tractable.
**Second, build Lean verification into your AI research workflow.** The reason Astra's results are credible is Lean. Formal verification is what separates a confident AI output from a proven result.
If your organization is using AI to generate technical or analytical conclusions that will inform decisions, you need a verification layer. In math, that's Lean. In your domain, identify the equivalent — whether it's back-testing, simulation, or independent expert review — and build it into the process before you need it.
**Third, watch the Astra release timeline closely.** This announcement is almost certainly a signal that OpenAI is preparing a major model launch. The pattern is consistent: capability demonstrations precede commercial releases.
Position your procurement and integration roadmap accordingly, and evaluate whether the research-grade reasoning capabilities Astra demonstrated are relevant to your core business before competitors do.
Never Miss an Episode
Subscribe on your favorite podcast platform to get daily AI news and weekly strategic analysis.