What’s trending in AI on 8 October 2026: On 6 October, OpenAI dropped 722 AI-written mathematics manuscripts into a public GitHub repository, claiming new results on research problems that had resisted human mathematicians, and on 7 October Sam Altman hailed what he called a “new era of discovery.” Today the White House hands the National Medal of Science to Jensen Huang and Lisa Su, whose chips make that kind of output cheap. But the most important number in OpenAI’s release is not 722. It is 162: the papers whose main result a computer has actually checked. The rest, by the repository’s own warning, could have issues, and mathematicians say reviewing them will take months. That gap between what AI can produce and what humans can confirm is the next big thing in AI, and it is already showing up in security, media and ordinary business work. Below: what OpenAI released, what is and isn’t proven, why the reaction is split, and a practical way to verify AI output in your own organisation.
Key takeaways
- Volume, not one breakthrough. An unreleased internal OpenAI model was given about 4,000 problems. About 9.3% made the cut, producing 372 “result families” written up as 722 manuscripts, at an average of roughly three hours of ChatGPT Pro thinking each.
- Only part of it is machine-checked. 162 manuscripts (about 22%) have their main result formalised in Lean, a proof checker. The others still depend on expert reading.
- Even OpenAI can’t fully explain it. A company spokesperson said OpenAI’s own mathematicians do not yet understand many of the results.
- Mathematicians are split. Some welcome publication over secrecy; others call it math by press release and want the model released so the work can be replicated.
- The lesson travels. When generation gets cheap, verification becomes the scarce resource. Businesses using AI need a verification budget, not just an AI budget.
1. What OpenAI actually released
The release, titled Sharing AI progress in mathematics, sits in the openai/math repository. OpenAI describes the source only as an internal frontier model and has not named it or made it available. The company says many proofs are formalised in Lean, that more formalisations will follow, and that it will fund workshops, conferences and programmes to help mathematicians understand major AI-produced results. It also says it is working to release the model responsibly.
Independent counts of the repository give a clearer picture of scale:
| Measure | Figure | Why it matters |
|---|---|---|
| Problems posed to the model | About 4,000 | Shows the industrial scale of the search |
| Result families kept as significant | 372 (about 9.3%) | Most attempts were discarded; the filter was OpenAI’s own |
| Manuscripts published | 722 | More papers than many research groups write in a decade |
| Families with a Lean scope document | 235 (about 63%) | A plan for formal checking exists |
| Manuscripts with main result formalised in Lean | 162 (about 22%) | The part a computer has actually verified |
| Average compute per result | About 3 hours of ChatGPT Pro thinking | Discovery-scale work at subscription-scale cost |
| Reasoning summaries published | 10 | Under 3% of families explain how the model got there |
Two details stand out. According to an OpenAI spokesperson, nearly every result came from a single prompt to a single agent, though some needed several attempts. And this is a sharp step up from August, when OpenAI showed ten solved open problems with Lean certificates at an estimated $2,000 of compute. Two months later, the unit of output is no longer a result. It is a catalogue.
2. What is claimed, and what is proven
The repository touches famous territory, which is exactly why careful reading matters. Reported highlights include a zero-free region for Dirichlet L-functions beyond real part 7/8, partial results on the Kakeya problem and a conditional result related to the Birch and Swinnerton-Dyer conjecture. Each sounds bigger in a headline than it is on the page:
- The zero-free region is not the Riemann hypothesis. It narrows where problem zeros can hide; it does not rule them out.
- “Conditional” means conditional. The Birch and Swinnerton-Dyer result depends on assumptions that remain unproven.
- A Lean check proves the logic, not the importance. Lean confirms each step follows from the stated assumptions. It cannot say whether a result is new, whether it was stated correctly, or whether it matters.
- Unformalised results are unverified. OpenAI’s own README warns that some of them could contain problems.
History explains the caution. In 2025, OpenAI walked back a claim that GPT-5 had solved Erdős problems after it emerged the model had mostly found existing solutions in the literature. Lean formalisation is the industry’s answer to that embarrassment, and for the 162 checked papers, Scientific American judged the results all but certain to be correct. For the other 560, nobody can say yet.
3. Why mathematicians are split
The reaction is not “AI good” versus “AI bad”. It is an argument about process, credit and who controls the tools of discovery.
| Who | Position |
|---|---|
| Bryna Kra, Northwestern | Says requests for explanatory papers were ignored and that results announced by press release do not nurture the field. |
| Andrew Sutherland, MIT | Wants the claims treated as unverified until the model can be replicated: “We should ask for receipts.” |
| Daniel Litt, Toronto | Argues that publishing is better than keeping results secret, and good for mathematics. |
| Terence Tao, UCLA | Has criticised the speed at which frontier labs are producing results. |
| Sébastien Bubeck, OpenAI | Frames more capable AI as a chance to expand what mathematicians can do. |
| IAS Advisory Group on Mathematics and AI | Says its involvement is not an endorsement and that the release starts, rather than ends, the review process. |
The process complaints are specific. The advisory group published disclosure guidelines on 29 September. OpenAI meets them in part, giving problem counts, average compute and some reasoning summaries, but not the model’s name, the prompts or per-result compute, and a spokesperson said the company is not bound by them. Researchers at an August meeting say they were told results would not land all at once; OpenAI says it was not aware of that assurance. Underneath is a power question: the papers sit in an OpenAI-controlled repository and only OpenAI can run the model that wrote them.
4. The pattern: generation outruns verification everywhere
Mathematics is simply the cleanest place to see a problem that is spreading through every field AI touches. When producing an answer becomes nearly free, the expensive part shifts to deciding whether the answer is right.
- Security. Anthropic said this week it is widening access to Project Glasswing, which Reuters reports found nearly 130,000 vulnerabilities between April and July, about 33,000 of them rated critical or high. Finding is no longer the hard part; triage and patching are. We saw the same effect break open-source bug bounties in AI Broke the Bug Bounty.
- Media. On 7 October Google opened a public SynthID website where anyone can check images and other media for its AI watermark, a verification tool built because generation got ahead of trust. OpenAI’s text watermark, covered in Your ChatGPT Text Now Has a Hidden Fingerprint, is the same idea for words.
- Science. Last month a swarm of about 950 Claude agents surfaced a new enzyme system in 21 hours, and reruns did not find it again, as we explained in 950 AI Agents, 21 Hours, One Discovery. A result you cannot reproduce is a lead, not a fact.
- Everyday work. A Gallup and Jobs for the Future study found 63% of workers who use AI say it makes them faster. Speed is easy to feel. Accuracy has to be measured, and most teams are not measuring it. Our look at why a 10x individual gain shrinks to 1.8x for a team found review queues are often where the gains disappear.
5. Why verification is the next big thing
For three years the AI race has been measured in what models can generate: longer context, better benchmarks, cheaper tokens. OpenAI’s release flips the scoreboard. If one model can write 722 papers faster than the world’s experts can read them, the bottleneck, and the value, moves to whatever turns output into trusted knowledge. Expect that shift to show up in four places:
- Machine-checkable formats. Lean for proofs, test suites for code, schemas and reconciliations for data. Work that a computer can check will be trusted first and adopted fastest. The Lean library behind this release reportedly runs to about 26 million lines.
- Provenance and “receipts”. Logs of which model, which prompt and which sources produced a result. Regulators, auditors and customers will increasingly ask for them, and the White House AI accord already points to independent evaluation, as we covered in The White House AI Accord Explained.
- Reviewer capacity as a resource. Expert attention becomes the scarce input. Teams that budget it deliberately will outrun teams that assume someone will check things later.
- New roles and products. AI output auditors, evaluation platforms and verification tools will grow alongside generators, much as testing grew alongside software development.
6. The verification ladder: how to check AI output at work
You do not need a proof assistant to apply the lesson. You need to decide, for each kind of AI output, how it gets checked before anyone relies on it. The verification ladder below sorts checks from cheapest to most expensive. Push every task as low as it can safely go, and never skip the rung its stakes require.
The receipts test: six questions before you trust an AI result
| # | Question | Green light | Red flag |
|---|---|---|---|
| 1 | Can a machine check it? | Tests pass, totals reconcile, data matches its schema | “It looks right” is the only check |
| 2 | Can someone else reproduce it? | Prompt, model version, inputs and logs are saved | Only one person, or one chat session, can recreate it |
| 3 | Do the sources exist and say that? | Every citation opened and confirmed | Links unchecked, quotes unverified |
| 4 | Can someone on the team explain it? | A person can walk through why it is right | Nobody understands it, as OpenAI admitted for many of its own results |
| 5 | What does being wrong cost, and who signs? | Rung on the ladder chosen and an owner named | No owner, or the stakes were never assessed |
| 6 | Is review capacity budgeted? | Weekly outputs and reviews are counted, and they roughly match | AI output grows every month; reviewer hours do not |
Start this week with one number. For a single team, count how many AI-assisted outputs went out last week and how many were checked at the right rung. That ratio, your verification rate, tells you more about AI risk than any model benchmark. If it is falling while output rises, you have your own 722-paper problem.
7. What to watch next
- Peer review and corrections. The first expert assessments of the unformalised papers will show whether the 22% Lean-checked share is the tip of something solid or the safe part of a mixed bag.
- The model itself. OpenAI says it is working on a responsible release. Until outsiders can run it, single-agent claims cannot be independently replicated.
- Disclosure standards. Whether OpenAI, and the rivals who will follow, adopt the advisory group’s guidelines in full: naming models, publishing prompts and reporting compute per result.
- The same pattern beyond maths. Watch for AI-scale outputs in drug discovery, chip design and legal research, and for who builds the checking layer for each. As we noted in AI Is Now Building AI, the faster the machines move, the more the checks matter.
Frequently asked questions
What did OpenAI release on 6 October 2026?
OpenAI published 722 mathematics manuscripts, grouped into 372 result families, in its openai/math GitHub repository. They were produced by an unreleased internal model that was given about 4,000 problems, with an average of about three hours of ChatGPT Pro thinking per result.
Did OpenAI’s AI prove the Riemann hypothesis?
No. One reported result is a zero-free region for Dirichlet L-functions beyond real part 7/8, which narrows where zeros can be but does not prove the Riemann hypothesis. Other headline results, such as one related to the Birch and Swinnerton-Dyer conjecture, are conditional.
How many of the AI-written proofs are verified?
About 162 of the 722 manuscripts, roughly 22%, have their main result formalised in the Lean proof assistant, which means a computer has checked the logic. The rest still need expert review, and OpenAI’s README warns that some unformalised results could have issues.
What is Lean and why does it matter?
Lean is a proof assistant: software in which mathematical proofs are written so a computer can check every step. It confirms a proof’s logic follows from its assumptions, but not whether the result is new, correctly stated or important.
Why are some mathematicians critical of the release?
Critics say OpenAI released hundreds of results at once without enough explanation, only partly followed an advisory group’s disclosure guidelines and has not released the model, so the work cannot be independently replicated. Supporters argue that publishing results is better than keeping them secret.
What does this mean for businesses using AI?
When AI makes output cheap, checking it becomes the bottleneck. Businesses should decide how each kind of AI output is verified, from automated tests to expert sign-off, track how much AI output is actually checked, and budget reviewer time alongside AI tools.
Sources
- OpenAI: Sharing AI progress in mathematics (6 Oct 2026)
- The Washington Post: OpenAI releases progress on more than 300 math research problems (7 Oct 2026)
- Progressive Robot: OpenAI posts hundreds more math results (7 Oct 2026)
- It Does What Now: OpenAI publishes hundreds of AI-written proofs of open maths problems (6 Oct 2026)
- Interesting Engineering: OpenAI’s largest math release tackles 4,000 problems with Lean proofs
- Better Stack: OpenAI Astra, ten open math problems solved with machine-checkable proofs
- Fox News live coverage: OpenAI math proofs, Project Glasswing, Gallup study and White House science summit (7 Oct 2026)
- Forbes: Trump will award National Medal of Science at D.C. tech summit (7 Oct 2026)
- TechCrunch: Google’s new SynthID website can identify AI-generated media (7 Oct 2026)
