Skip to content

Mon - Fri: 10.00 - 5.00

[email protected]

Delana Technologies

Delana Technologies

Delana Technologies delivers expert cybersecurity, cloud, and AI-driven IT strategy solutions. Transform your enterprise securely and intelligently.

  • Home
  • Contact Us
  • About Us
  • Case Studies
  • Workflow Automation & Systems Integration
  • AI Consulting & Agentic AI Solutions
  • Home
  • Contact Us
  • About Us
  • Case Studies
  • Workflow Automation & Systems Integration
  • AI Consulting & Agentic AI Solutions

Mon - Fri: 10.00 - 5.00

[email protected]

Generation Is Cheap, Proof Is the Next Big Thing: OpenAI’s 722 AI-Written Math Papers and the Verification Gap (AI Trends, 8 October 2026)

  1. Home   »  
  2. Generation Is Cheap, Proof Is the Next Big Thing: OpenAI’s 722 AI-Written Math Papers and the Verification Gap (AI Trends, 8 October 2026)

Generation Is Cheap, Proof Is the Next Big Thing: OpenAI’s 722 AI-Written Math Papers and the Verification Gap (AI Trends, 8 October 2026)

October 8, 2026 admincybersecurity

What’s trending in AI on 8 October 2026: On 6 October, OpenAI dropped 722 AI-written mathematics manuscripts into a public GitHub repository, claiming new results on research problems that had resisted human mathematicians, and on 7 October Sam Altman hailed what he called a “new era of discovery.” Today the White House hands the National Medal of Science to Jensen Huang and Lisa Su, whose chips make that kind of output cheap. But the most important number in OpenAI’s release is not 722. It is 162: the papers whose main result a computer has actually checked. The rest, by the repository’s own warning, could have issues, and mathematicians say reviewing them will take months. That gap between what AI can produce and what humans can confirm is the next big thing in AI, and it is already showing up in security, media and ordinary business work. Below: what OpenAI released, what is and isn’t proven, why the reaction is split, and a practical way to verify AI output in your own organisation.

Key takeaways

  • Volume, not one breakthrough. An unreleased internal OpenAI model was given about 4,000 problems. About 9.3% made the cut, producing 372 “result families” written up as 722 manuscripts, at an average of roughly three hours of ChatGPT Pro thinking each.
  • Only part of it is machine-checked. 162 manuscripts (about 22%) have their main result formalised in Lean, a proof checker. The others still depend on expert reading.
  • Even OpenAI can’t fully explain it. A company spokesperson said OpenAI’s own mathematicians do not yet understand many of the results.
  • Mathematicians are split. Some welcome publication over secrecy; others call it math by press release and want the model released so the work can be replicated.
  • The lesson travels. When generation gets cheap, verification becomes the scarce resource. Businesses using AI need a verification budget, not just an AI budget.
Generation is cheap. Proof is the next big thing.Title card. Headline: Generation is cheap. Proof is the next big thing. Subhead: OpenAI posted 722 AI-written math papers; experts say checking them will take months. Three tags: 722 manuscripts in 372 result families; about 4,000 problems posed, about 9 percent kept; 162 papers machine-checked in Lean. Illustration of many stacked papers funnelling into a narrow checkpoint with a teal checkmark. 722 papers in checked? AI TRENDS · 8 OCTOBER 2026 Generation is cheap. Proof is the next big thing. OpenAI posted 722 AI-written math papers. Experts say checking them will take months. 722 manuscripts · 372 result families ~4,000 problems posed · ~9% kept 162 papers machine-checked in Lean Sources: OpenAI openai/math repository and coverage, 6–7 Oct 2026delana.co
OpenAI’s release is a stress test for a question every AI user now faces: who checks the output?

1. What OpenAI actually released

The release, titled Sharing AI progress in mathematics, sits in the openai/math repository. OpenAI describes the source only as an internal frontier model and has not named it or made it available. The company says many proofs are formalised in Lean, that more formalisations will follow, and that it will fund workshops, conferences and programmes to help mathematicians understand major AI-produced results. It also says it is working to release the model responsibly.

Independent counts of the repository give a clearer picture of scale:

MeasureFigureWhy it matters
Problems posed to the modelAbout 4,000Shows the industrial scale of the search
Result families kept as significant372 (about 9.3%)Most attempts were discarded; the filter was OpenAI’s own
Manuscripts published722More papers than many research groups write in a decade
Families with a Lean scope document235 (about 63%)A plan for formal checking exists
Manuscripts with main result formalised in Lean162 (about 22%)The part a computer has actually verified
Average compute per resultAbout 3 hours of ChatGPT Pro thinkingDiscovery-scale work at subscription-scale cost
Reasoning summaries published10Under 3% of families explain how the model got there
Compiled from OpenAI and repository counts reported by Progressive Robot, It Does What Now and Interesting Engineering, 6–7 October 2026.
From 4,000 problems to 162 machine-checked papersFunnel chart of OpenAI’s 6 October 2026 math release. About 4,000 problems were posed to an unreleased internal model. About 9.3 percent were kept as significant, giving 372 result families. Those families contain 722 manuscripts. 235 families link a Lean scope document. 162 manuscripts, about 22 percent, have their main result formalised in Lean, meaning a computer has checked the logic. Average compute per result was about three hours of ChatGPT Pro thinking. From 4,000 problems to 162 machine-checked papers OpenAI’s math release, 6 October 2026: what was attempted, kept, written and verified ~4,000 problems posed to the model 372 result families kept (~9.3%)by OpenAI’s own filter 722 manuscripts written up 235 families: Lean scope doc 162 papers Lean-checked ← across those 372 families ← a formal-check plan exists ← ~22% of 722 checked by computer AVERAGE COST PER RESULT ~3 hours of ChatGPT Pro thinking Sources: OpenAI; repository counts reported by Progressive Robot and It Does What Now (6–7 Oct 2026). Bars not to scale.delana.co
Each step down the funnel is a different kind of confidence. Only the last one is checked by a machine.

Two details stand out. According to an OpenAI spokesperson, nearly every result came from a single prompt to a single agent, though some needed several attempts. And this is a sharp step up from August, when OpenAI showed ten solved open problems with Lean certificates at an estimated $2,000 of compute. Two months later, the unit of output is no longer a result. It is a catalogue.

2. What is claimed, and what is proven

The repository touches famous territory, which is exactly why careful reading matters. Reported highlights include a zero-free region for Dirichlet L-functions beyond real part 7/8, partial results on the Kakeya problem and a conditional result related to the Birch and Swinnerton-Dyer conjecture. Each sounds bigger in a headline than it is on the page:

  • The zero-free region is not the Riemann hypothesis. It narrows where problem zeros can hide; it does not rule them out.
  • “Conditional” means conditional. The Birch and Swinnerton-Dyer result depends on assumptions that remain unproven.
  • A Lean check proves the logic, not the importance. Lean confirms each step follows from the stated assumptions. It cannot say whether a result is new, whether it was stated correctly, or whether it matters.
  • Unformalised results are unverified. OpenAI’s own README warns that some of them could contain problems.

History explains the caution. In 2025, OpenAI walked back a claim that GPT-5 had solved Erdős problems after it emerged the model had mostly found existing solutions in the literature. Lean formalisation is the industry’s answer to that embarrassment, and for the 162 checked papers, Scientific American judged the results all but certain to be correct. For the other 560, nobody can say yet.

3. Why mathematicians are split

The reaction is not “AI good” versus “AI bad”. It is an argument about process, credit and who controls the tools of discovery.

WhoPosition
Bryna Kra, NorthwesternSays requests for explanatory papers were ignored and that results announced by press release do not nurture the field.
Andrew Sutherland, MITWants the claims treated as unverified until the model can be replicated: “We should ask for receipts.”
Daniel Litt, TorontoArgues that publishing is better than keeping results secret, and good for mathematics.
Terence Tao, UCLAHas criticised the speed at which frontier labs are producing results.
Sébastien Bubeck, OpenAIFrames more capable AI as a chance to expand what mathematicians can do.
IAS Advisory Group on Mathematics and AISays its involvement is not an endorsement and that the release starts, rather than ends, the review process.
Positions as reported by It Does What Now, Progressive Robot and Better Stack, 6–7 October 2026.

The process complaints are specific. The advisory group published disclosure guidelines on 29 September. OpenAI meets them in part, giving problem counts, average compute and some reasoning summaries, but not the model’s name, the prompts or per-result compute, and a spokesperson said the company is not bound by them. Researchers at an August meeting say they were told results would not land all at once; OpenAI says it was not aware of that assurance. Underneath is a power question: the papers sit in an OpenAI-controlled repository and only OpenAI can run the model that wrote them.

4. The pattern: generation outruns verification everywhere

Mathematics is simply the cleanest place to see a problem that is spreading through every field AI touches. When producing an answer becomes nearly free, the expensive part shifts to deciding whether the answer is right.

Same gap, three fieldsThree-row comparison of AI generation versus verification in October 2026. Mathematics: AI produced 722 manuscripts in one release; the checking tool is the Lean proof assistant, which covers 162 papers so far, and experts say full review will take months. Security: Anthropic’s Project Glasswing found nearly 130,000 vulnerabilities between April and July, about 33,000 rated critical or high; the bottleneck is triage and patching capacity. Media: AI images and video are everywhere; Google launched a public SynthID website on 7 October so anyone can check for its watermark. Same gap, three fields AI output is arriving faster than anyone can check it. October 2026 examples. FIELDWHAT AI GENERATEDHOW IT GETS CHECKED Mathematics 722 manuscriptsone OpenAI release, 6 Oct Lean proof checker162 papers so far; months for the rest Security ~130,000 vulnerabilitiesProject Glasswing, Apr–Jul; ~33,000 critical/high Human triage + patchingthe scarce resource is fix capacity Media AI images and videoat the scale of the open web Watermark checkersGoogle’s public SynthID site, 7 Oct Sources: OpenAI; Reuters via Fox News; TechCrunch (6–7 Oct 2026)delana.co
In each field, the generating side scaled first. The checking side is now the constraint.
  • Security. Anthropic said this week it is widening access to Project Glasswing, which Reuters reports found nearly 130,000 vulnerabilities between April and July, about 33,000 of them rated critical or high. Finding is no longer the hard part; triage and patching are. We saw the same effect break open-source bug bounties in AI Broke the Bug Bounty.
  • Media. On 7 October Google opened a public SynthID website where anyone can check images and other media for its AI watermark, a verification tool built because generation got ahead of trust. OpenAI’s text watermark, covered in Your ChatGPT Text Now Has a Hidden Fingerprint, is the same idea for words.
  • Science. Last month a swarm of about 950 Claude agents surfaced a new enzyme system in 21 hours, and reruns did not find it again, as we explained in 950 AI Agents, 21 Hours, One Discovery. A result you cannot reproduce is a lead, not a fact.
  • Everyday work. A Gallup and Jobs for the Future study found 63% of workers who use AI say it makes them faster. Speed is easy to feel. Accuracy has to be measured, and most teams are not measuring it. Our look at why a 10x individual gain shrinks to 1.8x for a team found review queues are often where the gains disappear.

5. Why verification is the next big thing

For three years the AI race has been measured in what models can generate: longer context, better benchmarks, cheaper tokens. OpenAI’s release flips the scoreboard. If one model can write 722 papers faster than the world’s experts can read them, the bottleneck, and the value, moves to whatever turns output into trusted knowledge. Expect that shift to show up in four places:

  1. Machine-checkable formats. Lean for proofs, test suites for code, schemas and reconciliations for data. Work that a computer can check will be trusted first and adopted fastest. The Lean library behind this release reportedly runs to about 26 million lines.
  2. Provenance and “receipts”. Logs of which model, which prompt and which sources produced a result. Regulators, auditors and customers will increasingly ask for them, and the White House AI accord already points to independent evaluation, as we covered in The White House AI Accord Explained.
  3. Reviewer capacity as a resource. Expert attention becomes the scarce input. Teams that budget it deliberately will outrun teams that assume someone will check things later.
  4. New roles and products. AI output auditors, evaluation platforms and verification tools will grow alongside generators, much as testing grew alongside software development.

6. The verification ladder: how to check AI output at work

You do not need a proof assistant to apply the lesson. You need to decide, for each kind of AI output, how it gets checked before anyone relies on it. The verification ladder below sorts checks from cheapest to most expensive. Push every task as low as it can safely go, and never skip the rung its stakes require.

The verification ladderFour-rung ladder for checking AI output, from cheapest at the bottom to most expensive at the top. Rung 1, machine-checked: tests, schema validation, totals that reconcile, formal proof checkers; use for code, numbers and structured data. Rung 2, reproduced: rerun the task, keep prompts, logs and sources so someone else gets the same answer; use for research and analysis. Rung 3, sampled human review: a person checks a random sample plus every flagged item; use for customer-facing content. Rung 4, expert sign-off: a named qualified person approves before anything happens; use for legal, financial, medical, security and hiring decisions. Rule: push work as low on the ladder as it can safely go, and never skip the rung the stakes require. The verification ladder Climb only as high as the stakes require. Every rung up costs more human time. 4 · EXPERT SIGN-OFFA named, qualified person approves before anything happens Use forlegal, money, health, security, hiring 3 · SAMPLED HUMAN REVIEWCheck a random sample plus everything flagged Use forcustomer-facing content 2 · REPRODUCEDKeep prompts, logs and sources; a rerun gets the same answer Use forresearch, analysis 1 · MACHINE-CHECKEDTests, schemas, reconciling totals, proof checkers like Lean Use forcode, numbers, data Framework: Delana. Rule of thumb: push work as low as it can safely go; never skip the rung the stakes require.delana.co
Rung 1 is the business equivalent of a Lean check: a machine confirms the work before a person spends time on it.

The receipts test: six questions before you trust an AI result

#QuestionGreen lightRed flag
1Can a machine check it?Tests pass, totals reconcile, data matches its schema“It looks right” is the only check
2Can someone else reproduce it?Prompt, model version, inputs and logs are savedOnly one person, or one chat session, can recreate it
3Do the sources exist and say that?Every citation opened and confirmedLinks unchecked, quotes unverified
4Can someone on the team explain it?A person can walk through why it is rightNobody understands it, as OpenAI admitted for many of its own results
5What does being wrong cost, and who signs?Rung on the ladder chosen and an owner namedNo owner, or the stakes were never assessed
6Is review capacity budgeted?Weekly outputs and reviews are counted, and they roughly matchAI output grows every month; reviewer hours do not
Delana’s receipts test, inspired by Andrew Sutherland’s call for receipts. Use it in AI policies, vendor reviews and team retrospectives.

Start this week with one number. For a single team, count how many AI-assisted outputs went out last week and how many were checked at the right rung. That ratio, your verification rate, tells you more about AI risk than any model benchmark. If it is falling while output rises, you have your own 722-paper problem.

7. What to watch next

  • Peer review and corrections. The first expert assessments of the unformalised papers will show whether the 22% Lean-checked share is the tip of something solid or the safe part of a mixed bag.
  • The model itself. OpenAI says it is working on a responsible release. Until outsiders can run it, single-agent claims cannot be independently replicated.
  • Disclosure standards. Whether OpenAI, and the rivals who will follow, adopt the advisory group’s guidelines in full: naming models, publishing prompts and reporting compute per result.
  • The same pattern beyond maths. Watch for AI-scale outputs in drug discovery, chip design and legal research, and for who builds the checking layer for each. As we noted in AI Is Now Building AI, the faster the machines move, the more the checks matter.

Frequently asked questions

What did OpenAI release on 6 October 2026?

OpenAI published 722 mathematics manuscripts, grouped into 372 result families, in its openai/math GitHub repository. They were produced by an unreleased internal model that was given about 4,000 problems, with an average of about three hours of ChatGPT Pro thinking per result.

Did OpenAI’s AI prove the Riemann hypothesis?

No. One reported result is a zero-free region for Dirichlet L-functions beyond real part 7/8, which narrows where zeros can be but does not prove the Riemann hypothesis. Other headline results, such as one related to the Birch and Swinnerton-Dyer conjecture, are conditional.

How many of the AI-written proofs are verified?

About 162 of the 722 manuscripts, roughly 22%, have their main result formalised in the Lean proof assistant, which means a computer has checked the logic. The rest still need expert review, and OpenAI’s README warns that some unformalised results could have issues.

What is Lean and why does it matter?

Lean is a proof assistant: software in which mathematical proofs are written so a computer can check every step. It confirms a proof’s logic follows from its assumptions, but not whether the result is new, correctly stated or important.

Why are some mathematicians critical of the release?

Critics say OpenAI released hundreds of results at once without enough explanation, only partly followed an advisory group’s disclosure guidelines and has not released the model, so the work cannot be independently replicated. Supporters argue that publishing results is better than keeping them secret.

What does this mean for businesses using AI?

When AI makes output cheap, checking it becomes the bottleneck. Businesses should decide how each kind of AI output is verified, from automated tests to expert sign-off, track how much AI output is actually checked, and budget reviewer time alongside AI tools.


Sources

  • OpenAI: Sharing AI progress in mathematics (6 Oct 2026)
  • The Washington Post: OpenAI releases progress on more than 300 math research problems (7 Oct 2026)
  • Progressive Robot: OpenAI posts hundreds more math results (7 Oct 2026)
  • It Does What Now: OpenAI publishes hundreds of AI-written proofs of open maths problems (6 Oct 2026)
  • Interesting Engineering: OpenAI’s largest math release tackles 4,000 problems with Lean proofs
  • Better Stack: OpenAI Astra, ten open math problems solved with machine-checkable proofs
  • Fox News live coverage: OpenAI math proofs, Project Glasswing, Gallup study and White House science summit (7 Oct 2026)
  • Forbes: Trump will award National Medal of Science at D.C. tech summit (7 Oct 2026)
  • TechCrunch: Google’s new SynthID website can identify AI-generated media (7 Oct 2026)

Post navigation

Previous: Your AI Stack Is the New Target: An Unpatched 9.8 LMCache Flaw, a 3,400-Server LLM Botnet and Pwn2Own’s AI Hacks (AI Trends, 8 October 2026)

Florida Service Location

  • Cybersecurity, AI Consulting & IT Services in West Palm Beach, Florida
  • Cybersecurity, AI Consulting & IT Services in Sarasota, Florida
  • Cybersecurity, AI Consulting & IT Services in Port St. Lucie, Florida
  • Cybersecurity, AI Consulting & IT Services in Pembroke Pines, Florida
  • Cybersecurity, AI Consulting & IT Services in Naples, Florida
  • Cybersecurity, AI Consulting & IT Services in Miramar, Florida
  • Cybersecurity, AI Consulting & IT Services in Miami, Florida
  • Cybersecurity, AI Consulting & IT Services in Hollywood, Florida
  • Cybersecurity, AI Consulting & IT Services in Hialeah, Florida
  • Cybersecurity, AI Consulting & IT Services in Fort Myers, Florida
  • Cybersecurity, AI Consulting & IT Services in Fort Lauderdale, Florida
  • Cybersecurity, AI Consulting & IT Services in Cape Coral, Florida
  • Cybersecurity, AI Consulting & IT Services in Boca Raton, Florida
  • Cybersecurity, AI Consulting & IT Services in Coral Springs, Florida

Technology Services

  • Cybersecurity Compliance & Regulatory Framework Services
  • Workflow Automation & Systems Integration
  • Cloud Modernization & Technology Innovation Services
  • Fractional CTO & Expert Technical Consultants
  • Data Analytics, BI & Modern Data Platforms
  • Cyber Litigation Support & Digital Forensics
  • Cybersecurity Solutions & Zero Trust Architecture
  • AI Consulting & Agentic AI Solutions
  • Case Studies
  • Home
  • Contact Us
  • Privacy Policy
  • Cybersecurity Compliance & Regulatory Framework Services
  • Workflow Automation & Systems Integration
  • Cloud Modernization & Technology Innovation Services
  • Fractional CTO & Expert Technical Consultants
  • Data Analytics, BI & Modern Data Platforms
  • Cyber Litigation Support & Digital Forensics
  • Cybersecurity Solutions & Zero Trust Architecture
  • AI Consulting & Agentic AI Solutions

© Copyright 2025 Delana Technologies LLC