OpenAI published 722 mathematical manuscripts organized into 372 families of related results, generated by an internal model and posted to a public GitHub repository. The team said it presented the model with about 4,000 open research problems, and each result used an average of roughly three hours of ChatGPT Pro thinking compute, according to the repository's documentation.

Highlighted results include progress on the Quasi-Riemann Hypothesis, an integer multiplication method faster than the longstanding n log n benchmark, and a uniqueness result for an elastic inverse problem that had been open since 1994, according to Latent Space's roundup of reactions to the release.

Not every result is verified. OpenAI's repository says not all manuscripts have accompanying Lean formal proofs, that "some of the unformalized results could have issues," and that the team will issue corrections as problems are found while preserving its release history.

Anthropic researcher Levent Alpoge called it "the most significant moment in mathematical history," according to Latent Space, while OpenAI's Will Depue said he expects "some results should not survive scrutiny." About one in five of the published results are disproofs or counterexamples rather than new proofs, Latent Space reported.

Three hours of compute per result is not a lot, which is exactly why the claim is getting scrutinized this hard. If a chunk of it survives independent verification, it changes what a working mathematician's job looks like. If it does not, the unformalized papers in that repository are the place to look for where it broke.