A public dispute between NYU mathematician Tristan Buckmaster and OpenAI has turned an apparent advance in AI-assisted mathematics into a test of research governance, provenance and credit.
Buckmaster and Anthropic mathematician Levent Alpöge announced results related to the Navier–Stokes existence and smoothness problem, one of the Clay Mathematics Institute’s Millennium Prize problems. Soon afterward, OpenAI published what it described as a full proof of the central problem, saying an unreleased next-generation model produced the work during an effort that began September 1.
The proofs’ ultimate mathematical standing remains a matter for expert verification. But the immediate operational issue is clearer: researchers, AI labs and universities need credible ways to establish who knew what, when, and how a model contributed to a result.

What the parties say
Buckmaster says that, while he and Alpöge were finalizing their work, they learned that information about their progress had reached OpenAI. He says OpenAI subsequently told them it had obtained a full proof after deploying a team and substantial compute. Buckmaster argued that the route taken was specialized enough that it would be unlikely to emerge merely from prompting a model with the problem statement.
He also raised a second concern: he had used OpenAI’s Codex extensively during the project. OpenAI’s policies allow model improvement using product interactions unless a user opts out. Buckmaster said this created a possible path through which his work could have influenced later model behavior, though he explicitly said he was not accusing anyone of accessing his data.
OpenAI said its researchers and agents did not see Buckmaster and Alpöge’s work before its public release and that no specific user data was accessed in solving the problem. It acknowledged that it could not completely rule out the possibility that de-identified data derived from product usage contributed to model improvements, while maintaining that its proofs differ materially from the other researchers’ work. Sébastian Bubeck, who leads mathematical research at OpenAI, called Buckmaster’s claims “false and inflammatory.”
Why this matters beyond mathematics
The dispute exposes a gap likely to recur as frontier models move from drafting and coding assistance into high-value discovery work. In conventional research, notebooks, drafts, correspondence, preprints and conference records can help establish priority. AI workflows complicate that record: prompts may be private, model outputs can be iterative, training-data policies can be broad, and a lab may run vastly more experiments once it identifies a promising line of attack.
For institutions, the risk is not limited to reputational conflict. Unclear provenance can complicate publication credit, patent strategy, research partnerships and employee mobility. It can also deter external experts from using a lab’s tools on commercially or professionally sensitive work if they do not understand whether their interactions may contribute to model training.
The reported compute scale matters too. OpenAI said its week-long effort used 300 billion output tokens; TechCrunch estimated that at $22.5 million using current Astra rates. Whether or not that estimate translates to OpenAI’s actual cost, it illustrates how access to concentrated compute can reshape priority contests once a research direction appears promising.
What to watch next
The first question is independent validation: a claimed solution to a Millennium Prize problem requires intensive scrutiny, not a launch announcement. The second is governance. Expect closer attention to opt-out defaults, data-retention disclosures, audit logs, documented prompt histories and clearer attribution standards for model-assisted discoveries.
AI labs that want outside researchers to use their systems for frontier work will need more than assurances after a dispute arises. They will need policies and technical records that can distinguish model capability, customer-derived learning and human-led insight before a result becomes a contest over credit.



