š¤ Are mathematicians officially obsolete? ā ļø UPDATE: Oct. 9
New Scientist, Oct. 7: OpenAI announces 722 mathematical discoveries in one go:
OpenAI has dropped 722 mathematical papers proving and disproving a range of problems ā a scale that academics say is as impressive as it is baffling.
...
Having achieved impressive depth, OpenAI has seemingly now opted for breadth, releasing hundreds of mathematical discoveries in one fell swoop. The company didnāt name the AI made the discoveries, saying only that it was an āinternal frontier modelā.
Francis Johnson at University College London worked for many years on Wallās D(2) problem ā one of the hundreds of puzzles solved in OpenAIās tranche of papers.
āI worked on this problem for 25 years. I produced two books on it. I, personally, gave up,ā says Johnson. āIām surprised that [AI] has done it quite so quickly, but Iām not surprised that itās done it.ā
Johnson recounts working with AI models in recent months and being āastonished with the sophistication and clarity of the analysis that they gaveā.
āIt puts us all in a very strange position,ā says Johnson. āLetās face it, the genie is out of the bottle now. Weāre going to have to live with it. Human beings are supposed to be adaptable, so weāre going to have to adapt. I think for the moment, we just stand back and be astonished.ā
But Kevin Buzzard at Imperial College London says we should exercise caution. He says the release included 30 papers that were relevant to his field of number theory, but only seven of those seemed impressive and only one was formally verified in Lean ā a type of computer analysis that can prove mathematical results beyond reasonable doubt.
āUnfortunately, acceptance of these results by the community will take time, and journalists are going to have to wait while the mathematicians do their job,ā says Buzzard. āThe six unformalised results will have to wait until either an expert is motivated to read and check the text, or a Lean formalisation is produced.ā
But while it is important not to get ahead of ourselves on the validity of the results, Buzzard says the general upwards trend in AI mathematics is astonishing. He says that, assuming the new results turn out to be mostly true, then we will get some kind of idea as to what the new normal is. āIt has been a long time since there were humans who were experts in all of mathematics, but now we seem to have machines with this property,ā he says.
...
Unusually for mathematical research, the papers were published on GitHub, an online database more commonly used to share computer code.
Oops. New Scientist came with an update on Oct. 8: OpenAI mistranslated mathematics into code for its Navier-Stokes proof.
The short version:
When OpenAI announced its surprise solution to the Navier-Stokes problem, it produced one proof for humans and one for computers ā however, they don't match.
The longer one:
On 8 September, OpenAI announced that it had found a solution to the Navier-Stokes problem, one of the most famous open problems in mathematics. It published the proof in two versions ā one written in ānatural languageā, meaning a combination of English and mathematical symbols, as a human mathematician would write, and another written in the computer code Lean. The Lean proof is meant to be a formalisation of the natural-language version, allowing a computer to mechanically verify that all of its logical statements are true. The problem is, say Hansen and his team, that they donāt match.
āThis formalisation process is trying to replace peer review,ā says team member Fabian Circelli, also at the University of Cambridge. āPeer review would mean that human eyes look at the proofs. But what weāve shown in this paper is that using this type of AI auto-formalisation canāt serve the same purpose.ā
To be clear, the researchers arenāt saying that OpenAI has failed to solve the Navier-Stokes problem. It is entirely possible that both the natural-language proof and the Lean one provide a solution, just as there are hundreds of valid proofs of Pythagorasās theorem. Instead, their point is a more subtle one: that OpenAIās model has āmistranslatedā when converting into Lean.
āWe are not saying that the natural-language proof is wrong,ā says Hansen. āNor do we say that it is correct.ā The issue is that OpenAI presents the two proofs as identical, stating on GitHub that āThis repository contains Lean 4 formalizations of the results presented in [the paper] āFinite time blowup for NavierāStokesāā.
This mistranslation occurs because the AI has to produce a Lean proof that ācompilesā, meaning that the computer code is fully self-consistent and doesnāt produce an error, says Hansen. If, in the process of auto-formalisation, the AI finds a section of the proof that doesnāt compile, it will attempt to find a workaround even if it means diverging from the proof as written in natural language.
The teamās specific claim hinges on part of the proofs called Lemma 8.6. In the natural-language proof, an equation in this part requires that a certain value be below m + 4, where m is a whole number. In the Lean proof, the equivalent value is required to be below m + 5, which is mathematically weaker.
...
OpenAI told New Scientist that it is aware of the mismatch between the natural-language proof and the Lean code, and that this doesnāt mean that either proof is invalid. It says it will rectify any errors in the natural-language proof as they are found, and will also continue the process of formalising the 722 maths papers the firm released this week, only some of which are accompanied by Lean proofs, which themselves havenāt been checked by hand.