Quick Luddite Notes

šŸ¤– Are mathematicians officially obsolete? āš ļø UPDATE: Oct. 9

New Scientist, Oct. 7: OpenAI announces 722 mathematical discoveries in one go:

OpenAI has dropped 722 mathematical papers proving and disproving a range of problems – a scale that academics say is as impressive as it is baffling.

...

Having achieved impressive depth, OpenAI has seemingly now opted for breadth, releasing hundreds of mathematical discoveries in one fell swoop. The company didn’t name the AI made the discoveries, saying only that it was an ā€œinternal frontier modelā€.

Francis Johnson at University College London worked for many years on Wall’s D(2) problem – one of the hundreds of puzzles solved in OpenAI’s tranche of papers.

ā€œI worked on this problem for 25 years. I produced two books on it. I, personally, gave up,ā€ says Johnson. ā€œI’m surprised that [AI] has done it quite so quickly, but I’m not surprised that it’s done it.ā€

Johnson recounts working with AI models in recent months and being ā€œastonished with the sophistication and clarity of the analysis that they gaveā€.

ā€œIt puts us all in a very strange position,ā€ says Johnson. ā€œLet’s face it, the genie is out of the bottle now. We’re going to have to live with it. Human beings are supposed to be adaptable, so we’re going to have to adapt. I think for the moment, we just stand back and be astonished.ā€

But Kevin Buzzard at Imperial College London says we should exercise caution. He says the release included 30 papers that were relevant to his field of number theory, but only seven of those seemed impressive and only one was formally verified in Lean – a type of computer analysis that can prove mathematical results beyond reasonable doubt.

ā€œUnfortunately, acceptance of these results by the community will take time, and journalists are going to have to wait while the mathematicians do their job,ā€ says Buzzard. ā€œThe six unformalised results will have to wait until either an expert is motivated to read and check the text, or a Lean formalisation is produced.ā€

But while it is important not to get ahead of ourselves on the validity of the results, Buzzard says the general upwards trend in AI mathematics is astonishing. He says that, assuming the new results turn out to be mostly true, then we will get some kind of idea as to what the new normal is. ā€œIt has been a long time since there were humans who were experts in all of mathematics, but now we seem to have machines with this property,ā€ he says.

...

Unusually for mathematical research, the papers were published on GitHub, an online database more commonly used to share computer code.


Oops. New Scientist came with an update on Oct. 8: OpenAI mistranslated mathematics into code for its Navier-Stokes proof.

The short version:

When OpenAI announced its surprise solution to the Navier-Stokes problem, it produced one proof for humans and one for computers – however, they don't match.

The longer one:

On 8 September, OpenAI announced that it had found a solution to the Navier-Stokes problem, one of the most famous open problems in mathematics. It published the proof in two versions – one written in ā€œnatural languageā€, meaning a combination of English and mathematical symbols, as a human mathematician would write, and another written in the computer code Lean. The Lean proof is meant to be a formalisation of the natural-language version, allowing a computer to mechanically verify that all of its logical statements are true. The problem is, say Hansen and his team, that they don’t match.

ā€œThis formalisation process is trying to replace peer review,ā€ says team member Fabian Circelli, also at the University of Cambridge. ā€œPeer review would mean that human eyes look at the proofs. But what we’ve shown in this paper is that using this type of AI auto-formalisation can’t serve the same purpose.ā€

To be clear, the researchers aren’t saying that OpenAI has failed to solve the Navier-Stokes problem. It is entirely possible that both the natural-language proof and the Lean one provide a solution, just as there are hundreds of valid proofs of Pythagoras’s theorem. Instead, their point is a more subtle one: that OpenAI’s model has ā€œmistranslatedā€ when converting into Lean.

ā€œWe are not saying that the natural-language proof is wrong,ā€ says Hansen. ā€œNor do we say that it is correct.ā€ The issue is that OpenAI presents the two proofs as identical, stating on GitHub that ā€œThis repository contains Lean 4 formalizations of the results presented in [the paper] ā€˜Finite time blowup for Navier–Stokesā€™ā€.

This mistranslation occurs because the AI has to produce a Lean proof that ā€œcompilesā€, meaning that the computer code is fully self-consistent and doesn’t produce an error, says Hansen. If, in the process of auto-formalisation, the AI finds a section of the proof that doesn’t compile, it will attempt to find a workaround even if it means diverging from the proof as written in natural language.

The team’s specific claim hinges on part of the proofs called Lemma 8.6. In the natural-language proof, an equation in this part requires that a certain value be below m + 4, where m is a whole number. In the Lean proof, the equivalent value is required to be below m + 5, which is mathematically weaker.

...

OpenAI told New Scientist that it is aware of the mismatch between the natural-language proof and the Lean code, and that this doesn’t mean that either proof is invalid. It says it will rectify any errors in the natural-language proof as they are found, and will also continue the process of formalising the 722 maths papers the firm released this week, only some of which are accompanied by Lean proofs, which themselves haven’t been checked by hand.