Namely, what sort of computational resources had been dedicated to the problem before? Because DeepMind suddenly throwing their weight at a problem that was previously the focus of a handful of random math grad students would make it hard to benchmark the ML advance that was made here -- like would it be possible to find a similar solution with a ton of compute and more traditional genetic algorithms?
I wouldn't be surprised if the answer were no. Protein folding competition was a pretty big space in biology before getting absolutely destroyed by Alpha Fold. But I also wouldn't be entirely surprised if the answer were yes. The amount of hype around LLMs right now is crazy, probably half the news I see about them turns out to be very exaggerated upon further evaluation.
ETA: the headline here is definitely exaggerated, because there was also human in the loop to refine what the LLM was generating. At a glance the technical article doesn't benchmark enough against alternatives to the LLM component in their workflow IMO. But it is entirely possible they tackled well established enough open problems, such that prior work already handled those control cases decently.
I'd love to know what someone in this space of mathematics thinks about the paper! Would it have generated much buzz if they got these same results like 3 years ago? Would it be accepted to Nature if they found they could accomplish something similar using their framework even if it was an RNN in the loop?