1. The test mechanism is to use prediction of sinusoidal series. While it's certainly possible to train transformers on mathematical functions, it's not clear why findings from a model trained on sinusoidal functions would generalize into the domain of written human language (which is ironic, given the paper's topic).
2. Even if it were true that these models don't generalize beyond their training, large LLMs' training corpus is basically all of written human knowledge. So then the goalpost has been moved to "well, they won't push the frontier of human knowledge forward," which seems to be a much diminished claim, since the vast majority of humans are also not pushing the frontier of human knowledge forward and instead use existing human knowledge to accomplish their daily goals.