This is also true of apparently restricted tasks like translation. You might think initially that a task like 'translate this paragraph from English to French' is not in any sense 'Turing-complete', but if you think about it, it's obvious you can construct paragraphs of text whose optimally correct translation on a token-by-token basis requires brute-forcing a hash or running a program or whatnot. Like grammatical gender: suppose I list a bunch of rules and datapoints which specify a particular object, whose grammatical gender in French may be male or female, and at the end of the paragraph, I name the object, or rather _la objet_ or _le objet_. When translating token by token into French... which is it? Does the model predict 'la' or 'le'? To do so, it has to know what the object is before the name is given. So it has an incentive from its training loss to learn the reasoning. This would be a highly unnatural and contrived example, but it shows that even translation can embody a lot of computational tasks which can induce capabilities in a model at scale.