One thing that stands in my mind from having read that is that the structural properties of the system matter in terms of its possible function, and I'm referring to the exact structure that includes the model trained on the non-gibberish text and treating that as the specific system I'm referring to. It seems when you say ChatGPT you're referring to the entire functional class of objects that implement the transformer architecture, which I'm not.
After thinking about it for a while the second thing I would note is that you're focused only on it's intrinsic function of predicting the next token, which as you say it always performs that function regardless of the training data, and I agree with that. Where we differ, I think, is I'm considering the specific structure which includes the model trained on non-gibberish data as an instrument (not a human!) that we use for it's extrinsic functions. Normally an individual or system's extrinsic functions are so narrow they're a lot easier to define, and we've never interacted with a system until now that the extrinsic function is like "understand what I say and reply accordingly". Think about it from the perspective of would you use or have any utility for GibberishGPT? No, same way I wouldn't read a book of pure gibberish.