If they are doing any of that, they're wasting a whole lot of processing power. Even old text transformers like GPT2 and BERT are capable of running offline without any of that nonsense. More open models like GPT-Neo can be fully audited to prove that there is no personal data in it's training stack.
There might be merit to what you're saying, but again, you've presented no proof of this. Common logic and the currently-available technology suggests that's unnecessary.