It’s pretty simple and not hypocritical to hold these two positions simultaneously:
- It’s legal and moral to train on data you have access to, regardless of copyright.
- Nobody is obligated to provide services to you so you can obtain that data from them.
It would be hypocritical if, say, ByteDance obtained synthetic data generated from GPT-4 and then OpenAI tried to prevent them from training on the data they already obtained. But all they are doing at the moment is temporarily pausing generating new data for them. OpenAI aren’t obligated to do this and OpenAI have never argued that other people are obligated to do it for them. So no hypocrisy.