Leaving the legal aspects of crawling aside, I think there is an important distinction here between 1. "can you procure it" and 2. "do you have enough money to process it all"
1. Yes, I think almost anyone can write code to procure the training corpus, in theory, and test it on a small scale
2. No, only the biggest labs and universities have enough resources to process such huge amounts of data and iterate on models with that scale. But that's just a matter of resources that can be overcome with partnerships between industry and academia that are common anyway. All the big labs already have huge efforts underway to reproduce GPT-X and it's just a matter of time before they catch up.