A big distinction is that you can built on top (fine-tune) thus released models as well as if they released the pre-training data.
Synthetic data can be more diverse if you sample carefully with seeded concepts, and it can be more complex than average web text. You can even diff against a garden variety Mistral or LLaMA and only collect knowledge and skills they don't already have. I call this approach "Machine Study", where AI makes its own training data by studying its corpus and learning from other models.
Mistral models are one example, they never released pre training data and there are many fine tunes.