More generally, the organic training data is almost always incomplete. Samples of text leave many assumptions and inferences hidden. Like a math problem, you see the statement, but don't know the answer until you work it out. Or like a puzzle, you see the pieces but don't know the big picture until you fit those pieces together.
That is how training text samples are like unsolved enigmas, we train our models on undigested text. Often the pieces are spread over many training examples that almost aways appear separately, never together have the chance to draw a conclusion from them. Search is needed, augmenting training examples with supporting data.
Neural nets are smart at inference time but dumb at training time. They don't make those connections when they train. Instead, we need to draw those connections out by generating new text. We need to benefit from inference-time smarts before training. That means we need to use current LLMs to write the dataset of next LLMs.
All the best LLMs today used a big piece of synthetic data, including GPT-4. Datasets like Orca, Phi-1.5, ShareGPT, etc. It's also the best way to create small models (<10B) that actually work, you need very high quality, high diversity synthetic data.