That's how you train neural network with synthetic data so it extracts actual meaning.
That's how humans also learn ie. adding numbers. First there is naive memoization, followed by more examples until you get it.
LLM training seems to be falling into memoization trap because models are extremely good at it, orders of magnitude better than humans.
IMHO what is missing in training process is this feedback explaining wrong answer. What we're currently doing with training is leaving out this understanding as "exercise to the reader". We're feeding correct answers to specific, individual examples which promotes memoization.
What we should be doing in post training is ditch direct backpropagation on next token, instead let the model finish its wrong answer, append explanation why it's wrong and continue backpropagation for final answer - now with explanation in context to guide it to the right place in understanding.
What all of this means is that current models are largely underutilized and unnecessarily bloated, they contain way too much memoized information. Making model larger is easy, quick illusion of improvement. Models need to be squeezed more, more focus needs to go towards training flow itself.