But there is another use case where LLM's can truly help with synthetic data: the more classical classification and regression problems - specifically gathering training data. I had this exact case at work two days ago: A large dataset with a small subset of labeled data. For a binary classifier, there was a huge imbalance in the data - the ratio was roughly 75-25%. I did not have the desire to do all this manually so I used an LLM to get a list that would even out the numbers(and get a 50-50 ratio). And using the data I had, plus the additional synthetic data, the accuracy of my small classifier ended up picture-perfect(given that my actual target was "85-90%" accuracy and the actual result was just shy of 99%).