Synthetic Data from Diffusion Models Improves ImageNet Classification
arxiv.org
arxiv.org
Does it give more lift than just using the training data from the diffusion model? You can search through LAION 5B in Clip embedding space and pull out lots of new training data without ever using a generative model.
Are they able to use prompts to generate augmentations around whatever invariant features - different colors or positions or scenes- in order to train better?
I've actually done something like this with GANs, but the gains came largely from generating plausible training images that were augmentations around some invariance, as I mentioned above.
… but this is 100% speculation, I have no clue. Thoughts?
We can do the same thing in text. A recent paper showed GPT-4 is better than most human labellers at tagging NLP datasets. I personally used it to generate samples for smaller models, basically the same thing this paper showed. You can even top that with a second LLM round to reflect on its generated data and filter the noisy examples.
In the end it still makes errors. And the errors GPT makes are hard to detect at scale.
> Suppose you want to optimize the ability for an LLM to do arithmetic. You prompt the LLM to generate a bunch of arithmetic questions, prompting it to show its working. Then you take that output, remove the intermediate steps, and train on the results.
Synthetic data doesn't create more learning signal, it just creates more data repeating the same learning signal with some bias assumptions that might as well be in your final model.
The difference is moreso in training times for retraining new models. If you've taken highly disparate data and compressed it into a particularly "juicy" dataset which is smaller, that could bring down training times.
That is to say, it's related to curriculum learning moreso than breaking information theory.
Isn't that the whole point of using models instead of, I don't know, huge lookup tables?
You can for example train a very large and good model and use that to generate more data to train a smaller model e.g. for a faster inference use case. That data follows a closer approximation of the underlying distribution than the small model has so it can still be used for convergence to the underlying distribution.
For some types of language model this signal could be generated though. If we have a language model which is able to generate pairs of the form (mathematical conjecture, proof attempt) in a formal language, then an automatic proof checker could generate the reward signal, i.e. proof correct / incorrect. The difficulty is probably to get the process bootstrapped since you need a certain amount of base proving capability to get the ball rolling.
AlphaGo samples from the latter distribution in a guided way (as the space of all go games is computationally intractable). It uses its learned distribution to do that guided sampling and uses the objective outcomes of the known distribution to inform its own learned distribution.
One way to think about this in the context of language modelling. Suppose I want to build a language model that says the word “goal” at least once every 2000 tokens generated. I could then repeatedly generate from the model and objectively score whether it has generated that word or not in each occurrence (the analogy of the finished go game). I then can use this objective scoring function to compete models against each other and do the alpha go style training. You can see here how the new training data is sampled from a different distribution than just regular language.
The design team explained that they couldn't blow the image up to the size of a poster because it would look too grainy as it was extremely low resolution.
The individual demanding the poster replied with a 'brilliant' solution: "Why don't you just take a picture of the low-res image, and then use that hi-resolution picture to blow up the image!"
The same applies here. You need more information to learn the signal better, by definition there is no new information in the model.
Does it, really? Analogies are great for explaining ideas, but aren't that great as logical arguments.
> You need more information to learn the signal better
You need more information than there is in the training set - which a well-generalizing model is supposed to be able to generate. Yes, there may be some bias if model doesn't generalize well, but that's why I put it in the assumption...
Yes. It's information theory.
You can add information, by making rules like the class doesn't change if the thing is a different color or size or position, or on a different background. But you can't automatically create new images from a distribution that has been learned and expect them to add information.
Humans also have far more ability than current model architectures to use memory (particularly in training), flexible context (Model context always directly follows the input distribution in training whereas we can flexibly combine different parts of our input data as we like ) and logic. We can also do live learning in ways current models can’t.
If all the books deal with France then Germans or English will never make an appearance because it's literally impossible to guess that they exist since they're not in the original books (the books here are training data if that's not obvious).
Humans learn from other humans because humans don't individually share the same information and model of the world. A science teach can speed up how you learn science by taking the compressed information and explaining it quickly (essentially what is happening in the post), but out scientific model is expanded when we have experiences that call into question the strength of our current model.
However both individually and as a society we are constantly taking in new information (sometimes more sometimes less) and using that to update our model.
What we're talking about here is approximating a distribution from a finite and comparatively small set of images. There's no "game" you can play to get new images, the training data is fixed.
The analogy might be say sampling images from a computer animation, which shows the object your interested in in different plausible poses and backgrounds, lighting, angles. There's a question of making sure this generalizes to "real" images, but that can help. But it relies on the animators knowledge of what the object can look like in different poses. That's how information gets added.
I am not quite sure I follow. Why would any of those things be true for all models?
Suppose you want to optimize the ability for an LLM to do arithmetic. You prompt the LLM to generate a bunch of arithmetic questions, prompting it to show its working. Then you take that output, remove the intermediate steps, and train on the results.
I think you would improve some benchmarks by doing this, possibly at the expense of others.
I'm sure you could do similar things with other types of model, but I think thissimple example shows the point adaquately.
In any case, information theory is not especially relevant here. It puts an upper bound on model performance but we are so incredibly far away from that bound it isn't worth worrying about.
Removing or editing the output of the model is providing new information to the model, what's improving the performance of the model in this scenario is that you are explicitly adding new information and fine tuning it on these new cases.
> information theory is not especially relevant here
It's extremely relevant because people seem to be arguing with about mathematical facts as though they were somehow opinions.
You cannot improve the performance of a model without adding new information to that model.
Deterministically removing data is not adding information, unless you are defining information in an extremely unusual way.
For the record, I tried something like what you suggest when I was doing my Master's. That was back in 2014, and I had to train a classifier for a machine learning class. I was given a training set while a separate validation set and test set were kept private (it was all set up in Kaggle as a private competition). To clarify, the idea was that you trained your classifier of choice on the training set, then labelled the validation set with your trained model and submitted the labelling to get a score that you could use to improve your model. The last day of the competition you had to make a choice and submit a final model, that would be evaluated on the test set, for which you had no information.
The problem was that the training data was not very much. There was more data in the validation set, but the data in the validation set wasn't labelled. So I tried to label the validation set with a model I trained on the training set. And, what would you know. My classifier scored 100% accuracy on the validation set. But when I submitted my trained model on the test set it did much worse, I think close to 60% or so.
Empirically demonstrated then: you can't dogfood a classifier to a better version of itself. When you train a classifier on some data, the classifier learns the underlying distribution of the data, with some amount of error. If you then label new data with the trained classifier and retrain the classifier on its own labelling, you end up multiplying the error.
Btw, that doesn't change with language models, large or small, and it doesn't make a difference whether the model has an unsupervised training step or not. As long as your model is, well, modelling, some unknown true distribution and incurring some error, reusing the trained model to generate new data will generate data with error.
So I don't think what you say can work and I'm curious to understand why you think it will. What are you saying will happen, exactly, if you dogfood an LLM's generations, like you suggest, that will make it improve?
In this case it is hardly an analogy. With respect to Information Theory, compression and machine learning are essentially the same thing [0]. If you can understand why taking a hi-res photograph of a lo-res image does not create more information (and therefore cannot be used to make a hi-resolution poster) then it's essentially the exact same logic explaining why outputs of a model cannot be used to improve that same model.
More to this point, which other comments have already pointed out, what's interesting about this post is it asks whether or not very, very large data sets can be reasonably compressed by models such that they can simulate training another model while requiring less data.
From an information theoretic perspective this latter case is both completely sound and of potentially big practical importance.
I do have to say, I'm a little surprised that basic information theory is no longer common knowledge among the HN community (more-so surprised at a resistance to it), especially among people interested in Machine Learning. Out of curiosity, do you have a background in ML and/or Comp Sci?
0. https://en.wikipedia.org/wiki/Data_compression#Machine_learn...
If by "a model [that] truly generalizes" you mean one that generalises outside its training distribution, then neural nets can't train such models. I don't reckon any machine learning approach can do that.
In any case this is ImageNet we're talking about that's been done to death many times over already. Far from learning more general models, after a certain point in time improvements in accuracy have meant that models are getting better at overfitting. Same with MNIST, where that happened a long time ago.
Obligatory reference to defend against this comment being knee-jerked to oblivion:
>> This stands in sharp contrast with what deep nets do, which I would call "local generalization": the mapping from inputs to outputs performed by deep nets quickly stops making sense if new inputs differ even slightly from what they saw at training time.
Francois Chollet, The limitations of deep learning
> Even with this data, you could not train a deep learning model to simply read a product description and generate the appropriate codebase.
Would be make the same claim now I wonder?
>> read a product description and generate the appropriate codebase
Chollet here is talking about creating an entire application, not short code snippets or larger functions. Producing short programs from specifications was already possible when Chollet wrote the above-linked article. It is still not possible to genera the entire codebase for an application by reading a product description. Not anywhere near that.
Example: You ask ai something, it takes 3 different tries to get what you want. AI Critiques itself, decides what it could do better and rolls that into training, perhaps it goes line by line and scores how 'helpful' that was to give certain negative/positive bias to new training datasets.
https://www.youtube.com/live/j0z4FweCy4M?feature=share&t=571...
1. The human brain must be doing something analogous to this, continuously, to learn from a relatively small numbers of samples.
2. Going forward, generating high-quality synthetic data looks likely to become a standard practice for training models.
3. Whoever has the largest, highest-quality synthetically generated datasets will have the best-performing AI models.
To your point though, three are definitely some proprietary pockets, like the Shutterstock/openai thing. Or the various RL layers ala chatgpt. But a lot of the value there is in the labeling, and I expect we'll see more competing versions that are open (we already are)