Laion-400M: open-source dataset of 400M image-text pairs
laion.ai
laion.ai
Nobody has sued yet though.
If this is true, how much longer until we start purposefully training AI models to be overfit and start returning inputs verbatim?
If you read an encyclopedia and use that knowledge to answer questions, a human isn't violating copyright, so an AI doing that is probably fine. If you look at all of Picasso's paintings and paint something in his style that isn't violating copyright, so training an AI to make Picasso-like paintings by training it on real Picassos is probably fine.
However looking at a painting and perfectly replicating it is a copyright violation. Same for reading a text and then writing down the same text. Both of these are just copies, not new works. So intentionally overfitting an AI to have it return the inputs with minimal changes probably makes the outputs subject to copyright from the owners of the training data.
Whether or not it's illegal is still an ongoing debate. FSF says it absolutely is illegal, but I think it's ultimately going to end up in court (though I'm just speculating).
1. Search for "monitor": https://rom1504.github.io/clip-retrieval/?back=https%3A%2F%2...
2. Note that there are several results with AOC model E2260SWDN.
3. Click the search button beside one of them.
4. Note that none of the results contain "E2260SWDN". Searching for "E2260SWDN" by itself is the same.
The back button doesn't work in this interface either. The URL changes, but it doesn't actually show that previous search.
It's impressive but if Google Images doesn't do it yet, it's almost certainly more to do with the realities of corporate policy, and scaling up & deploying it at GI scale than as a small toy demo on a static set of a few hundred million (as opposed to countless trillions of adversarially changing images). It's certainly not for lack of good NNs. Remember, Google has much better NNs than CLIP already! ALIGN was over half a year ago, and they have MUM and "Pathways" already superseding that.
As always, "the future is already here, it's just unevenly distributed"; it's just practical realities about lag - same way that the RNN machine translation models were massively better for years before Google Translate could afford to switch away from its old n-grams approach, or the speech transcription NNs were great long before they ever showed up on Google YouTube auto-captions, or BERT was kicking ass in NLP long before it became a major signal in Google Search, etc.
The results were all some sort of anime/cartoon drawings (perhaps just an issue with the Common Crawl results).
But the question remains — would a model trained on this data accurately identify a photo of two humans hugging?
I think I'm going to use this engine for all my future presentations ^^