Don't get me wrong, DALL-E is remarkable. But it's remarkable in almost the exact same way ELIZA was remarkable in 1966, and Markov Chain generative models were in the 1990s. All of these, given the context computational powers at the time, were both miracles and parlor tricks at the same time.
The trouble is that these demos are impressive, but not much has changed in how the world works. I work in AI/ML and every places I've seen the practical applications of these techniques has ranged the equivalent to adding sprinkles on a cake, to completely useless engineering nightmares that would ultimately create more value with their removal.
Yea, a 3.5 billion parameter model is neat, but we know that in 1970 when we couldn't imagine training such a thing. The problem is that we've made essentially no progresses other than showing when you pour incredible human and physical resources into a pot the result looks cool.
But when you do the accounting, and ask yourself "what has really changed with all these fantastic innovations in machine learning" the answer is surprisingly little.
Dall-E 2 is the type of cool that will be figure 12-3 in a undergrad textbook in 20 years. Students will go "oh that's cool" and turn the page.
20 years ago (2002) Deep Blue had beating reigning world chess champion Kasparaov was old news.
Unsolved problems were things like unconstrained speech-to-text, image understanding, open question answering on text etc. Playing video games wasn't a problem that was even being considered.
I was working in an adjacent field at the time, and at that point it was unclear if any of these would ever be solved.
> In the end they all fell to specific methods that did not provide general progress.
In the end they all fell to deep neural networks, with basically all progress being made since the 2014 ImageNet revolution where it was proven possible to train deep networks on GPUs.
Now, all these things are possible with the same NN architecture (Transformers), and in a few cases these are done in the same NN (eg DALL-E 2 both understands images and text. It's possible to extract parts of the trained NN and get human-level performance on both image and text understanding tasks).
> While the current progress on image related problems is great, if it does not lead to general advances then an AI winter will follow.
"current progress on image related problems is great" - it's much more broad than that.
"if it does not lead to general advances" - it has.
https://twitter.com/nickcammarata/status/1511861061988892675
And as said before, this is only the beginning.
As far as something to play with to get interesting ideas, and to take a first few cuts at implementing them, DALL-E is great. Then let's bring the human's technical skill and ability at curation into the mix.
Harry Roolaart
DALL-E 2 is useful now. It's easy to imagine it replacing 99designs, Getty Images, and most of the digital art services on fiverr. I can't imagine what is going to happen in the world of print-on-demand tshirts.
And of course porn... what a market.
I don't think anyone (yet) imagines all the humans at the NYT will be replaced by GPTX+DALL-E. We still have editors even though humans author the articles.
If you look closely at most of the images, they don't look quite right. There are usually artefacts, or slight misunderstandings of the brief etc. Its about 80% of the way there, but I think that last 20% is going to be a lot more difficult.
I'm sure there are some cool applications for this - maybe if you need a quick and cheap image for your newsletter, personal website or an experimental game for example.
For a serious commercial application I would think it would be easier and safer to pay someone.
I think this understates the market. For every New York Times article, there are millions of newsletters, analyst reports, tshirts, websites, logos, etc. Many of them are abstract eye candy and don't need photorealism.
Actually, even the high-budget articles often have pretty abstract art. I read this Wired article yesterday, and all the illustrations could easily have been generated:
https://www.wired.com/story/tracers-in-the-dark-welcome-to-v...
Now we can do that in a model that runs locally on your phone. This is a pretty big deal.
And calling it "curve fitting" is just disingenuous. There's quite a lot of evidence that "curve fitting" will scale for at least a few more orders of magnitude. Who knows, maybe the human brain is just "curve fitting".
No it hasn't. And the problems are clear - it's impossible to express anything in the rigid hierarchies that symbolic AI requires.
Representing "Britain" in geographic, language, political and economic hierarchies does not allow a model to do any reasoning about what "British sense of humour" means.
"Softer" structures that represent concepts as a "blob" in a multi-dimensional space is clearly a better approach (aka embeddings, and the even better representations that more complex models use are even better).
Representing "Britian" as blob in a multidimensional space that is adjacent to concepts like "satire" and "surreal" as well as people like "John Cleese" lets a model reason about what "British sense of humour" means without being specifically trained.
As GPT-J[1] says when prompted with "A good example of the British sense of humour":
A good example of the British sense of humour is found in George Orwell’s novel, The Lion and the Unicorn. It’s a satire on socialism that is much more sophisticated than anything in contemporary Leftist intellectual thought. It was written during the war, in 1944. The story opens with a visit to a pub in a fictional village in England. The pub is named the Unicorn, but Orwell (or George, as he calls himself) has decided to call it the Lion. He explains:
“The Lion is a pub, just like any other pub, where people drink in the evenings, and talk about their daily business and their hobbies and where they exchange ideas, views, points of view. But the difference is that nobody in the Lion ever argues about anything. They just sit there, saying nothing, drinking nothing, not even beer. People can come to the Lion and buy beer, and leave the Lion and not buy beer. Beer is freely on sale in the Lion, but nobody ever buys it.
(One should note that George Orwell's "The Lion and the Unicorn" is nothing to do with a pub where no one buys beer. BUT the joke is kind of exactly like a British sense of humour).
About usefulness - the CLIP part of the model is a ready made zero shot image classifier. It reduces the amount of work needed for simple image classification tasks to just naming the classes. The generative part is good enough for illustrations. It will make an average web designer have the powers of a graphical artist.
Unfortunately the models are restricted and expensive today. I hope to see a real open AI initiative to train such models and share the weights, but can't hope that from OpenAI.
That "count the objects" app that can tell you how many items you have in a photo seems like a very practical application that wasn't possible with traditional CV before ML.
Not far off from image recognition, save for the background segmentation ;)
machine is briefly described here, I've seen a more thorough breakdown somewhere out there... https://distributedmuseum.illinois.edu/exhibit/biological_co...