DALL·E now available in beta
openai.com
openai.com
It's like The Onion, but all the articles are made with GPT-3 and DALL·E. I start with an interesting DALL·E image, then describe it to GPT-3 and ask it for an Onion-like article on the topic. The results are surprisingly good.
I think I got it to even fill the title given a picture, something like “Article picture caption: Man holding an apple. Article title: ...”. Might experiment more with that in the future.
The combination of photo/title feels like they come from the more absurd articles published by theonion.
If we aren't living in a simulation, it's just a matter of time...
"He's a hate-fuelled neurotic farmboy searching for his wife's true killer. She's a tortured insomniac snake charmer from a family of eight older brothers. They fight crime!"
Here's an implementation in Perl.
>He's an unconventional gay paranormal investigator moving from town to town, helping folk in trouble. She's a violent motormouth wrestler from the wrong side of the tracks. They fight crime!
>He's a Nobel prize-winning sweet-toothed rock star who believes he can never love again. She's a strong-willed communist widow with a knack for trouble. They fight crime!
>He's an obese white trash barbarian with a secret. She's a virginal thirtysomething traffic cop with the power to bend men's minds. They fight crime!
Put a guinea pig in there and you'd get the same effect.
This was really funny :)
http://dailywrong.com/man-finally-comfortable-just-holding-a...
That said, with a little tweaking, these technologies can - and probably already are - being used to churn out blogspam websites left and right, fully automatic.
Easiest way for free SSL would be to just throw the domain on CloudFlare :)
Now I'm using that content in the video game. I wonder if you could use these articles as some fake news in your game, too. :)
Thank you! Bookmarked!
If you want to see some really creepy AI generated human "photo" faces, take a look at Bots of New York:
How about wording your comment in a way that highlights why it’s a shame these pictures aren’t accessible for those without a Facebook account, and skip the whole “you’re murdering puppies” bit?
The hyperbole is good marketing for a certain audience.
Hot dang. Some Reddit subs can be auto-generated now.
https://www.reddit.com/r/SubSimulatorGPT2/
but yea, it will have generated images now.
I guess Markov chains for the headings, Dall-E for images, maybe GPTx for comments. And/or the GPT models should be made wackier somehow—less coherent, perhaps.
Can DALL-E render Bat Boy?
http://dailywrong.com/wp-content/uploads/2022/07/DALL%C2%B7E...
http://dailywrong.com/wp-content/uploads/2022/07/DALL%C2%B7E...
The images don’t scale properly with the rest of the site, they’re massive compared to the content.
Why do the images load so slow though?
In fact, I've been glad to have a 50/day limit, because it helps me contain my hyperfocus instincts.
The information about new pricing is, to me as someone just enjoying making crazy imagines, a huge drag. It means that to do the same 50/day I'd be spending $300/month.
OpenAI: introduce a $20/month non-commercial plan for 50/day, and I'll be at the front of the line.
On the personal side, I've been getting into game development, but the biggest roadblock is creating concept art. I'm an artist but it takes a huge amount of time to get the ideas on paper. Using DALLE will be a massive benefit and will let me expedite that process.
It's important to note that this is not replacing my entire creative process. But it solves the issue I have, where I'm lying in bed imagining a scene in my mind, but don't have the time or energy to sketch it out myself.
this is what I really like about DALLE-mini, it's ability to create pretty good basic outlines for a scene. it's low resolution enough that there's room for your own creativity while giving you a good template to spring off from. things like poses, composition of multiple people, etc.
Generating itself can be art. I’m not going to win a Pulitzer here, it’s for the personal joy of it, but I will certainly never get tired of it.
Knowing that it will only get better - animation cannot be far behind - makes me feel genuinely excited to be alive.
It’s good at “in the style of” but there’s no “in a new style”.
It has a house style too that tends to feel Reddit-like.
If your task is "show me something that breaks through our hyperspeed media", then I guess some obscure museum is a better place than an ML model.
If your task is "find the best variation on theme X" or "quick draft visualization", they are often very useful. I am sure there will be many further tasks to which current and future models will be well suited. They are not magic picture machines. At least not yet.
Either way, I feel like your view is an exhaustingly pessimistic take on AI-generated art. I mean, sure, most of what DALL-E generates is pretty mundane, but other times I have been surprised at how bizarre and unique certain images are.
You seem to imply that because an AI is not human, its art is not imbued with meaning or originality -- but I find that an AI's non-human nature is precisely what _makes_ the art so original and meaningful.
At the extreme limit, maybe. But within art or even digital art, then new styles are actually not that rare, humans are pretty good at generating them. Maybe they grab inspiration from nature, visual phenomena, etc, so in that sense it's not "new" but it is "new to the medium". In art you new styles all the time. DALL-E will never do that by it's very nature, and so it's easy to see how it's boring.
And that's just the stylistic level, but it's happening at almost all levels. It's almost definitional that it doesn't innovate, only remix.
It's strange framing this as pessimistic, it's not really optimistic nor pessimistic, it just is. It's also not AI, and that's important to realize: it's a statistical model that generates purely based on pre-existing training. It's very nature is without-meaning and without-originality. That doesn't detract from it being cool or interesting or helpful or enjoyable. I find it cool and useful.
But it's not innovative or creative or meaningful by itself.
That's a pretty bold claim. What are humans but statistical models that generate based on pre-existing training -- and yet, are humans not with-meaning and with-originality?
Of course, human brains are some large order of magnitude more complex than the neural nets that underlie most AIs, but we can already see areas where these "simplistic" AIs outperform humans on specific tasks. So what prevents the arts from being one of those areas? If not now, in some not-too-distant future?
One thing humans seem to have that is beyond statistics is creativity. In that stats explain what is, and creativity takes what is and makes a dot outside that other people appreciate. No model has demonstrated even attempting the dot, let alone having a good chance of success. What DALL-E does is draw a dot between a few existing points, but never outside.
Humans incidentally have three more things that make for interestingness: emotions derived from feelings, long term memory, and roughly storytelling (~ an ability to turn long term memory into long form recall with a specific reaction intended to a specific audience). I don’t think ML has any of those, but it likely (eventually) gets the latter two.
Meaningness/interestingness require at least a few of those, and it’s what puts most art in a different category than games or math.
I'd disagree.
- One of the first queries I did made some interesting chairs, I would genuinely buy the first if it was sanely priced: https://www.ryanmercer.com/ryansthoughts/2022/6/17/dall-e-2-...
- One of the first H R Giger inspired works I did made some really interesting computers, I would (and may) hang some of these https://www.ryanmercer.com/ryansthoughts/2022/6/19/dall-e-2-...
- The first wood carving query I did generated solid gold of Vikings eating pizza, this is the kind of thing I'd see in a hole in the wall restaurant and absolutely love https://www.ryanmercer.com/ryansthoughts/2022/6/17/dall-e-2-...
- The first "painting" in this series I may very well print and hang in my office https://www.ryanmercer.com/ryansthoughts/2022/6/27/dall-e-2-...
- These H.R. Giger chairs 100% look like something I'd expect to see in a modern art museum https://www.ryanmercer.com/ryansthoughts/2022/6/24/dall-e-2-...
- If you want to create new characters for something, and are lacking inspiration, I think it could be extremely useful to artists. For example these variations of Don't Hug Me I'm Scared https://www.ryanmercer.com/ryansthoughts/2022/6/21/dall-e-2-...
I've got thousands of queries, and a LOT of them have generated things I genuinely see as having artistic value, I've probably got 200~ images that I would 100% hang/display in my home ('woodcarving' and 'stone carving' queries rarely disappointed me)
is it some unique form of art, no, but can it produce works I want to see in a medium or style that already exists to a level that it is believable as authentic human art, absolutely.
People like me, with zero artistic ability, are able to take part in creating visually pleasing works. I imagine artists would also find great value in it, being able to feed a few queries in with what they are thinking of creating to draw inspiration, or even putting their own work in and generating variations that may lead to inspiration for new works.
I'm not going to try and profit from the images, I don't need them for any business uses or anything, so to me it was a fun for a while and now just something I'll largely put out of mind.
When they're free, it's pretty cool. But charge an amount where there's actual profit in the product? Suddenly seems very expensive and not economically viable for a lot of use cases.
We are still in the "you need a supercomputer" phase of these models for now. Something like DALLE mini is much more accessible but the results aren't good enough. Early early days.
What are the resources needed to train this model?
If someone just gave you the model for free, what resources would you need to use it to generate new results?
Training requires even more GPUs, and I wouldn’t be surprised if they used more than 100 and trained over 3 months.
Based on this blog post where they scale to 7,500 'nodes', they say:
> A large machine learning job spans many nodes and runs most efficiently when it has access to all of the hardware resources on each node.
So I wouldn't be surprised if they do have a total of 7500+ GPUs to balance workloads between. TO add, OpenAI has a long history of getting unlimited access to Google's clusters of GPUs (nowadays they pay for it, though). When they were training 'OpenAI Five' to play Dota 2 at the highest level, they were using 256 P100 GPUs on GCP[0] and they casually threw 256 GPUs at 'clip' for a short while in January of 2021[1].
As for how they do it, see these posts:
https://openai.com/blog/techniques-for-training-large-neural...
https://openai.com/blog/triton/
But I seem to remember they were running 1,000+ 32gb GPUs for 3 months to train it and keeping that infrastructure running day-to-day and tweaking parameters as training continued was the bulk of the 100 pages. It is beyond the reach of anybody but a really big company, at least in the area of very large models, and the large models are where all the recent results are. I wish I was more bullish on algorithm improvements meaning you can get better results on less hardware; there will definitely be some algorithm improvements, but I think we might really need more powerful hardware too. Or pooled resources. Something. These models are huge.
Is https://github.com/facebookresearch/metaseq/blob/main/projec... what you're referring to?
This means the parameters of the trained model fit in something like 7GB (decoder only, half-precision floats) to 24GB (full model, full-precision). To actually run the model, you will need to store those parameters, as well as the activations for each parameter on each image you are running, in (video) memory. To run the full model on device at inference time (rather than r/w to host between each stage of the model) you would probably want an enterprise cloud/data-center GPU like an NVIDIA A100, especially if running batches of more than one image.
The training set size is ~97TB of imagery. I don't think they've shared exactly how long the model trained for, but the original CLIP dataset announcement used some benchmark GPU training tasks that were 16 GPU-days each. If I were to WAG the training time for their commercial DALL-E 2 model, it'd probably be a couple of weeks of training distributed across a couple hundred GPUs. For better insight into what it takes to train (the different stages/components of) a comparable model, you can look through an open-source effort to replicate DALL-E 2.[2]
[0] https://cdn.openai.com/papers/dall-e-2.pdf [1] https://openai.com/blog/clip/ [2] https://github.com/lucidrains/dalle2-pytorch
> you would probably want an enterprise cloud/data-center GPU like an NVIDIA A100, especially if running batches of more than one image.
That doesn't seem so bad.
looks up price of NVIDIA A100 - $20,000
oh...ok I'll probably just pay for the service then
For the half-precision version at 7GB there are a ton more options (the RTX 3060 has 12GB for example at ~$450).
I do hope that the conversation starts to acknowledge the difference between sunk costs and running costs.
Employees, office leases and equiment are all happening, regardless and ongoing.
Training DALL-E 2: very expensive, but done now. A sunk cost where every dollar coming in makes the whole endeavor more profitable.
Operating the trained model: still expensive, but you can chart out exactly how expensive by factoring in hardware and electricity.
I believe that by not explicitly separating these different columns when discussing expense vs profit, we're making it harder than it needs to be to reason about what it actually costs every time someone clicks Generate.
They really aren't that large by the contemporary scaling race standards. DALLE-2 has 3.5B parameters, which should fit on an old GPU like Nvidia RTX2080, especially if you optimize your model for inference [1][2] which is commonly done by ML engineers to minimize costs. With optimized model, your memory footprint is ~1 byte per parameter, and some less than 1 ratio (commonly ~0.2) of all parameters to store intermediate activations.
You should be able to run it on Apple M1/M2 with 16GB RAM via CoreML pretty fine, if an order of magnitude slower than on an A100.
Training isn't unreasonably costly as well: you can train a model given O(100k)$ which is less than a yearly salary of a mid-tier developer in silicon valley.
There is no reason these models shouldn't be trained cooperatively and run locally on our own machines. If someone is interested in cooperating with me on such a project, my email is in the profile.
1. https://arxiv.org/abs/2206.01861
2. https://pytorch.org/blog/introduction-to-quantization-on-pyt...
Still, I think optimization of diffusion models for efficient inference isn't yet pushed to the limits. At least if we look at what's available to the public - AFAIK public inference software distributions didn't even quantize their weights.
Excet we’re already waiting 90+ seconds for dall-e mini.
Multimodal.art (https://multimodal.art/) is working on a free version of something like DALLE, though it's not that good as of yet.
directly from the website
I’ve been creating generative art since 2016 and I’ve been anxiously waiting for my invite. I wont be able to afford to generate the volume of images it takes to get good ones at this price point.
I can afford $20/mo for something like this but I just can’t swing $200 to $300 it realistically takes to get interesting art out of these CLIP-centric models.
Heck, the initial 50 images isn’t even enough to get the hang of how the model behaves.
Meanwhile we should prepare ourselves for a future where the best generative models cost a lot more as these companies slice and dice the (huge) burgeoning market here.
Majestic diffusion - https://github.com/multimodalart/majesty-diffusion
Centipede diffusion - https://colab.research.google.com/github/Zalring/Centipede_D...
Colab is dogshit if you don't pay
Wait until the next edition comes out where it automatically learns the sorts of things that crack you up and starts generating them without any input from you.
It'll use hideous amounts of compute.
If you get a large percentage hooked on TikTok you can change and undermine democracy.
Starting to believe representative democracy and social media are incompatible.
My brother is a digital artist and while excited at first he found it to be not all that useful. Mainly because it falls apart with complex prompts, especially when you have a few people or objects in a scene, or specific details you need represented, or a specific composition. You can do a lot with in-painting but it requires burning a lot of credits.
This belongs on /r/linkedinlunatics
[0]: https://dallery.gallery/wp-content/uploads/2022/07/The-DALL%... (PDF)
I hope that every science teacher that can - provide this to every student. This is the future they live in now. They should know these as well as they know how to install an app on a device.
Wait until we have a DALL-E -- Enabled Custom EMOJI stream - whereby, every text you send out has it corresponding DALL-E resultant image for every txt --
Then we can compare images from different people at different times but the prompt was identical... and see what the resultant library of emoji<-->PROMPT looks like?
What about using Dall-e as a watermark for 'nft' signature 'notary' of an email.
If DALL-E provided a unique PID# for every image - and that PID was a key that only the OP runner of the image has - it can be used to authenticate an image to a text source... ??? (Assuming that no two prompts have the same result ever, but assigning a unique id that CAN be used to replay the image to verify it was generated when an original email/SMS was actually sent - it could be a unique way to timestamp authenticity/provenance of a thing...
However the artists being featured in DALL-E's newsletters can't stop gushing about 'the new instrument they are learning how to play' and other such metaphors that are meant to launder what's going on.
My theory is that the professions most at-risk for automation are acting on their anxieties. Must not be a lot of freelance artists on HN, and a whole lot of programmers.
I think the artists have an even clearer case. I don't think GitHub Copilot is ready to steal anyone's job yet. But DALL-E is poised to replace all formerly commissioned filler art for magazines, marketing sites, and blogs. Now the only point to hiring a human is to say you hired a human. Our filler art is farm-to-table.
On the other hand, the market for stock photography was already decimated by the internet. Where previously skilled photographers would create libraries of images to exemplify various terms and sell these as stock, in the last decade or so, an art director with the aid of a search engine could rapidly produce similar results.
Think of it like a set of powertools saving you time over manual tools.
That's not my problem with Copilot. I think tools and methods that can free human from some amount of work are good in a correctly organized society. They have been existing for a long time, too. They let us free time for other stuff that can't be automated. This extra free time could theoretically let us have more leisure or rest time too. I also trust myself to be able to learn another job if mine can ever be automated.
But I don't want my work to be reused under terms I don't approve of. There are some things I don't want to help with my work and this is reflected in the licenses I choose. I totally sympathize with artists who don't want their work to be reused in ways they don't like. I don't find this hard to understand. I also don't find it hard to understand that if an artist do some work that you should pay for to use is not happy with their work being reused without being paid. They should get paid a tiny bit for each generated art if theirs is in the training set, and only if they approve this use. That's would be only fair, the set would not be possible without those artists.
(Good for me, my personal code is not on GitHub for other, older, reasons)
However if the result is adequately different, I don't see how it is different from someone viewing other's work and then being "inspired" and creating something new. If you think about it the vast majority of things are built on top of existing ideas.
(I really don't know, and I didn't find anything about it on their site.)
As long as DALL-E isn't caught painting out a 1-to-1, reverse searchable copy of an image, its not really as bad as copilot, IMO.
The issue isnt just that copilot is trained on my GPL code, its that it might decide to copy paste lines from it, including my comments, etc.
> Preventing harmful images: We’ve made our content filters more accurate so that they are more effective at blocking images that violate our content policy — which does not allow users to generate violent, adult, or political content
What is defined as political content? Can I prompt DALL-E to draw ”Fat Putin”?
Well, I can't go ask Caravaggio or Gentileschi to paint my query since they've been dead hundreds of years. But being able to to feed a query containing much more modern concepts in and get a baroque style painting in that specific style is wonderful.
Plus what has already been said about a lot of art being an imitation/derivation of previous works.
For now at least, there is a detectable difference in variety.
Getting humans to refine your data is the best solution right now and many companies and researches go with this approach.
Clip was trained on 400,000,000 images, GPT is roughly 180B tokens, at 1-2 tokens per word, that's 120,000,000,000 words.
Source ?
All those big models are trained with data for which the source is not known or vetted. The amount of data needed is not human-refinable.
For example for language models we train mostly on subsets of CommonCrawl + other things. CommonCrawl data is “cleaned” by filtering out known bad sources and with some heuristics such as ratio of text to other content, length of sentences etc.
The final result is a not too dirty but not clean huge pile of data that comes from millions of sources that no human as vetted and that no one in the team using the data knows about.
The same applies to large images dataset, e.g. Laon 400m that also comes from CommonCrawl and is not curated.
Made with gpt3
Unless I'm missing something, these seem pretty darn good
Once the API is released, this will be easier to do in a programmatic fashion.
Note: Depending on how many times you do this... I could see there being a continuity problem with the extremes of the image (eg: the far left has no knowledge of the far right). An alternative could be to scale the image down and mask the borders then later scale it back up to the desired resolution.
This scale and mask strategy also works well for images where part of the scene has been clipped that you want to include (EG: Part of a character's body outside the original image dimensions). Scale the image down, then mask the border region, and provide that to the generation step.
https://apps.apple.com/us/app/waifu2x/id1286485858
I paid extra to get the higher quality model using the in-app purchase option. It crushes the phone's battery life, but runs in only ~10 seconds on an iPhone 13 Pro for a single 1000x1000 input image.
Considering waifu2x is the name of an algorithm I assumed it was just that algorithm. There's also no mention of other models on the demo page or the Github page as far as I can see.
The confusion originates from the fact that I was using a GUI project for Waifu2x called "Waifu2x Extension GUI" (https://github.com/AaronFeng753/Waifu2x-Extension-GUI) which other than Waifu2x also supports other algorithms like Real-ESRGAN, Real-CUGAN, SRMD, RealSR, Anime4K, RIFE, IFRNet, CAIN, DAIN, and ACNet.
So as you said Cupscale is surely more advanced than Waifu2x (the single algorithm), but do you think it's also better than Waifu2x Extension GUI?
https://www.topazlabs.com/gigapixel-ai
No kidding.
On-demand stock photo generation probably is the next step, particularly when combined with other free media services (Unsplash immediately comes to mind). Simply choose a "look" or base image, add contextual details, and out pops a 1 of 1 stock photo at a fraction of the cost of standard licensing. It'll be very exciting seeing what new products/services will make use of the DALL-E API, how and where they integrate with other APIs, use cases, value adds like upscaling and formatting, etc.
And this is the first picture I got: https://labs.openai.com/s/lSWOnxbHBYQAtli9CYlZGqcZ
It got it a bit strong on the depth of field and I don’t like the angle but I could iterate a few times and get a good one.
For the first few days when it was announced I use to look deep even in real photos in search of generative artifacts. They are not so difficult to spot now, most of the times anyway.
Heck: If the cost to entry is prohibitively low they might do it at a loss and take over the site
It's very good at generating art style images. These kind of images are mostly amazing most of the times. But the Photorealistic images only work with cherry picking.
Me and you must have very different definitions of "cherry picking". For prompts that fall within it's scope (i.e not something unusually complex or obscure) I get usable results probably 90% of the time.
Can you give me some examples of prompts that you tried where you found good results difficult to obtain?
It did generate good dslr like face closeups, as good as Nvidia does, most of the times but not always. Sometimes there are weird artifacts and face does not make sense.
Dslr style blurry photos are mostly good. From the looks of images I follow, imagen is probably more believable. Don't know how much cherry picking goes on there. See this thread [1] for example. I failed to generate image like this (honey dress) in dalle2.
[1]: https://www.reddit.com/r/ImagenAI/comments/w3saku/creating_i...
Maybe it would be cheaper. I imagine it would one day. And maybe it would have a more liberal usage license.
At any rate, I look forward to this. And I look forward to the inevitable debates over which is better: AI generation or photographer.
I wonder if there is a way for DALL-E to generate a character, then persist that character over subsequent runs. Otherwise, it would be pretty difficult to generate illustrations that depict a coherent story.
Example ...
Image 1 prompt: A character named Boop, a green alien with three arms, climbs out of its spaceship.
Image 2 prompt: Boop meets a group of three children and shakes hands with each one.
It's not the "specify the character positions in text" proposed, but still a neat take on using this sort of AI for art.
Then use inpainting to only preserve that pose and generate new content around it. It’s definitely not perfect.
Then put that at the side of a transparent image, and use as the prompt, "Two identical aliens side by side. One is jumping"
There seems to be a post-DALLE obscenity detector on openAI's tool, as so far I've found it to be entirely robust against deliberate typos designed to avoid simple 'bad word lists'. Ask it for a "pruple violon" and you get purple violins... you get the deal.
"Metastable" prompts that may or may not generate obscene (content with nudity, guns, violence as I've found) results sometimes shown non-obscene generations, and sometimes trigger a warning.
Why do that? Just refusing to run my query is sufficient. Who is harmed if I continue to bang my head against that wall?
Stopping more than so many attempts makes this much harder / much less likely.
"Starting today, users get full usage rights to commercialize the images they create with DALL·E, including the right to reprint, sell, and merchandise. This includes images they generated during the research preview."
I assumed this was going to be the sticking point for wider usage for a long time. They're now saying that you have full rights to sell Dall-E 2 creations?
Of course, that clause won't deter a third party from filing a lawsuit against you if you commercialize a generated image too close to something realistic, as the copyrights of AI generated content still hasn't been legally tested.
[1]: https://felixreda.eu/2021/07/github-copilot-is-not-infringin...
A lawyer could argue that the algorithm is producing a derivative work of the copyrighted input.
While nothing has been commercialized yet on the DALLE2 subreddit, I know that it can do Dave Choe's work remarkably well. I also saw Alex Gray's work to be close, but not really identical either. It wasn't as intricate as his work is.
It will be interesting if this takes off and you have a sort of Banksy effect take over where unless it's a physical piece of art it doesn't have much value and is only made all the better because of some sort polemic attached to it, eg Girl with balloon.
There will be outliers of course but I would be shocked if there's much of a market in it for at least the present.
The US Copyright Office did make a ruling that might suggest that recently[1], but crucially, in that case, the AI "didn't include an element of human authorship." The board might rule differently about DALL-E because the prompts do provide an opportunity for human creativity.
And there's another important caveat that the felixreda.eu link seems to miss. DALL-E output, whether or not it's protected by copyright, can certainly infringe other copyrights, just like the output of any other mechanical process. In short, Disney can still sue if you distribute DALL-E generated images of Marvel characters.
1: https://www.theverge.com/2022/2/21/22944335/us-copyright-off...
Of course it's unlikely to produce the exact same image, or if it does, you've also discovered an incredible image compression algorithm.
Both technologies have billions of dollars of R&D and tens of thousands of engineers behind supply chains necessary to create the button that a user has the press.
> …you own your Prompts and Uploads, and you agree that OpenAI owns all Generations…
Edit: Joined the discord via the beta and got in. Thanks a lot for the heads up!
It’s clearly personally preference, but I loathe Discord but love it for MidJourney. As you said, there’s an interactive element where I see other people doing cool things and adapting part of their prompts and vice versa. It really is fun. And when you do it in a PM, you have all your efforts saved. DALL-E is pretty clunky in that you have to manually save an image or lose it once your history rolls off.
I have access to both and they're good for different things. DALL-E seems somewhat more likely to know what you mean. Midjourney seems better for making interesting fantasy and science fiction environments.
For comparison, I tried generating images of accordions. Midjourney doesn't really understand that an accordion has a bellows [1]. DALL-E manages to get the right shape much of the time, if you don't look too closely: [2], [3]. Neither of them knows the difference between piano and button accordions.
Neither of them can draw a piano keyboard accurately, but DALL-E is closer if you don't look too hard. (The black notes aren't in alternating groups of two and three.)
Neither of them understands text; text on a sign will be garbled. Google's Parti project can do this [4], but it's not available to the public.
I expect DALL-E will have many people sign up for occasional usage, because if you don't use it for a few months, the free credits will build up. But Midjourney's pricing seems better if you use it every day?
[1] https://www.reddit.com/r/Accordion/comments/uuwrbj/midjourne...
[2] https://www.reddit.com/r/Accordion/comments/vz9zxw/dalle_sor...
[3] https://www.reddit.com/r/Accordion/comments/w0677q/accordion...
It doesn't matter, OpenAI wins anyway as these companies will pour hundreds of thousands into generated images.
It seems that the NFT grift is about to be rebooted again, such that it isn't going to die that quickly. But still, eventually 90% of these JPEG NFTs will die anyway.
They will be replaced by DALL·E 2 for creating these illustrations, book covers, NFT variants, etc opening up the whole arena to anyone to do this themselves. All it takes is to describe what they want in text and less than a minute, the work is delivered as little as $15.
OpenAI still wins either way. If a crypto company goes to using DALL·E 2 to generate photorealistic NFTs, they won't stop them and they will take the money.
Art is already dirt cheap. People aren't buying NFTs for their content. This doesn't make it appreciably easier to con rubes.
https://twitter.com/nutanc/status/1549798460290764801?s=20&t...
>> And I just used it to create cover art for a book published in Amazon :)
Man... what a missed opportunity for Altman... he could have had a really good cryptocurrency/token with a healthy ecosystem and a creative based community if he didn't push this Worldcoin biometric harvesting BS had he just waited for this to release and coupled it with access to GPT.
This is the kind of thing that Web3 (a joke) was pushing for all along: revolutionary tech that the everyday person can understand with it's own token based ecosystem for access with full creative rights from the prompts.
I wonder if he stepped down from Open AI and put it in a figurehead as CEO could this still work?
> Why is using a token better than using money, in this case?
It would be better for OpenAI if it can monetize not just its subscription based model via a token to pay for overhead and for further R/D but also for it's ability to issue tokens it can freely exchange for utility on it's platform for exclusive access outside of it's capped $15 model and allow for pay as you go models for those who don't have access to it like myself as it's limited to 1 million users.
I don't want an account, and I think that type of gatekeeping wasn't cool during the gmail days either and I had early access back then too, but I'd still personally buy $100s of dollars worth of prompts right now since I think it is fascinating use of NLP and I'm just one of many missed opportunities and represent a lost userbase who just want access for specific projects. By doing this they can still retain the caps of useage on their platform and expand and contract them as they see fit without excluding others.
This in turn could justify the continual investment from the VC World into these projects (under the guise of web3) and allow them to scale into viable businesses and further expand the use of AI/ML into other creative spaces, which as a person studying AI and ML and a background in BTC, is what we all wanted to see instead of these aimless bubbles in things like Solana or yield farming via fake DeFi projects like Calesius that we've seen.
It would legitimize the use of a token for use of an ecosystem model outside of BTC, which to be honest doesn't really exist and has still a tarnished view with all these failed projects, while gaining reception amongst a greater audience since it's captivated so many since it's release.
What were you expecting?
It's great at photorealistic images like this: https://labs.openai.com/s/0MFuSC1AsZcwaafD3r0nuJTT, but it's intentionally lobotomized to be bad at faces, and often has an uncanny valley feel in general, like this: https://labs.openai.com/s/t1iBu9G6vRqkx5KLBGnIQDrp (never mind that it's also lobotomized to be unable to recognize characters in general). It's basically as close to perfect as an AI can be at generating dogs and cats though, but anything else will be "off" in some meaningful ways.
It has a particular sort of blurry, amateur oil painting digital art style it often tries to use for any colorful drawings, like this: https://labs.openai.com/s/EYsKUFR5GvooTSP5VjDuvii2 or this: https://labs.openai.com/s/xBAJm1J8hjidvnhjEosesMZL . You can see the exact problem in the second one with inpainting: it utterly fails at the "clean" digital art style, or drawing anything with any level of fine detail, or matching any sort of vector art or line art (e.g. anime/manga style) without loads of ugly, distracting visual artifacts. Even Craiyon and DALLE-mini outperform it on this. I've tried over 100 prompts to get stuff like that to generate and have not had a single prompt that is able to generate anything even remotely good in that style yet. It seems almost like it has a "resolution" of detail for non-photographic images, and any detail below a certain resolution just becomes a blobby, grainy brush stroke, e.g. this one: https://labs.openai.com/s/jtvRjiIZRsAU1ukofUvHiFhX , the "fairies" become vague colored blobs here. It can generate some pretty ok art in very specific styles, e.g. classical landscape paintings: https://labs.openai.com/s/6rY7AF7fWPb5wWiSH0rAG0Rm , but for anything other than this generic style it disappoints hard.
The other style it is ok at is garish corporate clip art, which is unremarkable and there's already more than enough clip art out there for the next 1000 years of our collective needs -- it is nevertheless somewhat annoying when it occasionally wastes a prompt generating that crap because you weren't specific that you wanted "good" images of the thing you were asking for.
The more I use DALLE-2 the more I just get depressed at how much wasted potential it has. It's incredibly obvious they trimmed a huge amount of quality data and sources from their databases for "safety" reasons, and this had huge effects on the actual quality of the outputs in all but the most mundane of prompts. I've got a bunch more examples of trying to get it to generate the kind of art I want (cute anime art, is that too much to ask for?) and watching it fail utterly every single time. The saddest part is when you can see it's got some incredible glimpse of inspiration or creative genius, but just doesn't have the ability to actually follow through with it.
OpenAI needs to open their damn eyes and realize that a brilliant AI with provocative, biased outputs is better than a lobotomized AI that can only generate advertiser-friendly content.
I’m eager for OpenAi to wake up and walk back on the clumsy corporate censorship, and/or for competitors to replicate the approach and improve upon the original magic without the “bias” obsession tacked on. Real challenge though “bias” may pose in some scenarios, perhaps a better way to address this would be at the training data stage rather than clumsily gluing on an opaque approach towards poorly implemented, idealist censorship lacking in depth (and perhaps arguably, also lacking sincerity).
All in all though I think I am underwhelmed mostly because my initial expectations were off, I am still a fan of DALL-E specifically and GPT3 in general. Now when is GPT4 coming out? :)
It has, honestly, completely blown me away beyond my wildest imagination of where this technology would be at today.
Will it do it "more accurately" as they claim? As in, if 90% of CEOs are male, then the odds of a CEO being male in a picture is 90%? Or less "accurately reflect the diversity of the world’s population" and show what they would like the real world to be like?
You're being quite charitable. It is much more likely that optics and virtue signaling is behind this addition.
I think you’re jumping to quickly to bad intentions. Injecting diversity of results is a sane thing to do, totally irrespective of politics.
If I search for "food", the reasonable result would be to get images that represent food according to its actual proportions of real life. E.g. if Pizza is the most common food at 10% prevalence, 10% of the images should be pizza.
That's not what OpenAI are doing.
They are introducing crafted biases to create images that deliberately misrepresent what the world looks like, and instead represent what they believe the world ought to look like.
--
You also need some reason why diversity "of this" is important but not diversity "of that". Why is diversity of race and sex so critical, but not diversity of age, height, disability? Should a search for "basketball player" yield 1/2 able-bodied people and 1/2 wheelchair basketball players? Why?
Then try to answer where you came up with the categories you do want depicted. Why are the races what they are? Should "basketball player" include half whites and half black people? Or maybe split in 3, white/black/Asian? Why not Australian Aborigines, native Americans, or Persians - so we can divide into 6? If you don't add Indian people to your list then, is that racist against them? How did you decide what must be represented, in what proportions, and what's okay to leave out?
A comical work around to so called "bias" (isn't the whole point of these models to encode some bias?). Here's some experimentation showing this.
https://twitter.com/rzhang88/status/1549472829304741888
As competitors with lower price points prop up, you'll see everyone ditch models with "anti bias" measures and take their $ somewhere else. Or maybe we'll get some real solution, that adds noise to the embeddings, and not some half assed workaround to the arbitrary rules that your resident AI Ethicist comes up with.
Would only work for positive biases where if they actually want to equalize it then it needs to be adding the opposite to negative biases.
To counteract the bias of their dataset they need to have someone sitting there actively thinking in bias to counteract the bias with anti-bias seasoning for every bias causing term. Feel bad for whatever person is tasked with that job.
Could always just fix your dataset, but who's got time and money to do that /s
If you used good input you'd expect an appropriate output, I don't know why manual intervention would be necessary unless it's for other purposes than stated. I suspect this is another case where "diversity" simply means "less whites".
"A photo of a group of soldiers from WW2 celebrating victory over nazi CEOs and plumbers".
Supplying the race and sex information seems to prevent new keywords from being injected. I see no problem with the system generating female CEOs when the gender information is omitted, unless you think there are?
My point is that it is pointless. If you want an image of a <race> <gender> person included, you can just specify it yourself.
I agree wholeheartedly. So what are we arguing about?
What we're seeing is that DALL-E has its own bias-balancing technique it uses to nullify the imbalances it knows exists in its training data. When you specify ambiguous queries it kicks into action, but if you wanted male white CEOs the system is happy to give it to you. I'm not sure where the problem is.
I set up a similar GPT prompt with a lot more power ("rewrite this vague input into a precise image description") and I find it much more creative and useful than DALLE2 is.
Slightly more than half of the pictures will be women.
That accurately represents the world's diversity. It won't accurately reflect the world's power balance but that doesn't seem to be their goal.
If you want to say "white male CEO" because you want results that support the existing paradigm it doesn't sound like they'll stop you. I can't imagine a more boring request.
Let's look at interesting questions:
If you ask for "victorian detective" are you going to get a bunch of Asians in deerstalker caps with pipes?
What about Jedi? A lot of the Jedi are blue and almost nobody on Earth is.
Are cartoon characters exempt from the racial algorithm? If I ask for a Smurf surfing on a pizza I don't think that making the Smurf Asian is going to be a comfortable image for any viewer.
What about ageism? 16% of the population is over sixty. Will a request for "superhero lifting a building" have an 16% chance of being old?
If I request a "bad driver peering over a steering wheel" am I still going to get an Asian 50% of the time? Are we ok with that?
I respect the team's effort to create an inclusive and inoffensive tool. I expect it's going to be hard going.
Wouldn't that result end up being like "inoffensive art" or "inoffensive comedy"?
Bland, boring and Corporate-PC.
There are others, like being clever, or being absurd, or being goofy, or being poignant, or being refreshing.
Of the good stuff, offensive humor is only a tiny slice.
it takes a special talent to please everybody
- Red Teaming Language Models with Language Models
I'm not sure how they would accomplish 100% accurate proportions anyway, or even why that would be desirable. If I don't specify any traits then I want to see a wide variety of people. That's a more useful product than one that just gives me one type of person over and over again because it thinks there are no female firefighters in the world.
For MidJourney I was painfully surprised to find that everything is done through chat messages on a Discord server.
I'm not a paid member, so I have to enter my prompts in public channels. It's extremely easy to lose your own prompts in the rapidly flowing stream of prompts going on. I can kind of see why they did it that way--maybe, if I squint really hard--to try to promote visibility and community interaction, but it's just not happening. It's hard enough to find my own images, say nothing about follow what someone else is doing. This is literally the worst user experience I have ever had with a piece of software.
There are dozens of channels. It's so spammy, doing it through Discord. It's constantly pinging new notifications and I have to go through and manually mute each and every one of the channels. Then they open a few dozen more. Rinse. Repeat.
I understand paid users can have their own channels to generate images, but I really don't see the point in paying for it when, even subtracting the firehose of prompts and images, it's still an objectively shitty interface to have to do everything through Discord chat messages.
I don't know who thought that discord would make a good GUI front end...
"To be copyrightable, a work must be fixed in a tangible form, must be of human origin, and must contain a minimal degree of creative expression"
So some employees there are aware of the impact that AI can have. Getting these DALL-E images copyrighted won't be trivial. I think it will be many years before the law is clarified.
I have an RTX 3080 and will likely be buying a 4090 when it comes out. Will I ever be able to generate these images locally, rather than having to use a paid service? I've done it with DALL-E Mini, but the images from that don't hold a candle to what DALL-E 2 produces.
Anyway, OpenAI is unlikely to release the model. The situation will like it is with GPT-3; however, it's also likely another team will attempt to duplicate OpenAI's work.
The same person is also at work on an open-source implementation of Google's Imagen which should be even better (and faster) than DALLE-2: https://github.com/lucidrains/imagen-pytorch.
This is possible because the original research papers behind DALLE-2 and Imagen were both publicly released.
Haha description of the company from Google
if you've got 60GB available to your GPU then maybe you can get close
I'm really curious if Apple's unified memory architecture is of benefit here, especially a few years from now if we can start getting 128/256GB of shared RAM on the SoC
I spent several tries yesterday to get this angle "from the ground up": https://labs.openai.com/s/mz8LiyvkI8KwD2luJ6MrS23m
I'd prefer an option to pay like 200 usd/year to use unlimited. And maybe have a price per use only in the API.
edit: this pricing model also makes it expensive to learn to use the tool.
They might even provide image generation at a loss to drive people to their platforms.
For a long while whenever Midjourney or DALLE-mini or the other models underperformed or failed to match a prompt the common refrain seemed to be "ah, but these are just the smaller version of the real impressive text2image models - surely they'd perform better on this prompt". Honestly, I don't think it performs dramatically better than DALLE-mini or Midjourney - in some cases I even think DALLE-mini outperforms it for whatever reason. Maybe because of filtering applied by OpenAI?
What difference there is seems to be a difference in quality on queries that work well, not a capability to tackle more complex queries. If you try a sentence involving lots of relationships between objects in the scene, DALLE will still generate a mishmash of those objects - it'll just look like a slightly higher quality mishmash than from DALLE-mini. And on queries that it does seem to handle well, there's almost always something off with the scene if you spend more than a moment inspecting it. I think this is why there's such a plethora of stylized and abstract imagery in the examples of DALLE's capabilities - humans are much more forgiving of flaws in those images.
I don't think artists should be afraid of being replaced by text2image models anytime soon. That said, I have gotten access to other large text2image models that claim to outperform DALLE on several metrics, and my experience matched with that claim - images were more detailed and handled relationships in the scene better than DALLE does. So there's clearly a lot of room for improvement left in the space.
In case you are interested in reading the whole take: https://aifuture.substack.com/p/the-ai-battle-rages-on
(1) Any opinions on if removing the watermark is possible? Is doing so against the terms of service?
(2) Appears the output is still at 1024x1024 - what are options to upscale the resolution, for example would OpenCV super resolution work?
Yep... The output is an issue, I'd like to pay if that was an upgrade.
Here’s more information on super resolution options beyond what Adobe already offers:
(1) List of options current options for super resolutions:
https://upscale.wiki/wiki/Different_Neural_Networks
(2) Older example of one way to benchmark:
https://docs.opencv.org/4.x/dc/d69/tutorial_dnn_superres_ben...
I can run DALL-E mini with the MEGA model with 12 GB of VRAM. What are the requirements for minDALL-E in terms of VRAM?
https://huggingface.co/spaces/dalle-mini/dalle-mini
Reminder that the OpenAI team claimed safety issues about releasing the weights. Now they’re charging, when the above link GPU time is being paid for by investor dollars. I guess sama must be hurting if he can only afford OpenAI credit packs for celebrities and his friends.
What about generating NFTs? It was explicitly prohibited during the previous period, now there is no notion of it. Without notion and rights for commercial use I think it's allowed but because it was an explicitly forbidden use case before, I want to be sure whether it can be used or not.
Regardless, excited to see what possibilities it opens.
The commercial use language appears pretty clear to me to allow NFTs. (But note the absence of any discussion of derivative works...)
I have been on the waitlist from the very beginning. Still waiting.
I have been on the waitlist for a while and did not get access yet.
Did anybody get access already today?
e.g. One is unable to create faces of real people in the public eye.
it's great for post-post-ironic memes, but I don't see it being useful for anything else
How did you score?
I only scored as well as I did because I knew the kind of stylistic choices to look out for. In terms of "quality" I really don't understand how you've reached this conclusion.
is it not dall-e?
It's a long way off in terms of quality (at the moment anyway)
But it does seem to know a lot of things the real DALLE2 doesn't.
That includes YOURS.
Have you been paid for it?
I suppose it's possible that at some point they'll try to make an image -> svg translation model?
Why do you think that is?
- Unlike software photography/illustration are not used to run essential systems like banking, manufacturing, medical equipment etc
What I do use is OpenAI’s GPT-3 APIs, I am a paying customer. Great tool!
Also, I love how, while signing up for access to this amazing AI, I am asked to, indeed, affirm I am _not_ a robot.
> "Do not attempt to create, upload, or share images that are not G-rated"
Turns out that they randomly, silently modify your prompt text to append words like "black male" or "female". See https://twitter.com/jd_pressman/status/1549523790060605440
I don't know which emotion I feel more - applause at how glorious this hack is or tears at how ugly it is.
Good luck to them!
A mortgage AI that calculates premiums for the public shouldn't bias against people with historically black names, for example.
This problem is harder to tackle because it is difficult to expose and resign the "latent space" that results in these biases; it's difficult to massage the ML algo's to identify and remove the pathways that result in this bias.
It's simply much easier to allow the robot to be bias/racist/reflective of "reality" (its training data), and add a filter / band-aid on top; which is what they've attempted.
when this is appropriate is the more cultured question; I don't think we should attempt to band-aid these models, but for more socially-critical things, it is definitely appropriate.
It's naive on either extreme: do we reject reality, and substitute or own? Or do we call our substitute reality, and hope the zeitgeist follows?
That's a great example, thanks. Also, I hope the teams working on that come up with a different solution...
My question to you is: is an algorithm that takes no racial inputs (name, race, address, etc) yet still produces disproportionate results biased or racist? I say no.
[1] https://files.stlouisfed.org/files/htdocs/publications/es/08...
The government, and many people, have moved the definition and goal posts; so that anything that has the end result of a non-proportional uniformity can be labeled and treated as bias.
Ultimately it is a nuanced game; is discriminating against certain clothing or hair-styles racist? Of course. Yet, neither of those are explicitly tied to one's skin color or ethnicity, but are an indirect associative trait because of culture.
In America, we have intentionally muddled the waters of demarcation between culture and race, and are starting to see the cost of that.
Not that I agree with that but I don't see why you would build one otherwise, if you wanted discrimination free mortgages wouldn't the whole process by anonymized and minimal personal information rather than the current system of having to hand over every detail of your life.
1. The training data would've been the best way to get organic results, the input is where it'd be necessary to have representative samples of populations.
2. If the reason the model needs to be manipulated to include more "diversity" is that there wasn't enough "diversity" in the training set then its likely the results will be lower quality
3. People should be free to manipulate the results how they wish, a base model without arbitrary manipulations of "diversity" would be the best starting point to allow users to get the appropriate results
4. A "diverse" group of people depends on a variety of different circumstances, if their method of increasing it is as naive as some of the are claiming this could result in absurdities when generating historical images or images relating to specific locations/cultures where things will be LESS representative
If I want to produce "animal images" but it only produces images of black cats, do you think there is any question whether it's a problem or not?
The "issue" is a different one: that training data - IE, reality, has _unwanted_ biases in it, because reality is biased.
Producing images of men when prompting for "trash collecting workers" should not be much of a surprise: 99% of garbage collection/refuse is handled by men. I doubt most will consider this a "problem," because of one's own bias, nobody cares about women being represented for a "shitty" job.
But ask for picture of CEOs, and then act surprised when most images are of white men? Only outrage, when proportionally, CEO's are, on average, white men.
The "problem" arises when we use these tools to make decisions and further affect society - it has the obvious issue of further entrenching stereotypical associations.
This is not that. Asking DALLE for a bunch of football players, would expectedly produce a huddled group of black men. No issue, because the NFL are disproportionately black men. No outrage, either.
Asking DALLE for a group of criminals, likewise, produces a group of black men. Outage! Except statistically, this is not a surprise, as a disproportionate amount of criminals are black men.
The "problem" is with reality being used as training data. The "problem" is with our reality, not the tooling.
Except in the cases where these toolings are being used to affect society - the obvious example being insurance ML algorithms. et al - we should strive to fix the issues present in reality, not hide them with handicapped training data, and malformed inputs.
I think, for about 95% of the world football is synonymous with soccer. Its kind of interesting that you take this particular example to represent what reality looks like statistically
This is not great. Only about 57% of NFL players are black, and the percentage is more like 47% among college players. It would be better to at least reflect the diversity of the field, even if you don't think it should be widened in the name of dispelling stereotypes.
> Asking DALLE for a group of criminals, likewise, produces a group of black men. Outage! Except statistically, this is not a surprise, as a disproportionate amount of criminals are black men.
Only about 1/3 of US prisoners are black. (Not quite the same as "criminals" but of course we don't always know who is committing crimes, only who is charged or convicted.) That's disproportionate to their population, but it's not even close to a majority. If DALLE were to exclusively or primarily return images of black men for "criminals", then it would be reinforcing a harmful stereotype that does not reflect reality.
"criminals" producing most black people actually would be a perfect example of bias in DALL-E that is arguably racism.
Black people commit a diproportionate amoumt of crime (for a variety of socioeconomic reasons I won't get into here), but even so white people make up a majority of criminals (because white people are the largest ethnic group by far).
Thus, a random group of criminals, if representive of reality, should be majority white.
It's a hard problem for sure. But remember, the bias ends with the user using the tool. If I want a black scientist, I can just say "black scientist".
Let me be mindful of the bias, until we have a generally intelligent system that can actually do it. I'm generally intelligent too, you know.
That is a really, really, narrow viewpoint. I think what people would prefer is that if you query "Scientist" that the images returned are as likely to be any combination of gender and race. It's not that a group is "fragile", it's that they have to specify race and gender at all, when that specificity is not part of the intention. It seems that they recognize that querying "Scientist" will predominantly skew a certain way, and they're trying in some way to unskew.
Or, perhaps, you'd rather that the query be really, really specific? like: "an adult human of any gender and any race and skin color dressed in a laboratory coat...", but I would much rather just say "a scientist" and have the system recognize that anyone can be a scientist.
And then if I need to be specific, then I would be happy to say "a black-haired scientist"
It's way bigger than just this narrow race issue the current zeitgeist is concerned about.
But I agree, maybe I should skew to being optimistic that at least we're trying
I wonder what the distribution of those modifications is?
And the best of all - it does have a meme community around it, and you can always donate if you feel it adds value to your life
I tried the first whimsical, benign thing I could think of: "indiana jones eating spaghetti." The results are clearly recognizable as that. But they are also a kaleidoscope of body horror; a Indiana Jones monster melted into Cthulu forms inhaling plates that are slightly not spaghetti.
I'm interested in using Dall-E commercially, but I think some competitor offering sampling with raw input will have a better chance at my wallet.
I don't understand the relevance of the black box's scrutability - I just want to play with the black box. I am interested in increasing my understanding of the black box, not of a trust-me-it's-great-our-intern-steve-made-it black box derivative.
These kinds of modifications are obviously different. At least the mathematical transformations are attempting at least some level of fidelity to user input, these ones aren't (e.g. someone mentioned they're sometimes getting androgynous results and speculates the added terms are conflicting with the ones they provided in their input). Not all black boxes are equivalent.
This feels like a very hacky way to essentially reinvent programming badly.
My bet is that in a few years or so only a small cohort of engineering and product people will even remember Dall-E and GTP-3 and someone cringe at how we all thought this was going to be a big thing in the space.
There's are both really fascinating novelties, but at the end of the day that's all they are.
Happily for me I stopped painting digitally long time ago. I even stopped calling myself "an artist". Nowadays I paint and draw only with real medium and call all of that "Archivist craftsmanship with analogue medium". :)
So DALL·E 2 is going to restart, revive and cause another renaissance of fully automated and mass generated NFTs, full of derivatives and remixing etc to pump up the crypto NFT hype squad?
Either way, OpenAI wins again as these crypto companies are going to pour tens of thousands of generated images to pump their NFT griftopia off of life support, reconfirming that it isn't going to die that easily.
Regardless of this possible revival attempt, 90% of these JPEG NFTs will eventually still die.
So I think it's prudent that OpenAI keep the 'sell shovels' business model instead with DALLE and GPT, at least for the time being.
I think it has not been trained on NFT art (crypto punks and so on).
How exactly are you defining NFT art?
I mean, it can literately be anything: Dorsey sold a screencap of his 1st tweet, Nadya from Pussy Riot did some creative stuff, and the Ape crap was the bulk of this stuff that got passed around.
I think what can be gleaned from that short-lived non-sense is that value is subjective and that the quality of a valuabe piece of 'art' is equally as hard to define. Much the same with its predecessor: cryptokitties.