Run your own DALL-E-like image generator
reticulated.net
reticulated.net
Here's the code:
from diffusers import DiffusionPipeline
ldm = DiffusionPipeline.from_pretrained(
"CompVis/ldm-text2im-large-256"
)
output = ldm(
["a painting of a raccoon reading a book"],
num_inference_steps=50,
eta=0.3,
guidance_scale=6
)
output["sample"][0].save("image.png")
This took 5m25s to run on my laptop (no GPU configured) and produced a recognisable image of a Raccoon reading a book. I tweeted the resulting image here: https://twitter.com/simonw/status/1550143524179288064AI art notes in general:
* this will permanently raise the standard of "programmer art" in prototype UIs, for example.
* I expect large-scale AI image creation social games will take off very soon
* utilities to listen to your conversations and illustrate people's points with poetic interpretations on screens next to you with captions.
* this will/should be integrated with emoji / emoji kitchen-like systems - imagine iterating on a custom emoji-like response in snapchat!
* trademark/copyright wars over input data may ruin thisBut I would also suggest [2] which is the one used in the article already deployed for free
[1] https://discord.gg/midjourney
[2] https://huggingface.co/spaces/multimodalart/latentdiffusion
Disclaimer: I subscribed and when I came to my senses 5 hours had passed and I had a gallery of weird pictures...
For example, this one [1] works very well.
Google Colab comes with a fair amount of GPU time for free.
[1]: https://colab.research.google.com/drive/1TBo4saFn1BCSfgXsmRE...
The ”insane amount of memory” turned out to be 32 GB or more. I thought we are talking about 512 GB or something like that.
128GB = 4x 32GB DDR4 UDIMM costs less than 360€.
That's a lot less than the GPU.
Yeah they're technically cheaper than a GPU, but your costs are still gonna be up and they can't replace a GPU because regardless you'll need an nvidia card, unless there's a workaround for CUDA requirements.
Sigh, when will I have a nightmare generator of my own?
This makes me hella curious about the type of stuff you were requesting.
Of all of the celebrities, it seemed reticent to draw Fred Rogers as, well, Fred Rogers, despite me being able to make some fairly appalling melds of others, like Rod Serling and Marilyn Monroe.
Additionally, someone has put in an overly-touchy anti-vore filter.
I also tried uploading a film still, and it said I may not use realistic faces. So I used another AI tool to remove the face. It still wouldn't budge.
https://labs.openai.com/policies/content-policy
There are a number of other rules that can only be discovered by combing their Discord or running into them headfirst.
It puzzled me for a while, but then it made sense:
"A hairy ball with lights at the end is manipulated by a human"
Not a great prompt, but it seemed like a good start. I imagined fiber optic strings attached to a ball and a human doing some stuff around it like a wizard or something. But then I realized that "hairy ball" might lead to something completely different :)
Okay, no problem, I changed it to "tumbleweed" but that triggered the drug filter so I gave up and tried something else :)
They give you the seeds, that are used to populate the images, so you can fine tweak an image by adjusting slight details in a prompt. It is really awesome.
There are languages ecosystem that solved these issues and never have a problem with it. Look at it, and copy their tooling.
On the other hands, I lost a lot of time with npm, and also a lot with pip.
It's a shame, I think the field is one where reproduction of results should be really welcome and feasible.
Yeah, a docker image would be nice. Even a Dockerfile, even though that in itself may not guarantee reproducibility if you try to build it later. (And may have issues with gpu drivers etc). But at least it documents all assumptions about the setup.
Someone sets it up, leases out instances etc..
Or someone works to ensure that everything gels together and provides support etc..
Or in other words: 'products'
I got access to DALL-E a few weeks back (after a waitlist… come on) and tried it out. They want all this personal info to even access it, and then they have this oppressive “content policy” that removes anything remotely fun.
For example I tried “Trump riding a velociraptor on mars fighting aliens” because why not? Sounds hilarious. Turns out any query with the word Trump is banned and I got a warning about “repeated violations might remove my access”. It’s not just trump, anything remotely non-corporate-friendly is heavily filtered. Don’t you want to just make images of a cute teddy bear made of pizza instead?! I’m just so tired of it all. It’s like the puritans of the 1990’s won, but they’re corporations now.
Personally I think they can shove that nonsense up their ass.
They just want to avoid bad PR. That’s sound safe.
https://labs.openai.com/policies/content-policy
Highlights include:
-All images must be "G-rated"
>Any prompt referencing politics, violence, sexuality, or even the concept of "health" is banned
>This does occasionally include the very same LGBT themes their "hate" rule ostensibly protects, although not consistently, so there's no way to know whether an LGBT prompt will cause an account strike
>In practice the list of bannable offenses is much longer than this, and includes everything from the concept of death to anything violence or politics-adjacent (e.g. nothing about war, conflict, or any kind of weapon)
>OpenAI refuses to share this list because it's part of a "contextual" filter, even though it demonstrably bans words
They advertise DALL-E 2 as artistic, but it can't make art with restrictions like this. The best it can do is corporate content farming, and even then there's no way I'd make part of my marketing pipeline depend on a service this fickle.
How gracious of them…
Come on man
Pepe is banned. Why? It doesn't matter. Someone will cut around, and they'll get my compute instead. And I'll get my autogenerated pepes, regardless of whether some non productive AI "ethicist" thinks a cartoon frog meme is offensive.
Sometimes even one is a problem.
I don't think I've talked to a career ML engineer yet that has had a positive work experience with conda, that I can recall. Usually the question has been a good way to trauma bond, which is always good.
It seems like this is more an academic kind of thing?
Other people's infrastructure always seems bad to me, as I'm sure mine does to others. The cycle and circle of life! :D
You’re implying it isn’t like that for every language/framework. I’d love to know what you’re working in where the dependencies aren’t always a problem. Even Carmack has complained about this.
Of course, ML is a bit different in that it often needs more drivers/gpu setup. But there are loads of ecosystems that don't assume you have package X installed on your system, like python's does.
Tensorflow JS running on GPU/CUDA in Node.
Sane? Absolutely not.
Though I suppose it's not really Node's fault that developers are importing modules like "leftpad".
What's really fucking insane though is that there are modules like "trim-newlines" [0] that exist merely to trim \r and \n from the beginning and end of a string...and that this is such a hard task to get right that it's in version 4.0.2...and that a previous version had a security vulnerability [1].
(I'm a Patreon supporter but otherwise have no connection)
However, it would be very cool if deep learning tools weren't dependent on CUDA and could be run on, say, any GPU with OpenCL.
I'm now tied to CUDA, but I didn't need to be. I was starting from scratch.
The longer before AMD can ship something which actually works, the more entrenched NVidia+CUDA become.
At this point, I lost hope.
- I found out I could only use it for compute headless. WTF?!?!?! (https://www.phoronix.com/news/Radeon-ROCm-Non-GUI). If it was driving a monitor, my machine would crash hard. There wasn't even an error message.
- A lot of other stuff didn't work and just resulted in odd crashes, or worse performance than CPU. I don't know why.
- Within 9 months, AMD discontinued support for my card. I raised this as a warranty issue (suitability for advertised purpose), but that obviously would go nowhere without a lawsuit. I had a very expensive brick.
- AMD support channels were non-existent. There literally was no way to reach anyone.
I bought an NVidia card, and it's been working well ever since.
ROCm is not well-supported because it's absolute garbage. You have *less* work reinventing wheels, since it's been invented once. You have more work to get community support and network effects, since you're starting out behind. Fundamentally, though, that can't start to happen if your system doesn't work at all.
I agree with you they're trying, but they're trying incompetently.
If ROCm was half the speed of CUDA, and wasn't integrating into the latest-greatest frameworks, but it was stable and working, I'd make it work. It wasn't anywhere close to stable and working.
Who is going to give away their trained model (that presumably cost millions) to the public? And getting the blame for any nefarious usage that will eventually follow?
What I find strange is that StabilityAI doesn't outright ban using actual living people in prompts (where I can imagine wide usecases for abuse). At the same time both OpenAI and StabilityAI block sexual content/nudity/violence. I don't see why such content is supposed to be harmful.
Anyway looking forward to the model, thanks for informing me about this project.
(Caveat: I have updated neither repo since about November because it took a faff to get them working and I do not want to touch them again.)
It just looks surprisingly like its mixing and matching the top returned image searches from an index.
I'm only saying it looks like that - not that it is ofcourse. I don't want to undermine anyones works here, I was just wondering.
Most networks have to do with large language models, text embeddings, and image embeddings.
One good quote from that era: "You should be no more concerned that a computer can beat you in chess than that a car can beat you in a race."
Technology changes the experience of being human. Chess used to be the marker of human intelligence; that passed. Now some forms of creativity will too.
The originals look like malformed mutants at 256x256 with nonsensical key bits and strange twisted handles.
Note that this does not replace a good pixel art artist either - for things like walls, anything that needs to connect together or tile like a dungeon wall or castle, or a cohesive art style, this will not do. But for rapidly identifiable different quest item drops it's not terrible.
"Within the next year or two, you’ll be able to make content in real time: 30 frames a second, high resolution. It’ll be expensive, but it’ll be possible. Then, in 10 years, you’ll be able to buy an Xbox with a giant AI processor, and all the games are dreams.".
As someone who used to take quite a lot of psychedelics, there's something quite terrifying about the promise of this premise - it takes me back to the wrong sort of trips where the ever-unfolding strata of reality became too much bear and I'd end up mentally cowering beneath the unrelenting bigness of it all.
The infinite dreamspace is unholy-big and nt somewhere I'd much choose to get lost.
Or maybe I totally will...
- ed - Notwithstanding the obvious realisation that I could just take the helmet off, or remove the contact lenses, or whatever we have in a couple of decades.
You will be able to visualize your own nightmares when you are awake.
I don't feel the need to check for reality. I've seen screenshots from games that look more realistic than some pictures I've taken (Forza on max settings at the right angle might as well be a photograph) and this trend will only continue.
This tool can be considered more of an automated version of meme communities, where dedicated members will spend an hour photoshopping muppets into historic events and making the pictures look absolutely believable. The only novelty I see is that the computer now does a lot (but not all) work for you.
There are nice opportunities here. If you need a stock photo of something very specific, you'll soon be able to generate one with the right query and the right AI. Small companies can generate fancy brands without paying professional designer's fees, especially if all they need is a billboard and not a whole suite of office supplies. You can generate your own posters and decorations featuring interesting landscapes and scenes in any style you want.
The current iterations of these algorithms are quite limited in many aspects and sometimes uncanny or even horrifying, but I look forward to a future where I can imagine something, describe it, and have it rendered into digital art just like I pictured, without having to spend decades on honing my skills as an artist.
In the case of DALL-E and GPT-3, I believe they will undermine human creativity and usher in a new era where fewer people care about slowly crafted skills like painting that require years or decades of practice and patience to master if the barrier to just have the art in front of you in seconds becomes so low.
It might get to the point that people will divide themselves on ideological grounds that AI art is "impure" and attack each other if they cannot prove its origin. I'm not saying that I'm one of those people that would join the pushback, given that a coming explosion in AI art is all but inevitable, but I'm describing what I think the new technology is going to cause the masses to believe - the population that aren't enthusiastic tech evangelists. When the expectations of millions of people are set by DALL-E and the like, I don't think the prospects will be universally positive.
Look at the number of people on HN asking if certain comments were written by GPT-3; they seem to appear weekly. I don't think implying that you didn't actually write the comment you posted will be taken very well by some people outside of an insular circle like HN once the general public becomes fully aware of AI, and could very well grow into a well-known insult if there ever comes to be a rift in opinions around AI art.