Drawing a childrens story with DALL-E 2
twitter.com
twitter.com
I had to mess around with the prompt a little but GPT-3 just did this:
Prompt:
Hi i'm Elsie, a first grader going to school for the first time.
This is a six panel comic speaking in the first person about how my day went.
I'll include a suggestion for the illustrator to draw each frame in square brackets:
Generated text: [Elsie is standing in front of her house with her backpack on, looking a little nervous.]
Hi, I'm Elsie. Today is my first day of school. I'm a little nervous, but I'm excited too.
[Elsie is walking to school with her mother.]
I walk to school with my mom. We talk about what my day might be like.
[Elsie is standing in front of her school with her mother, who is giving her a hug and kiss goodbye.]
My mom gives me a hug and a kiss, and then I'm off!
[Elsie is in her classroom, sitting at her desk.]
I sit at my desk and the teacher starts talking. I learn a lot of new things today.
[Elsie is standing in front of her house again, this time with a big smile on her face.]
I had a great first day of school! I can't wait to go back tomorrow.
Any time i mess with this stuff i get lost trying different variations. Here's a few more (# marks the prompt) https://pastebin.com/deC0RxDbhttps://lithub.com/will-this-bionic-font-help-you-read-faste...
Curious if that helps
1. https://www.psychnewsdaily.com/new-study-shows-all-caps-hard...
That said i’m not writing books.
Not quite ready to replace comic artists.
https://docs.google.com/presentation/d/e/2PACX-1vT4XWNx2SdEg...
Lesson learned is that if you don’t prompt it with a premise it will probably copy it from something in the training data. In this case the words are novel but the theme comes from a book of the same name.
Is it possible to tell DALL-E to use the same girl in each panel? That is the only detail preventing me from being convinced this isn't a real children's book.
So you can generate a picture say 'picture of a boy playing with a ball' then go into edit mode (think this is new) and erase ball, then change prompt to 'picture of a boy playing with toy car'.
It will keep the original elements of boy not erased and put in a toy car. The character is kept the same but doesn't work as well as you might think as the body, face pose are the same.
Still, you can see this getting better over time.
Only got access this morning, hard not be impressed.
As a comparison, stock photos and video of people are often available as a scene showing the same folks in a bunch of scenarios and angles.
You kind of need this to build something of any complexity out of the art.
Denoising diffusions have been successfully applied to produce coherent videos.
In fact, there are even some benefits to train on video because images share information between them, which make context extraction easier.
Counterintuitively, because of the shared information between frames, their computations can also be partially shared, it can be almost free to generate multiple frames at the same time : A pair of RGB image can be considered as a 6D image, a 5-frame sequence a 15D image, and the neural network can learn to compress the various image channel together. The subsequent neural network layers can be kept 256D (for example) the same size as they were for a single frame.
Of course this trick won't work if the pair of image are too different but you can always stack images side by side to have a bigger film-roll image and you use a standard attention mechanism to attend to various parts.
https://old.reddit.com/r/dalle2/comments/vfi8lj/homer_simpso...
For original characters, it's tricky, you have to use a simple design that can be describe in text, in this example it's the prominent red hoodie and dark pants that make it look like the same character across pictures:
https://www.youtube.com/watch?v=J_laffNOoQw
Another problem is that DALL·E has a hard time telling different characters apart, e.g. requesting Captain America and Iron Man in the same image will give you two characters that have attributes of both Captain America and Iron Man mixed together:
https://39669.cdn.cke-cs.com/rQvD3VnunXZu34m86e5f/images/a39...
If you wanted to work within the current feature set, I think inpainting or 'uncropping' is probably the way to go. For example, generate a character until you get one that you like; generate a bunch of variants of that one face in various positions/angles/emotions (either through the variation feature or by reprompting); now, cut the faces or figures out of each one, and use them for each illustration as inpainting or uncrop with a text prompt describing that scene.
So you'd generate your target girl, get a bunch of her faces (uncertain, sad, facing left, facing right), cut them out for an outline, and then inpaint/uncrop: "Shawna felt sad to be away from home and Mom all day"+[sad-face.jpg]+"a brown-haired young white girl in a gingham dress standing alone in a colorful cozy kindergarten classroom"; and so on for each scene.
The user describes a character to the character model and gets candidate illustrations. The user can then save a character by giving it a name and use the name to generate variations of the same character. I'm imagining that either you just have a lot of sliders to vary the vector that the character is drawn from or maybe it's possible to combine ideas (e.g. Character X + running).
Then the user just illustrates the scenes they want, the characters, combines them together, and gets a useful image.
it is an alpha version of the future hence good content for hacker news, not necessarily something that should go into a real children's book ;)
"Girl with brown hair, small nose, blue eyes, white shirt and blue skirt goes inside school" for example.
Just need to train it on that girl somehow?
This is the infinite universes thing right?
Besides that, this is pretty good - certainly a few things that don't seem quite ideal (the fifth frame doesn't seem to capture the meaning of the text especially well), but enough to make you think that in not too long, it is fully plausible that DALL-E will be fully able to act as an illustrator for this use case. On the one hand, pretty exciting. On the other, certainly harrowing for children's book illustrators.
I wonder how long it will be before I can use this for my marketing emails. I sell dog treats, and I have a fairly simple template with an image at the top. That's almost always a photo of my products and/or my dogs. How long before I can just ask for an image from DALL-E for something like "Dog sitting next to grill in back yard with American flags and other patriotic decorations" for the top of my Fourth of July email? How long before I can feed it a picture of my products and have it generate photorealistic images of dogs eating them, thus replacing the photographer I use for product shoots? It's an exciting prospect for me as a small business owner - lots of time and money saved in an area where I'm not an expert - but definitely pretty scary for people who create visual imagery of any kind.
Not sure it really matters, half of all children books are filled of incoherent stories and imagery, but I think it's mostly adults who notice that.
1. OpenAI claim ownership of the content produced.
2. The content policy forbids what you're describing. (commercial use is ruled out)
Content policy: https://labs.openai.com/policies/content-policy
Sharing & publication policy: https://openai.com/api/policies/sharing-publication/
I'm not 100% sure, but I don't think it has limitations on how you're allowed to use it?
I suspect it will turn out to be much more work for the artist to find what it is you actually liked about the generated output and rework the pictures based on that, rather than the usual process where the client describes what they need and the artist uses their experience and understanding of the human mind to decide how to represent that.
I guess I’m saying that the art in “art” is not the superficial skill of making pretty pictures, but the concurrently honed skill of making meaningful choices. I’m not sure you can just bypass learning one without compromising the other.
It can already generate several independent images containing known characters that are in the training data (Obama, Pikachu, etc.)
In this case it seems that you could fine-tune the generative image model with previous frames to provide that continuity. After all that's what we do when we read the panel, we instantly store the previous one in memory so that we can actually recognize the difference in the next panel.
(e.g. why does the little girl keep changing hair, face, clothes, age, ....)
For example, an incoherent story. I think that will piss off most people.
I'm reading your comment as saying that there is some sine qua non about art that professional driving lacks - but I think that may just be your bias as an artist. If software produces results that are as good or better than yours, what is lost by replacing artists with software?
I'm also not writing this just to dig at professional artists and drivers. I think my career too is in the process of being replaced by software. I have doubts and misgivings about whether this will be entirely good, but it seems clear that it is happening and that most (all?) professions are or will soon be in a similar state.
computers are useless, they can only provide answers
Children's book illustrators, though, really? How do we benefit if we make it impossible for that to be a living? Is there any benefit to you or your children if they look at pictures generated from a machine learning algorithm based on past images instead of by a human artist that was paid for their creation? Is there any benefit to artists? Is there any benefit to society as a whole, or to any individuals except the owners of the algorithms generating the pictures and the publishers who save a relatively minor [1] cost?
Personally I doubt ML-illustrated children's books would catch on as anything more than a novelty simply because generating illustrations in this way seems so dystopian, impersonal and tacky, but the way people react to new technology is something I find very hard to predict.
[1] Typically between £6,500 and £8,500 per book according to https://www.peopleofpublishing.com/post/how-much-does-a-chil...
Imagine taking the cost of illustrating a children's book from 10k USD (plus who knows how much time) to paying a 10 dollar per month fee to create unlimited illustrations? What would that enable? How many more children's books would get produced? How much more happiness for children and would be authors?
My mother used to tell us stories about characters she had made up. With a little more time, if this technology existed, she could have illustrated her stories and shown them to us. Or, if this technology gets invented tomorrow, I could create generate the illustrations from what I remember of my mom's stories and share it with her. Why wouldn't I want this? Because professional illustrators would be devalued?
Let's not forget the children who lack interesting content or concerned parents who would create it for them. These kids could benefit from a reddit or imgur of children's books and get an infinite scroll of content. If software becomes good enough the infinite scroll could be generated automatically.
I predict DALL-E art to be on the level of those freaky auto-generated kids' videos for a long time. People see those as successful too. Millions of Views with very little Cost.
That's the perfect word for this, because we have to ask what this would look like if it metastisized.
Algorithm illustrates the book -> Algorithm writes the book -> Algorithm generates the audio to read the book to your kids for you -> Algorithm mines trends and engagement stats to decide which books it will write.
It's the total annihilation of literature, and all the way more creative people get cut out and the bottom line pads the pockets of the publishers who own it. Humanities without humanity, i.e. nothing.