Comparing Adobe Firefly, Dalle-2, and OpenJourney
blog.usmanity.com
blog.usmanity.com
1st prompt: https://i.postimg.cc/T3nZ9bQy/1st.png
2nd prompt: https://i.postimg.cc/XNFm3dSs/2nd.png
3rd prompt: https://i.postimg.cc/c1bCyqWR/3rd.png
This is using some of the popular prompts you can find on sites like prompthero that show amazing examples.
It’s been serious expectation vs. reality disappointment for me and so I just pay the MidJourney or DALL-E fees.
In a nutshell:
1. Use a good checkpoint. Vanilla stable diffusion is relatively bad. There are plenty of good ones on civitai. Here's mine: https://civitai.com/models/94176
2. Use a good negative prompt with good textual inversions. (e.g. "ng_deepnegative_v1_75t", "verybadimagenegative_v1.3", etc.; you can download those from civitai too) Even if you have a good checkpoint this is essential to get good results.
3. Use a better sampling method instead of the default one. (e.g. I like to use "DPM++ SDE Karras")
There are more tricks to get even better output (e.g. controlnet is amazing), but these are the basics.
I learned this mostly by experimenting + browsing civitai and seeing what works + googling as I go + watching a few tutorials on YouTube (e.g. inpainting or controlnet can be tricky as there are a lot of options and it's not really obvious how/when to use them, so it's nice to actually watch someone else use them effectively).
I don't really have any particular place I could recommend to discuss this stuff, but I suppose /r/StableDiffusion/ on Reddit is decent.
I'd like to have a go at making one myself targeted towards single objects (be it car,spaceship, dinner plate, apple, octopus, etc). Most checkpoints are very heavily leaning towards people and portraits.
People generally suggest 30+ images. I’ve found - at least with people - the more the better. My wife’s model is trained on ~80 images of her.
You can finetune it on your own material, or choose one of the hundreds of public finetuned models. You can guide it in a precise manner with a sketch or by extracting a pose from a photo using controlnets or any other method. You can influence the colors. You can explicitly separate prompt parts so the tokens don't leak into each other. You can use it as a photobashing tool with a plugin to popular image editing software. Things like ComfyUI enable extremely complicated pipelines as well. etc etc etc
Honestly? Probably YouTube tutorials.
I'm going to sound like an entitled whiny old guy shouting at clouds, but - what the hell; with all the knowledge being either locked and churned on Discord, or released in form of YouTube videos with no transcript and extremely low content density - how is anyone with a job supposed to keep up with this? Or is that a new form of gatekeeping - if you can't afford to burn a lot of time and attention as if in some kind of Proof of Work scheme, you're not allowed to play with the newest toys?
I mean, Discord I can sort of get - chit-chatting and shitposting is easier than writing articles or maintaining wikis, and it kind of grows organically from there. But YouTube? Surely making a video takes 10-100x the effort and cost, compared to writing an article with some screenshots, while also being 10x more costly to consume (in terms of wasted time and strained attention). How does that even work?
Best just to dive in if you're interested IMO. Otherwise you'll get lost in all the new jargon and ideas. Great place to start is the A1111 repo, lot of community resources available and batteries included.
I can't see how covering this on YouTube, instead of (vs. in addition to) writing text + some screenshots and diagrams, makes any kind of sense.
This is the level we're generally working at - first or second party to the authors of the research papers illustrating implementations of concepts, struggling with the Gradio interface, things going straight from commit to production.
It's way less frustrating to follow all of the authors in the citations of the projects you're interested in than wasting your attention sorting through blogspam, SEO, and YT trash just to find out they don't really understand anything, either.
This is where video demonstrations come in handy. Since many concepts are novel, it's uncommon to find anyone who deeply understands them, but it's very easy to find people who have picked up on some tricks of the interfaces, which they're happy to click through. I think gradio/automatic1111 makes learning harder than it needs to be by hiding what it's doing behind its UI, while eg- comfyui has a higher initial learning curve but provides a more representational view of process and pipelines.
For all the promise of control and customization SD boasts, Midjourney beats it hands down in sheer quality. There's a reason like 99% of ai art comic creators stick to Midjourney despite the control handicap.
Neither of the existing models gives actually passable production-quality results, be it MJ or SD or whatever else. It will be quite some time until they get out of the uncanny valley.
> There's a reason like 99% of ai art comic creators stick to Midjourney
They aren't. MJ is mostly used by people without experience, think a journalist who needs a picture for an article. Which is great and it's what makes them good money.
As a matter of fact (I work with artists), for all the surface-visible hate AI art gets in the artist community, many actual artists are using it more and more to automate certain mundane parts of their job to save time, and this is not MJ or Dall-E.
And vice versa, which is the exciting part to me - only a matter of time!
If you’re ok with basic aesthetics it’ll work but if you want something a bit less cringe or that will stand out in marketing it won’t cut it.
Default Midjourney is one thing and that’s mid…
Opposite of what ? OP posts results from a tuned model.
>For all the promise of control and customization SD boasts, Midjourney beats it hands down in sheer quality.
The results are comparable, but MJ in this comment https://news.ycombinator.com/item?id=36409043 hallucinates more (look at the roofs in the second picture). And it cannot be fixed, maybe except for an upscale making it a bit more coherent. Until MJ obtains better tooling (which it might in the next iteration), it won't be as powerful. I'm not even starting on complex compositions, which it simply cannot do.
>OP posts results from a tuned model.
Yes, which is the first step you should do with SD, as it's a much smaller and less capable model.
If you want the power, it’s there. But nearly bone stock SD in auto1111 is going to get to any of these examples easily.
Show me the civitai equivalent for MJ or Dalle2. It doesn’t exist.
Ok...? Read what i wrote carefully. Your 6 sliders won't produce better images than midjourney for your prompt on the base SD model.
Also I see nothing wrong with using different models for different purposes.
But yes SD can be a bit of a pain to use. Think of it like this. SD = Linux, Midjourney = Windows/MacOS. SD is more powerful and user controllable but that also means it has a steeper learning curve.
We could just as easily say "hosting your own email can be set up in a few minutes if you know what you're doing". I could do that, but I couldn't get local SD to generate comparable images if my life depended on it.
screenshot of the options interface: https://stash.cass.xyz/drawthings-1687292611.png
Here, I've uploaded it to civitai: https://civitai.com/models/94176
There are plenty of other good models too though.
1. Start with a good base model(s) from which to train from.
2. Have a lot of diverse images.
3. Ideally train for only one epoch. (Having a lot of images helps here.)
4. If you get bad results lower the learning rate and try again.
5. After training try to mix your finetuned model with the original one, in steps of 10%, generate X/Y plot of it, pick the best result.
6. Repeat this process as long as you're getting an improvement.
For training I mostly used scripts from here: https://github.com/bmaltais/kohya_ss
The main problem here is that essentially during inference you're using a bag of tricks to make the output better (e.g. good negative embeddings), but when training you don't. (And I'm not entirely sure how you'd actually integrate those into the training process; might be possible, but I didn't want to spend too much time on it.) So your fine tuning as-is might improve the output of the model when no tricks are used, but it can also regress it when the tricks are used. Which I why I did the "mix and pick the best one" step.
But, again, I'm not an expert at this and just did this for fun. Ultimately there might be better ways to do it.
3. Train for only 1 epoch - interesting, any known rationale here?
5. I just read somewhere else that someone got good results from mixing their custom model with the original (60/40 in their case) - good to hear some more anecdotes that this is pretty effective. Especially the further training after merging, sounds promising!
I've also been using kohya_ss for training LoRAs so great to hear it works for you for models as well. On your point about the inference tricks, definitely noted but I did notice that you can feed some params (# of samples, negative embeddings, etc) to the sample images generated during training (check the textarea placeholder text). Still not going to have all usual the tricks but it'll get you a little closer.
update: I've edited the post to include these results as well
But launching it and then just letting it stagnate indefinitely and get worse every day compared to its increasingly popular competitors seems like the worst of all worlds, and I can't see what is the OA strategy there.
The way I see it, they don't need txt2img at this moment - GPT-4 ensures they're the top #1 name both in the industry and in AI-related news stories. But it doesn't mean they won't come back to it. Couple observations:
- OpenAI isn't a "release early, release often" shop. They might be already working on something, but they'll release it only when it is a qualitative improvement over everyone else (or at least Dall-E).
- A bunch of hobbyists is doing all their work for free anyway. Stable Diffusion itself may not be SOTA, but the totality of hundreds of different fine-tunes on Civitai very much is. With all those models being shared in the open and relatively easy/cheap to recreate, it would make sense for OpenAI to just stand by and watch, and only invest resources once hobbyists hit a plateau.
- Looking at those Civitai models, it seems to me that OpenAI could beat txt2img SOTA easily, at any moment, by taking (or re-creating, depending on the license) the best five to ten SD derivatives, and put them behind GPT-4, or even GPT-3.5, fine-tuned to 1) choose the best SD derivative for user's prompt, and 2) transform user's prompt to set of parameters (positive & negative prompts, diffuser algo, numeric params) crafted with choice from 1) in mind. It's a black box. On the Internet, no one can tell you're an ensemble model.
- They could even be doing it as we speak - addition of function calls is aligned with this direction, fine-tuning for good prompt generation is mostly a txt2txt exercise, and again, hobbyists around the world are busy building a high-quality human-curated data set of {what I want}x{model + positive prompt + negative prompt + diffuser + other params} -> {is this any good?}. If I were them, I'd just mine this and not say anything.
- Overall, I think that in txt2img space, currently the hard part isn't the "img" part, but the "txt" part. OpenAI has a huge advantage here, and as long as its true, they're in position to instantly overtake everyone else in this space. That is, they have an "Ultimate attack" charged and ready, and are patiently waiting for a good moment to trigger it.
- Didn't they hint that GPT-4 successor will be multimodal? That could end up being their comeback to txt2img. And img2txt. And a bunch of other modalities.
EDIT: As if on cue, the very thing I was speculating about above is being discussed wrt. LLMs right now:
- https://news.ycombinator.com/item?id=36413296 - GPT-4 is 8 GPTs in a trench coat
- https://news.ycombinator.com/item?id=36413768 - 3-4 orders of magnitude efficiency (size vs effect) improvement in code generation, if your training data isn't garbage
And in both threads, people bring up older papers and discuss the merits of combining smaller specialized models into a more generic whole.
1. Yes, they are. Look at the constant iterative rollouts of GPTs 2. Most of which is useless to them, not that they have made any use of it 3. the fact that it would be so easy to improve, and they haven't, only emphasizes my point. 4. sure, that could be useful. Except there's zero integration or mention. (They haven't even opened up the vision part of GPT-4 yet.) 5. the fact that it would be so easy to improve, and they haven't, only emphasizes my point. 6. why wait for GPT-5 possibly years from now?
So they're "on the list". So whenever journalists and bloggers write articles about text2image, they're listed as a player in this space. For vast majority of such articles, neither the authors nor the audience will be able to tell that OpenAI's offering is far behind and that they're basically keeping a token presence in the space.
At least that's my hypothesis. I'm neither a domain expert or a business expert - I just feel that, for OpenAI, having laymen view them as an industry leader in AI in general, is worth the price of keeping Dall-E available. In fact, as more and more users realize there are better models available elsewhere, that price goes down, while the effect on laymen audience stays the same.
(Note: the term "laymen", as I use it here, specifically includes most entrepreneurs, managers and investors, in tech or otherwise. If I'm being honest in myself, I belong to that category too; it's in fact this conversation and some recent threads that made me realize just how weak OpenAI is in image generation space.)
> Look at the constant iterative rollouts of GPTs
You mean some unannounced ones, or the pinned models? Because AFAIK GPT-3.5 had two updates after release (the turbo model and the current one), and GPT-4 had one. I mean public releases; for example, how often they updated GPT-4 back before it was public, e.g. when Microsoft was building Bing Chat, is not relevant in this context.
Also compare that with how, going by HN submissions alone, every other day someone releases some improved LLaMA-derived LLM.
> 2. Most of which is useless to them, not that they have made any use of it 3. the fact that it would be so easy to improve, and they haven't, only emphasizes my point. 4. sure, that could be useful. Except there's zero integration or mention. (...) 5. the fact that it would be so easy to improve, and they haven't, only emphasizes my point.
There's little for them to gain by openly using all that work now. At the moment, they can just keep an eye on what's posted to Civitai, paying particular attention to how different model derivatives respond to prompts (think e.g. CyberRealistic vs. Deliberate) and why, and build up a training corpus of prompts and settings, helpfully provided by the community, complete with quality rating. They can do that using a small fraction of resources they have available - so that when the time comes, they can use their full resources to quickly train and deploy a model that blows everyone else out of the water.
Also, as an organization, they can focus only on so many things at a time. GPT-4 is buying them some space, and I believe they're currently focusing primarily on their cooperation with Microsoft, and/or other things involving LLMs. Given the relative usefulness and potential of LLMs vs. image generation, both short and long-term, doing more than bare minimum in image generation right now might be too much of a distraction for an organization this size.
> (They haven't even opened up the vision part of GPT-4 yet.)
They're in the lead. They're not in a hurry. They're likely giving Microsoft a head start.
> 6. why wait for GPT-5 possibly years from now?
Why do it earlier? What could they possibly gain by jumping back into text2image space now? At this point, compared to LLMs, text2image seems neither profitable not particularly relevant for x-risk, so whichever way you cut it, I can't see why would they want to prioritize it.
Which is fair enough, when you are a (relatively) small company competing with the likes of Google and Meta you really need to focus.
I would still easily put both ahead of the base models. You won't match the quality of those models without finetuning. When you do fine-tune, it'll be for a particular aesthetic and you won't match them in terms of prompt understanding and adherence.
It's like each of these has a hidden giant pile of negative prompts, or additional positive prompts, that greatly narrow down the range of output. There are contexts where the Dall-E 'spoopy haunted house ooooo!' imagery would be exactly right… like 'show me halloweeny stock art'.
That haunted house prompt didn't explicitly SAY 'oh, also make it look like it's a photo out of a movie and make it look fantastic'. But something in the more 'competitive' AIs knew to go for that. So if you wanted to go for the spoopy cheesey 'collective unconscious' imagery, would you have to force the more sophisticated AIs to go against their hidden requirements?
Mind you if you added 'halloween postcard from out of a cheesey old store' and suddenly the other ones were doing that vibe six times better, I'd immediately concede they were in fact that much smarter. I've seen that before, too, in different Stable Diffusion models. I'm just saying that the consistency of output in the 'smarter' ones can also represent a thumb on the scale.
They've got to compete by looking sophisticated, so the 'Greg Rutkowskification' effect will kick in: you show off by picking a flashy style to depict rather than going for something equally valid, but less commercial.
Can't wait to have something like StableDiffusion but for LLMs.
If stable diffusion didn’t launch Dall-e 2 would have been still valuable.
If I create a Mickey Mouse using photoshop would adobe be liable for it?
Sure likely more machines run Linux on servers but that’s like saying your body has more bacteria than your own cells. Technically correct but actually bullshit.
Regarding image generation in Photoshop I can confirm two things:
- It is excellent for in and out painting with a few exceptions*
- It remains poor for generating a brand new image
*Photoshop's generative fill is very good at extending landscapes, it will match lighting and according to the release video can be smart enough to observe what a reflection should contain even if that is not specifically included in the image (in their launch demo they showed how a reflection pool captured the underside of a vehicle.)
Where generative fill falls apart: Inserting new objects that are not well defined produces problems. Choosing something like a VW Beetle will produce a good result as it is well defined, choosing something like "boat", "dragon", or even "pirate's chest": will produce a range of images that do not necessarily fit the scene - this is likely because source imagery for such objects is likely vague and prone to different representations.
1st note about Firefly: Anything that is likely to produce a spherical looking shape tends to be blocked - likely because it resembles certain human anatomy. This is problematic when doing small touch ups such as fixing fingers.
A special note about photoshop versus other systems: Photoshop has the added problem of needing to match the resolution of the source material. Currently it achieves this from combining upscaling with resizing - this means that if one is extending an area with high detail, that detail cannot be maintained and instead is softer/blurrier than the original sections. It also means that if one extends directly from the border of an image, then a feathered edge becomes visible which must be corrected by hand.
I currently test the following AI generators, feel free to ask me about any of these: StableDiffusion (Automatic and InvokeAI), OpenAI's Dall-E 2, MidJourney, Stability AI's DreamStudio, and Adobe Firefly.
Seems clear to me that Midjourney has by far the best "vibes" understanding. Most models get the items right but not the lighting. Firefly seems focused on realism which makes sense for a photography audience.
https://twitter.com/fanahova/status/1639325389955952640?s=46...
If you want to play around with OpenJourney (or any other fine-tuned StableDiffusion model). I made my own UI with a free tier at https://happyaccidents.ai/.
It supports all open-sourced fine-tuned models & loras and I recently added ControlNet.
The flaw with these comparisons is that you really shouldn't use the same prompt with different generators. If you want to get best results you do have to play with the prompts and do a bunch of iteration to kind of explore the latent space and find what you're looking for. The first super long prompt looks like it's tuned for stable diffusion for instance. Different generators also have different syntax (e.g. with stable diffusion you can surround a phrase with parens to give it extra emphasis).
Generally, this model is much better than Dall-E 2, and it beats Firefly in some areas (I didn't try Midjourney or Stable Diffusion). Firefly usually produces photos with significantly fewer visual mistakes (like the wrong number of fingers or messed up faces) than the Bing Dall-E. But the latter usually understands prompts much better and more often produces something that matches it well. Firefly also doesn't "know" a lot of pop culture or history things, e.g. Marilyn Monroe, or what Coca-Cola is.
I used stable-diffusion-xl-beta-v2-2-2 model, copypasted prompts from the blog post, one-shot for each prompt. I chose style presets that closely matched the prompt (added as suffixes in image filenames).
Literally all of the examples have floor to ceiling windows across the entire length of the wall…
Generating a “word bubble” is going to look terrible in every major diffusion model. Cohesive words and writing in image models is still highly specialised.
Most image AI tools are terrible with words.
I am curious, what images did you try generating with midjourney?
And the inevitable booby cheesy rendered forest fairy.
I don't think they're terrible at all. They absolutely can make original art with decent production values.
They can't write text yet, but I'm sure that's coming soon.