Midjourney AI Guide
enchanting-trader-463.notion.site
enchanting-trader-463.notion.site
One recent feature the guide missed is the permutation and repeat features [1]. They're quite helpful for power users that want to explore multiple styles quickly.
Last week I tried putting together a short film using GPT-4 and Midjourney v5. I was stunned by the cinematic frames Midjourney v5 was able to create:
I (human) wrote the prompts for Midjourney, though.
I've gotta say, I've seen worse storytelling and cinematography come from actual, serious humans who were getting paid to do it. And the Balenciaga thing is obviously a joke, because that whole genre only works because the style of those videos is beyond parody and sails right over the uncanny valley. This is different. This is interesting. I like it.
I'm glad you pointed out the music. The music was also AI generated with a tool called AIVA [1]. I'd never composed a piece of music before, and I was pretty surprised by what I could "create". I spent 30~60 minutes max creating the score.
Some parts of their product still feel janky, but as an overall concept, it's quite fascinating. One of the interactions I enjoyed was that AIVA creates scores with different tracks (layers). So I was able to edit tracks I don't like (e.g., change a Piano track to Brass) or have AIVA completely regenerate certain sections of the score (e.g., redo the bridge, regenerate the chorus sections).
One difference from Midjourney is that there's no text-based prompting. Instead, you "prompt" through music inspiration.
So, basically, if I want it to compose Baroque music, I can give it Vivaldi, Bach, and maybe a little early classical, tell it to go to work, and end up with something that sounds like it came out in 1765?
I wonder what the limitations are on that whole "musical prompting" deal.
Your video was better in my opinion because it has a real story. All the Balenciaga videos out there are really just realistically rendered parodies with little or no emotion to them.
It's worth highlighting that, even in v5, the best results are obtained with fairly complicated prompts.
These results, for instance: https://imgur.com/RD7fF3M
Are what came up via: The interior of a future-brutalist megastructure. Single image comics style sci-fi robotic architecture, 70s, jet era Jetson cantilever modular space station modules, flag patch on walls, pastel glowing symbols, sci-fi towers architecture, foggy, pastel volumetric light and atmosphere, glass and steel pastel, The 5th element, The Matrix, Blade Runner, photography award, ultra elegant wide angle, volumetric light at mid day nice sky, trending on artstation, Unreal engine hyper realistic photography magazine cover, no people --v 5 --q 2 --upbeta
MJ does really well with simple prompts.
Specific words and phrases that accurately translate what you want into the latent space are way more important than complex prompts.
If you've played around with MJ much you'll know what I mean. Some words and phrases overpower everything else in the prompt. I've found that good MJ prompts are very vibe based.
Did we? Go grab a brush and some paint and create. The joy always come from the inside. If you can't seem to find it anymore inward is where you should look for it.
There are prompt books and prompt sites where you can buy them a dollar at a time. The hustle culture has built this fog of fake complexity and hidden methods in order to prop up their little cottage industry. Every AI Guy on TikTok has a way to "maximize your productivity with GPT" so you can "Start an AI company".
We haven't sucked the joy out of creating things, we've just tried to further commodify the process of creating and in turn spun a web of myths and lore where there didn't need to be any.
So... we've sucked the joy out of it.
Which should be obvious because a picture paints a thousand words…
Then there's the whole "+masterpiece, +best quality, -bad hands, -extra limbs" superstition that comes from Danbooru-based models, but people coming on board just assume it's necessary everywhere.
But some of the most thought-provoking results I get from Midjourney are with short, quasi-macaronic prompts [1]. I know enough botanical and zoological nomenclature to conjure fantastic lifeforms close to what's in my imagination, by contriving Latin binomials that just sound right. I can also give certain subjects the cultural flavor I want by making up nonsense words that sound like Swahili, French, Finnish or whatever, in situation where actually including the language name, or real words from the language, take the image in the wrong direction.
I had some GPU hours left to burn up the other day and even figured out how to make a short Midjourney animation with dozens of images in a smooth sequence. There are ways of very carefully stepping around local bits of the latent space.
[1] https://www.unite.ai/the-nonsense-language-that-could-subver...
Are you talking about stuff like the trick of adding "artstation," or "in the style of..." or "taken with a blah blah lens/camera?" That all sounds like it generally comes from bias in the training data set to me. It doesn't make it any less valid an observation or any less useful, but I wonder if A) that is the case, and B) if it's hinting at the idea that training these things on source data that was initially generated by humans is going to create some kind of limitation in what they can do.
For example, I tried some weird stonehenge prompts like "mcdonalds in stonehenge" and "stonehenge as mcdonalds store". Stonehenge completely overpowers McDonalds in generated images.
Another example, when I prompt with "squad of space marines" I get tabletop warhammer figurines. However, when I prompt with just "space marines" I get really nice art. It's easy to accidentally slip into a different part of the latent space.
The same thing tends to happen with particular adjectives and descriptors as well (E.g. colour, material, texture). This is a good thing if you can figure out words that translate your idea into the latent space well, but it can make controlling the output tricky sometimes.
I've tried using weights with ::, but that feels like a bandaid fix most of the time.
One thing I've noticed is that sometimes, the exact same prompt will generate very different images. Once, I got an image of a blue sports car, off to the side of the road, a kind of craftsman-ish looking interior of a house, and two pictures of some random outdoor forest type area. I don't remember the prompt I used, unfortunately, because I ended up writing it off as a glitch, but it would be interesting to see if anybody else has had it happen.
> words that translate your idea into the latent space well,
So, this phrase gave me a sort of a thought: what if certain words or phrases, like, say, "McDonald's" and "Stonehenge," are just so far apart in the model space (again, likely due to biases in the training set, or just the fact that there aren't many McDonald's' restaurants at Stonehenge), that the more interesting or common or unusual one serves as a kind of attractor and dominates the generation process most of the time?
Do you know if these effects are documented anywhere in the prompt engineering literature?
Simple prompts can definitely get good results, but they tend to be much more generic than what you can get with complex prompts.
Not sure if this is just my prompts or something general.
https://docs.google.com/spreadsheets/d/1cm6239gw1XvvDMRtazV6...
checking the V5 tab I almost immediately found Akira Toriyama, and the art is ... well, sure on one of the four drawings there are some Goku-like hair on ...some cheeks (not the face ones) drawn in the style of Terry Gilliam?
Then of course, for older artists the match is much stronger, and for some of the modern and famous ones too. But it seems to me that the whole method could be heavily biased.
ControlNet for Stable Diffusion may be an exception to this.
Both StabilityAI and the open source community are working on improvements to Stable Diffusion.
Keep in mind StabilityAI is also pursuing LLMs and the host of other model types, whereas text to image is Midjourney's single core competency and value prop. Midjourney is very focused on staying ahead.
edit: I wanted to add that the extensive training costs can be prohibitive for the OSS community to fully participate. Coordination via groups such as LAION can help, but gone are the days of individual OSS participants contributing directly to core foundational model training.
Stable Diffusion's advantage is in the huge amount of open source activity around it. Most recently that resulted in ControlNet, which is far more powerful than anything Midjourney can currently do - if you know how to use it.
Ambreen Butt
Jan Cox
Constance Gordon-Cumming
Dai Xi
Jessie Alexandra Dick
Dong Qichang
Dong Yuan
Willy Finch
Constance Gordon-Cumming
Spencer Gore
Ernő Grünbaum
Guo Xi
Elena Guro
Adolf Hitler
Prince Hoare
William Hoare
Fanny McIan
Willy Bo Richardson
Shang Xi
Wang Duo
Wang E
Wang Fu
Wang Guxiang
Wang Hui
Wang Jian
Wang Lü
Wang Meng
Wang Mian
Wang Shimin
Wang Shishen
Victor Wang
Wang Wei
Wang Wu
Wang Ximeng
Wang Yi
Wang Yuan
Wang Yuanqi
Wang Zhenpeng
Wang Zhongyu
Xi Gang
Xie Shichen
Xu Xi
I have this list because I recently made a site[1] that displays the 4 images from a prompt of "Lotus, in the style of <paintername> <birth-death dates> [nation of origin]" for every painter listed on wikipedia's "List of painters" -- except for those in the above list.
The fact that they banned both Xi and Jinping separately to prevent Xi Jinping was surprising to me. Twice as banned as Adolf Hitler.
[1] https://lotuslotuslotus.com - small chance you get an NSFW image if you hit upon Fernando Botero or John Armstrong, perhaps there's more.
Anyway, I suspect with more competition many of the restrictions on sites like Midjourney will eventually be removed.
Nope, that doesn't stop models from being trained on your art. It makes it somewhat more difficult for people to prompt specifically for your style, but your art still influences output, and there may be other ways (e.g., titles of specific works) to deliberately and specifically evoke it in particular.
I already had a discord bot I wrote by hand before for downloading the images.
I thought of the project in the morning while my kid was getting ready for school and had it running jobs before we were out the door, worked a little before I left for work, and and a little more after work, and it was done before dinner time.
ChatGPT is incredible.
Maybe that just means banning images that are meant to fool people, and obvious satire would be ok, but they might be ok erring on the side of caution.
(Nothing in the talk was this explicit, but this was my read of the subtext)
The main difference with Stable Diffusion is that you can fine tune with your own dataset. There's img2img, and a bunch of other tools. But the base model it's really worse than competitors right now.
I started testing out the official Stable diffusion API and it already gives you way more control than the Dall-e API and seems to produce less horrifying images that are better quality but I feel like Dall-e understands the prompts better. 7/10
I would love to try mid journey but I uninstalled discord years ago and have no plans to ever reinstall it ever again. So I’ll wait for the API access if they ever do it. 0/10 (only for being discord only)
The base tech and models straight from Stability AI give pretty crap results if you just plainly describe a scene.
MJ in contrast provides a great results out of the box. Say anything, get beautiful picture of something. From there you need to figure out specifically what you actually want.
However, if you really want to iterate on just specific details of a scene with a specific layout, with specific characters in specific poses, with specific style elements, MJ is too chaotic to control at that fine detail. So to is SD out of the box. But, if you take the time to learn how to install and use ControlNet, highly specialized models/LORA/textual inversions from wildly varying sources, in-painting, latent upscaling, hook up Photoshop/Krita/Blender integrations, ect, ect… you can eventually get very precise control of SD’s results. And, then new better tech releases next week! :D
Stable Diffusion does more things though, like in-painting where you can erase part of an image and then have it recreated. I've seen videos of people doing impressive things with in-painting and extensive regeneration of each portion of an image until it's just right. Seems like a ton of work though. Still I've had some fun using it to modify images or extend images.
I'd be happy to be corrected if anyone knows the details on this in SD.
Its better in terms of having a no-user-configuration service available that gets you from zero to decent results with nothing more than prompting.
SD is better in available specialized customized models (finetuned checkpoints and the various kinds of mix-and-match auxiliary models and embeddings that can be used with them), not having banned topics because it is self-hostable, available tooling and UIs that expose tuning parameters and incorporate support for integrating techniques like guided generation with various types of ControlNet models, animation, inpainting, outpainting, prompting-by-region, etc., with image generation.
Especially when the only interface is via Discord! I continue to refuse to join Discord. It's an archivists nightmare.
I hate this hostile, intrusive UX pattern. I verified my email, I never use VPN, and my browser's ad blockers were disabled. Yet they claim "security issue". Discord can shove their dishonesty.
And yes I tried one of those free number websites but didn't work because millions of other people use those same numbers.
I don't get anything like this from other services on the same device.
I hate to be one of those guys who bandies the term "UX" around, but yeah, that's terrible UX.
I'm willing to put up with it for v5, though. My local install of SD just doesn't measure up, at least not at present.
Bing image creator is also really good and free with a hundred generations per day
In short when they did user testing with a Web UI, people would write a basic prompt "dog", "cat" and then get stuck.
Discord makes it social by default, you can see what people are prompting and get inspiration.
I also found it a bit weird, but now that ive onboarded its great having a multi platform native experience (ie even chatgpt still has no mobile apps).
The social angle is not what's fueling Midjourney's success. The quality of its output certainly is.
Are there SD applications that have UX as good as that?
But I’m not saying one is better than the other, I think they each have pros/cons, just wanted to list numbers for the curious.
I’m guessing you are referring to archive.org style public archiving.
But, I’ll mention anyway that for your own generations there is a web interface to search and bulk download your images. And, a web interface is in the works. You lose out on the social aspect of prompting with friends. But, gain some UX for working solo. Or, you can work solo in a DM convo with the bot.
My experience has been exactly the opposite.
I created my own Discord "server" (which is not truly a server per se; more like your server account [1]), and then installed MJ bot. This way I can keep all my creations in a single place, with the added benefit that I have my own channels for things like compilation of links, prompts database, separate channels for different types of creations, and so on.
I even invited a couple of family members to join my Discord server, so they can use MJ if they want, without having to share my Discord account (but of course it still counts towards my MJ credits).
All in all, I have been incredibly (and surprisingly) pleased with Discord experience.
There must be a way of creating a Discord bot wrapper which holds your credentials, but on a quick search I wasn't able to find anything.
What a big bummer. That would be the ideal environment.
I found the most interesting results can be found by
a) combining artists i.e.: by Rafal Olbinski in the style of frank frazetta
b) combining artists and photographers i.e.: by Rafal Olbinski photographed by Helmut Newton
c) turning artists into photographers i.e.: photographed by Frank Frazetta
d) limits limits limit i.e.: only in orange and yellow
some of the things I did can be found https://instagram.com/f_r_a_n_z_ai?igshid=YmMyMTA2M2Y= here (but most in the vanishing Stories as I did not but them into squares yet)
> "If you're good at prompt engineering"
Now that's funny. Promptbase.com thinks writing silly prompts is engineering!
Prompt phrases like "photography award" are silly. May as well add "best image in the universe" or "really really really good" etc.
Calling it "prompt engineering" is laughable. Selling prompts even more laughable.
In addition to being able to create completely novel, high quality images, they can do so at speeds that no human could ever hope to match. And image models of this quality have only been around for a year, imagine where they'll be in five.
I empathize will all the artists who feel cheated out of these models training on all their data, but the sad reality is these AI models are just far too useful to ever go away. The world's standards for art and text have gone up faster than they ever have in world history over the course of the last 10 months
This basically means you get thousands of different nostalgic triggering visuals but it's a visual trick the same way marvel movies all follow the same premise and eventually satiate the space.
Given the cost saving potential of AI art generation, I suspect it'll only be a matter of time before it becomes mainstream in the west.
Instead of having an "us" against "them" mentality, there might be a future where human artists can work alongside AI to push human creativity to a new level.
Firstly, this ignorantly trivializes the value, even commercially, of artistic talent (hint, tool usage pretty a small part of it.) The most important ingredient and biggest time sink in bigger high-level projects is the artistic minds that decide what goes on the screen to begin with... how they're made is an implementation detail. The images these algorithms pump out are dazzling to amateurs, but it's not close to precise, reliable or consistent enough to make content for these projects. People who do this stuff at a high level know that this tech will be relagated to supporting tools-- tools that copy one element manually made to many different contexts, mood board or story board panel generators, photo or compositing filters, etc. for the foreseeable future. Insisting otherwise is like saying github copilot is imminently about to replace developers, as if coding is the only important thing that developers contribute to software projects. Sure, it will eventually replace a lot of lower-end utility developers because the higher level developers will be so much more productive, but that's a very different thing. Speaking of that...
Secondly, in the market for the higher-volume, lower artistic effort commodity commercial art, tool mastery IS the big selling point, and that market will take a giant hit. These are people with mortgages, kids approaching college age, maybe carrying for sick relatives or relying on employer-sponsored health care for insulin. And it's not like you're getting laid off and can get another job– your entire category of employment is toast. Someone flamboyantly dancing on your grave in public and smugly telling you to find a new profession is pretty fucking good reason to get defensive or offended.
There's also a "terrain" in the latent space, with valleys of popularity and mountain ranges of novel intersections. Generated images can flip from one valley to the other with a small change in prompt, and it's hard to stay in the intersection. Fine tuning helps alter this terrain - that's mostly what fine tuning does. Rather than teaching genuinely new things, it mostly warps and tweaks what's already there to make it easier to target, at the cost of making everything else a tiny bit harder to get to. Teaching the model new things requires including lots of regularization images to ensure it doesn't forget all the other stuff, which is more expensive than fine tuning or dreambooth.
I've also found that some words and phrases will overpower the rest of the prompt, so it can be hard to get to specific areas of the latent space.
I think they're really good for mood board kind of work and exploring ideas, but you'll probably still need an artist to create something that is specifically what you want.
The UI/UX of a proper workflow is still very janky (even in user-friendly UIs like auto1111), but you can get very specific, complicated scenes with a mixture of ControlNet + inpainting.
> I've also found that some words and phrases will overpower the rest of the prompt, so it can be hard to get to specific areas of the latent space.
I've also had this problem (particularly with multiple colored objects, like "blue eyes, brown hair"). Apparently Cutoff (https://github.com/hnmr293/sd-webui-cutoff) is very good at addressing this "leakage", but I haven't implemented it yet.
Right now, things are really good for mood board kind of work, but there's a lot of tech already available for enabling a lot of back-and-forth "work" on images to get them into exactly what you're imagining. There aren't a lot of good UIs for them all yet though; as soon as they get a little more user-friendly, I'd expect another boom in generative AI for artists (this time, with artists being the primary benefactor).
One shot random prompt stuff is better on Midjourney and I'm still subscribed to it, but I would say I do 99% of my playing around inside StableDiffusion these days.
Using a simple drawing of single colour shapes as the original image and a good prompt, one can get quite far in terms of composition.
The desert dome city still blows my mind every time I think about it.
SD with Multidiffusion or another multiple prompts with (potentially overlapping or nested) regions tool, plus some inpainting, can do a lot here.
So can SD with ControlNet, to go beyond text prompting.
This is a funny sentiment because to me, the exact opposite has occurred. Now that everyone can create work of comparable technical ability, never has it been more clear that taste and other conceptual skills are lacking in those who haven’t spent hours upon hours engaged with and making art.
I’m sure many will decry this as snobbery, but I think this is also obvious in the reliance on various famous artists names in prompts. I expect a similar phenomenon will emerge with AI generated code where things that look superficially impressive are terrible in other dimensions (architecture, efficiency, security, etc)
We got jet engines in airplanes only 25 years after the first flight. Imagine where we'll be 100 years after!
(There's plenty of progress to be made in gen AI, but don't expect the speed of progress to be constant)
-90s tech journalists, amateur web 'designers', and other non-coders who assumed they understood software
It’s going to make humans hella productive.
For artists that means less staring at blank screens and fishing for inspiration.
The future is that anyone will be able to whip up anything without needing that particular skill that is a blocker. Instead they can pull from collective intelligence and piece something new together.
If an artist wants to make games but can’t code they can make games like a boss in the future because they can funnel their productivity into all the things they were never capable of before.
It’s literally infinite possibilities. Like I’m a software engineer. I’m sure these AI systems will be boss tier at a lot of stuff I know.
That’s good. That means I can finally stop doing the things I hate the most when it comes to making software and do the fun part: solve problems. Just like you use a calculator and formulas to solve problems and not doing every calculation by hand from scratch..
Anyway just my 2 cent. All this doomer talk is being pushed by people to get people afraid. Honestly if you’re at the forefront of technology you should be striving to shape the future for better where everyone can make hella money.
Imagine how much stuff there will be to sell to everyone, and everyone can channel their true potential and creativity into building cool stuff. You like cars and want to make body mods? Can’t code? Don’t know cad? No worries someone has set up a service which helps you easily get those molds delivered to you. And if that’s your thing you can put that workflow together and make money.
Am I crazy for thinking that way?
the main second order effect is that the middle class of knowledge workers ceases to exist
and the third order effect is that democracy no longer works and we either turn into saudi arabia or feudalistic medieval europe
I'd rather GPUs were banned
Excited for what V6 will bring.
sorry aspiring prompt engineers, AI can do that too
My experience with LLaMA has shown me this is definitely not the case.