Midjourney web experience is now open to everyone
midjourney.com
midjourney.com
Well good thing we have Flux out in the open now, both midjourney releasing web version or ideogram releasing there 2.0 on the same day after 2 weeks of flux won't redeem them as much. Flux Dev is amazing, check what SD community is doing with it on https://www.reddit.com/r/StableDiffusion/ . It can do fine tuning, there are Loras now, even control net. It can gen casual photos like no other tool out there, you won't be able to tell they are AI without looking way too deep.
https://fastflux.ai/ for instant image gen using Schnell (but its fixed on 4 steps and is mainly a tech show off of inference engine by runware.ai)
https://www.segmind.com/ has API support with lots of options, I am using it to generate and set wallpaper using an AHK script. It's very very slow though.
https://replicate.com/black-forest-labs/flux-schnell/example...
https://huggingface.co/spaces/black-forest-labs/FLUX.1-schne...
https://getimg.ai/text-to-image
There are other tools now if you Google 'Flux image generator online'
I think there is a line somewhere between 4-bit to 8-bit that will hurt performance (for both diffusion models and LLM). But I doubt the line is between 8-bit to 13-bit.
(Another case in point: you can use generic lossless compression to get model weights from 13bit down to 11bit by just zip exponent and mantissa separately, that suggests the effective bit rate is lower than 13bit on full-precision model).
But yes, I do believe that we will find proper lossless quants, and eventually (for real this time) get "only a little bit of loss" quants, but I don't think that the current 8 bits are there yet.
Also, quantized models often have worse GPU utilization which harms tokens/s if you have the hardware capable to run the unquantized types. It seems to depend on the quant. SD models seem to get faster when quantized, but LLMs are often slower. Very weird.
If we start from peering into quantization, we can show it is by definition lossy, unless every term had no significant bits past the quantization amount.
so our lower bound must that 0.03% error mentioned above.
However this seems to be model size dependent, ex. Llama 3.1 405B is reported to degrade much quicker under quantization
Flux's power isn't necessarily in its ability to produce realistic images, so much as its increased size gives it a FAR superior ability to more closely follow the prompt.
Also see https://civitai.com/models/652699/amateur-photography-flux-d...
and
Or as someone once said, if they took porn off the internet, there'd only be one website left, and it'd be called "Bring Back the Porn!".
What happened with 1.5 and newer? I’m out of the loop.
I liked ideogram's approach. It looked like their training data was not censored as much (it still didn't render privates). The generated images are checked are tested for nudity before presenting to user, if found, image is replaced by a placeholder.
If you check SD reddit, community seems to have jumped ship to flux. It has great prompt adherence. You don't have to employ silly tricks to get what you want. Ideogram also has amazing prompt adherence.
Flux is also censored, but in a way that doesn't usually break anatomy. It took about a week before people figured out how to start decensoring it.
Crazy this took them so long, and also crazy that they got so far through a very confusing Discord experience.
Presuming your bot requires people to interact with it in public channels, your CSRs can then just sit in those channels watching people using the bot, and step in if they’re struggling. It’s a far more impactful way to leverage support staff than sticking a support interface on your website and hoping people will reach out through it.
It’s actually akin to the benefit of a physically-situated product demo experience, e.g. the product tables at an Apple Store.
And, also like an Apple Store, customers can also watch what one-another are doing/attempting, and so can both learn from one another, and act as informal support for one-another.
I don't use Discord on my office laptop, and that was very odd experience
That's like the killer use case for image generating AIs
Can you give an example? Midjourney is heavily censored so it seems like it has a lot of restrictions.
When you went to the Discord, you immediately got the endless stream of great looking stuff other people were generating. This was quite powerful way to show what is possible.
I bet one challenge for getting new users engaged with these tools is that you try to generate something, get bad results, get disappointed and never come back.
No hosting of the generated pictures, just send them via discord message and forget them. No S3 or big cloud lambda functions.
Easy to start to make a minimal working prototype.
What a strangely deranged view of the world some people have.
It is deranged that requiring a phone to access a website is seen as a non-issue, I agree.
I have about 40+ google accounts for reasons, so I don't begin to understand the aversion some have to registering a burner google/discord/facebook/etc account under their hampster's name, but many of my closest friends are just like you so I respect it anyway, whatever principle it is.
Maybe because several of those tend to ask for phone verification nowadays, and phone burner services tend to either not work or look so shady it looks like a major gamble to give them any payment information?
I tried their new web experience, and... it's just broken. It doesn't work. There's a showcase of other people's work, and that's it. I can't click the text-box, it's greyed out. It says "subscribe to start creating", but there's is no "subscribe" button!
Mindblowing.
https://storage.googleapis.com/dream-machines-output/378f0f1...
Why should Google's or Discord's policy dictate my participation in the web?
* You have to save your (hopefully unique!) email/password in a password manager which is effectively contradictory to your "I won't use a cloud service" argument.
* The company needs to build out a whole email/password authentication flow, including forgetting your password, resetting your password, hints, rate limiting, etc etc, all things that Google/Apple have entire dedicated engineering teams tackling; alternatively, there are solid drop-in OAuth libraries for every major language out there.
* Most people do not want to manage passwords and so take the absolute lazy route of reusing their passwords across all kinds of services. This winds up with breached accounts because Joe Smith decided to use his LinkedIn email/password on Midjourney.
* People have multiple email addresses and as a result wind up forgetting which email address/password they used for a given site.
Auth is the number one customer service problem on almost any service out there. When you look at the sheer number of tickets, auth failures and handholding always dominate time spent helping customers, and it isn't close. If Midjourney alienates 1 potential customer out of 100, but the other 99 have an easier sign-in experience and don't have to worry about any of the above, that is an absolute win.
Effectively you mean that people have multiple Google accounts?
Especially since those companies can wield this enormous power by removing my access to this service because I may or may not have violated a policy unrelated to this service.
There has to be a better way.
I want to believe, but sadly there's no market for it. unless someone wants to start a privacy minded alternative to auth0, and figure out a business model that works , which is to say, are you willing to pay for this better way? are there enough other people willing to pay a company for privacy to make it a lucrative worthwhile business? because users are trained to think that software should be free-as-in-beer but unfortunately, developing software is expensive and those costs have to be recouped somehow. people say they want to pay, but revealed preferences are they don't.
I’m very not impressed by this deep, extended critique of machine learning researchers using common security best practices on the grounds that those practices involve an imperfect user experience for those requiring perfect anonymity…
Password managers don't have to be cloud services. The user gets to choose.
Some of us like to use a different email address for each account on purpose.
It is a bit of extra work but that's just how it is nowadays.
There's too much unnecessary connected PII data generated by such mechanisms.
Plus, you may not mind not being anonymous to Midjourney but mind not being anonymous to some other service (like Google).
If that even still works...
Because that's what you're asking for, as it'd be trivial to reset the rate without accounts.
i assume they still have ip/browser fingerprint based rate limiting, but you can type in a prompt and get an answer with zero login or anything else.
Most of the current AI companies are going to fail or get bought up, so you have no idea where that account information, all your prompts and all the answers will eventually go. After social media, I don't really see any reason why anyone would trust any tech company with any sort of information, unless you really really have to. If I want to run an LLM, then I'll get one that can run on my own computer, even if that mean forgoing certain features.
https://gondolaprime.pw/pictures/sporks-generative.jpg
It's a rather arbitrary metric. All it proves is that the token "spork" wasn't in midjourney's training data. A FAR better test for adherence is how well a model does when a detailed description for an untrained concept is provided as the prompt.
I'd say a better test for adherence is how well a model does when the detailed description falls in between two very well known concepts - it's kinda like those pictures from the 1500s of exotic animals seen by explorers drawn by people using only the field notes after a long voyage back.
Prompt: A hybrid kitchen utensil that looks like a spoon with small tine prongs of a fork made of metal, realistic photo
https://gondolaprime.pw/pictures/flux-spork-like.jpg
The combination of T5 / clip coupled with a much larger model means there's less need to rely on custom LoRAs for unfamiliar concepts which is awesome.
EDIT: If you've got the GPU for it, I'd recommend downloading a copy of the latest version of the SD-WEBUI Forge repo along with the DEV checkpoint of Flux (not schnell). It's super impressive and I get an iteration speed of roughly 15 seconds per 1024x1024 image.
- Descriptive -> describing a difficult concept that is most certainly NOT in the training data
- Hybrids -> fusions of familiar concepts
- Platonic overrides -> this is my phrase for attempting to see how well you can OVERRIDE very emphasized training data. For example, a zebra with horizontal stripes.
etc. etc.
( Prompting: An iron-gated door, set into a light stone arch, all deep set into the side of a gentle hill, as if the entrance to a forgotten crypt. The hill is lightly wooded, there is foliage in season. It is early evening. --v 6.1 )
And result: https://cdn.midjourney.com/5b56f713-3d64-471f-8c3c-08a0247e6...
The style matches exactly what I'd want too, it's captured "Fantasy RP book illustration" extremely well, despite that not being in the prompt!
Midjourney’s request does not comply with Google’s ‘Use secure browsers’ policy. If this app has a website, you can open a web browser and try signing in from there. If you are attempting to access a wireless network, please follow these instructions.
You can also contact the developer to let them know that their app must comply with Google’s ‘Use secure browsers’ policy.
Learn more about this error
If you are a developer of Midjourney, see error details.
Error 403: disallowed_useragent
Unable to process request due to missing initial state. This may happen if browser sessionStorage is inaccessible or accidentally cleared. Some specific scenarios are - 1) Using IDP-Initiated SAML SSO. 2) Using signInWithRedirect in a storage-partitioned browser environment.
EDIT:
I tried again from scratch in a new tab and this time it worked. So, temporary hiccup.EDIT2: I have all the images I created on Discord in the web app - very nice!
Same feeling you get looking at the airbrushed art on a state fairground ride.
That's a great way to describe it. A lot of articles and youtube pics are using these images lately and they all give that sort of vibe.
The docs say to hover over the uploaded image or use the test tube icon from the sidebar, neither of which seem to be available on mobile.
Everybody can now be an artist and have creativity for free now.
What a time to be alive.
This is a net positive for the world.
First attempt:
https://cdn.midjourney.com/686f364f-188f-4065-813f-b45d21972...
Tweaking to specify spokes around the central axis:
https://cdn.midjourney.com/d18d3ad3-3fd0-461e-a3a8-5fc46f0ba...
I think this is the most popular one right now online for Flux.
https://cdn.midjourney.com/ed234782-ac5c-41d7-99e9-20ce122cd...
They definitely struggle with getting the land moving up the wall and often it treats the cylinder like a window onto earth at ground level but I think with tweaking the weighting or order of terms and enough rolls of the dice you could get it.
"Create a highly detailed image of the interior of a massive cylindrical space habitat, rotating along its axis to generate artificial gravity. The habitat's inner surface is divided into multiple sections featuring lush forests, serene lakes, and small, picturesque villages. The central axis of the cylinder emits a soft, ambient light, mimicking natural sunlight, casting gentle shadows and creating a sense of day and night. The habitat's curvature should be evident, with the landscape bending upwards on both sides. Include a sense of depth and scale, highlighting the vastness of this self-contained world."
flux kind of got the idea but the gravity is still off.
https://github.com/black-forest-labs/flux/blob/main/model_li...
From the license: "We claim no ownership rights in and to the Outputs. You are solely responsible for the Outputs you generate and their subsequent uses in accordance with this License. You may use Output for any purpose (including for commercial purposes), except as expressly prohibited herein. You may not use the Output to train, fine-tune or distill a model that is competitive with the FLUX.1 [dev] Model."
This license seems to indicate that the images from the dev model CAN be used for commercial purposes outside of using those images to train derivative models. It would be a little weird to me that they'd allow you to use FluxDev images for commercial purposes IF AND ONLY IF the model host was Replicate.
They also now just added a "Commercial friendly" tag to the endpoints above.
Yes, it's weird.
Or at least a year ago it was like that.
Midlibrary.io recognizes 5500 styles Midjourney knows.
With that said, I loved Midjourney a year ago but I am at the point I have seen enough AI art to last several lives.
AI art reminds me of eating wasabi. It is absolutely amazing at first but I quickly get totally sick of it.
"Enable JavaScript and cookies to continue"...