FLUX is fast and it's open source
replicate.com
replicate.com
Kind of like those video2video (or img2img on each frame I guess) models where they enhance the image outputs from video games:
https://www.theverge.com/2021/5/12/22432945/intel-gta-v-real... https://www.reddit.com/r/aivideo/comments/1fx6zdr/gta_iv_wit...
I want a picture of frozen cyan peach fuzz.
Prompt: frozen cyan peach fuzz, with default settings on a first generation SD model.
People _seriously_ do not understand how good these tools have been for nearly two years already.
https://pollinations.ai/p/frozen_cyan_peach_fuzz?seed=1
https://pollinations.ai/p/frozen_cyan_peach_fuzz?seed=2
https://pollinations.ai/p/frozen_cyan_peach_fuzz?seed=3
Disclaimer: I'm behind Pollinations.AI
Disclaimer: I'm not behind any
Imagine if instead of generating the RGB image directly the model would generate something like that, but with richer descriptive embeddings on each segment, and then having a separate model generating the final RGB image. Then it would be easy to change the background, rotate the peach, change color, add other fruits, etc, by editing this semantic representation of the image instead of wrestling with the prompt to try to do small changes without regenerating the entire image from scratch.
It seems sensible to extract features and reason about things the way a human would, but it turns out its easier to scale pattern matching purely done by computer.
https://www.cs.utexas.edu/~eunsol/courses/data/bitter_lesson...
The second general point to be learned from the bitter lesson is that the actual contents of minds are tremendously, irredeemably complex; we should stop trying to find simple ways to think about the contents of minds, such as simple ways to think about space, objects, multiple agents, or symmetries. All these are part of the arbitrary, intrinsically-complex, outside world. They are not what should be built in, as their complexity is endless; instead we should build in only the meta-methods that can find and capture this arbitrary complexity. Essential to these methods is that they can find good approximations, but the search for them should be by our methods, not by us. We want AI agents that can discover like we can, not which contain what we have discovered. Building in our discoveries only makes it harder to see how the discovering process can be done."
My point being, optimization or splitting up int subs, before handing over the problem to the machine, makes sense.
- We don't have limitless CPU cycles
- Thus we need to split things into sub-problems
If so that might still be amenable to the bitter lesson, where Sutton is saying human heuristics will always lose out to computational methods at scale.
Meaning something like:
- We split up the thought to vision problem into N sub-problems based on some heuristic.
- We develop a method which works with our CPU cycle constraint (it isn't some probe -> CPU interface). Perhaps it uses our voice or something as a proxy for our thoughts, and some composition of models.
Sutton would say:
Yeah that's fine, but if we had the limitless CPU cycles/adequate technology, the solution of probe -> CPU would be better than what we develop.
But i think we're onto something!
Voice to image indeed might give better results than text to image, since voice has some vibe to it (intonation, tone, color, stress on certain words, speed and probably even traits we don't know yet) that will color or even drastically influence the image output.
With image generation on the other hand, which starts from a handful of words, we can first do some text processing into categories, such as objects vs people, color vs brightness, environment vs main object, etc.
At the very least image generators should output layers, I think the style component is already possible with the img2img models.
If you can train a neural network that goes from a to b and a network that goes from b to c, you can usually replace that combination with a simpler network that goes from a to c directly.
This makes sense, as there might be information in a that we lose by a conversion to b. A single neural network will ensure that all relevant information from a that we need to generate c will be passed to the upper layers.
It's probably not a viable idea, I just wish for more composable modules that lets us understand the models' representation better and change certain aspects of them, instead of these massive black boxes that mix all these tasks into one.
I would also like to add that the text2image models already have multiple interfaces between different parts. There's the text encoder, the latent to pixel space VAE decoder, controlnets, and sometimes there a separate img2imgstyle transfer at the end. Transformers already process images patchwise, but why does those patches have to be even square patches instead of semantically coherent areas?
compositing.
do this with ai today where each layer you want has just the artifact on top of a green background.
layer them in the order you want, then chroma key them out like you’re a 70s public broadcasting station producing reading rainbow.
the ai workflow becomes a singular, recursive step until your disney frame is complete. animate each layer over time and you have a film.
Only the FLUX.1 [schnell] is open-source (Apache2), FLUX.1 [dev] is non-commercial.
> Any software is source-available in the broad sense as long its source code is distributed along with it, even if the user has no legal rights to use, share, modify or even compile it.
According to the OSI definition, you also need a right to modify the source and/or distribute patches.
> I don’t know any closed source apps that let you view the source.
A lot of them do, especially in the open-core space. THe model is called source-available.
If you're selling to enterprises and not gamers, that model makes sense. What stops large enterprises from pirating software is their own lawyers, not DRM.
This is why you can put a lot of strange provisions into enterprise software licenses, even if you have little to no way to enforce these provisions on a purely technical level.
Certain usages may be covered by trademark protection, as an "OSI Approved License":
<https://opensource.org/trademark-guidelines>
It's based on the Debian Free Software Guidelines (DFSG), which were adopted by the Debian Project to determine what software does, and does not, qualify to be incorporated into the core distribution. (There is a non-free section, it is not considered part of the core distribution.)
<https://www.debian.org/social_contract#guidelines>
Both definitions owe much to the Free Software Foundation's "Free Software" definition and the four freedoms protected by the GNU GPL:
- the freedom to use the software for any purpose,
- the freedom to change the software to suit your needs,
- the freedom to share the software with your friends and neighbors, and
- the freedom to share the changes you make.
https://pollinations.ai/p/a_donkey_holding_a_sign_with_flux_...
https://pollinations.ai/p/a_donkey_holding_a_sign_with_flux_...
https://pollinations.ai/p/Minimalist%20and%20conceptual%20ar...
It's incredible how fast it is. We generate 8000 images every 30 minutes for our users using only three L40S GPUs. Disclaimer: I'm behind Pollinations
[1]https://substackcdn.com/image/fetch/w_1456,c_limit,f_webp,q_...
Not sure I have an opinion on that, technology marches on etc, but it is interesting.
In the context of the point I made, it's definitely more similar to piracy, since the point was about taking advantage of something that if not free people would not pay for.
That differentiator is gone, and as such won’t pay for it anymore. They’ll just use the same AI as you.
This destroys the existing market of the artist.
To be clear, my comment isn’t meant as a judgment, just as market analysis.
Here are two videos that explain that well. I don't think I would ever be capable of designing with that degree of purpose given a generative AI tool.
[1] https://www.youtube.com/watch?v=mVyXUMJLzE0 [2] https://www.youtube.com/watch?v=5eymH15AfAU
The older I get the more concerned I get that the larger the team that makes decisions the worse the decisions are, whats the word for this? Is there any escape? Teamfortress 2, took years and teams to build, but it was just perfect.
I heard they had a flat structure which is even more confusing as to how they attained such an excellent product.
I have taste just no skill in drawing. I don't need an artist, i need a graphics designer and now i can replace a graphics designer with GenAI.
Plenty of Artists can draw very well, but what they learn in the industry is to learn to draw for someone else in an aligned art style etc. That has nothing to do with Art.
Very few people earn there living with being artists.
Its the same thing with all the other people. Look at masterpieces of woodworkers etc. They look interesting, nice but they normall just work for someone else doing their craft not their art.
All I wanna know is the prompt that was used to generate the art speaking of which i wanna know how to create cartoony images like that OP
"A hand-drawing of a scientific middle-aged man in front of a white background. The man is wearing jeans and a t-shirt. He is thinking a bubble stating "What's in a ReAct JSON prompt?" In the style of European comic book artists of the 1970s and 1980s."
Finding the right seed and model configuration is the more difficult part.
starving artists are going to famish now, not sure how to feel about it
I’ve found it odd how there’s a segment of the population that hates a shallow depth of field now, as they’re so used to their phone pictures. I got in an argument on Reddit (sigh) with someone who insisted that the somewhat shallow depth of field that SDXL liked to do by default was “fake.”
As in, he was only ever exposed to it through portrait mode and the like on phones and didn’t comprehend that larger sensors simply looked like that. The images he was posting that looked “fake” to him looked to be about a 50mm lens at f/4 on a full frame camera at a normal portrait distance, so nothing super shallow either.
In the real world, if I see a person at the beach, I can look at the person and see them in perfect focus, I can then look at the ocean behind them and it is also in perfect focus. If you are an AI generating an image for me, I certainly don't need you to tell me on which parts of that image I'm allowed to focus, just let me see both the person and the ocean (unless I tell you to give me something artsy :)).
Our eyes work the same way. Of course, just like the camera's aperture can be set our pupils will be pretty contracted on a beach.
Of course, you should be able to tell the AI to generate it how you want - that's the goal, after all. Having at least a somewhat shallow depth of field by default makes sense though.
I'll say I'm okay with DOF - it just feels (subjectively to me) like its incredibly exaggerated in Flux. The workarounds have mostly been prompt based adding everything from "gopro capture" to "on flickr in 2007" but this approach feels like borderline alchemy in terms of how reliable it is.
crazy not even a year has past since Emad's downfall a local open source and superior model drops
which just shows how little moat these companies have and are just lighting cash on fire which we benefit from
> which just shows how little moat these companies have
Flux was developed by the same people that made Stable Diffusion.
Did they scrape public facebook posts? Snapchat? Vkontakte? Buy private images from onedrive/dropbox? If I put as the second word a female name, it almost always triggers nsfw filter. So I assume images in the training set are quite private.
See for yourself (autoplay music warning):
people: https://vm.tiktok.com/ZGdeXEhMg/
food and stuff: https://vm.tiktok.com/ZGdeXEBDK/
signs: https://vm.tiktok.com/ZGdeXoAgy/
[edit] Looking at these images feels uneasy, like I am looking at someones private photos. There is not enough "guidance" in a prompt like "IMG00012.JPG forbid" to account for these images, so it must all come from the training data.
I do not believe FLUX 1.1 pro has radically different training set than these previous open models, even if it is more prone to such generation.
It feels really off, so, again, is there any info on training data used for these models?
These two reddit threads [1][2] explore this convention a bit.
DSC_0001-9999.JPG - Nikon Default
DSCF0001-9999.JPG - Fujifilm Default
IMG_0001-9999.JPG - Generic Image
P0001-9999.JPG - Panasonic Default
CIMG0001-9999.JPG - Casio Default
PICT0001-9999.JPG - Sony Default
Photo_0001-9999.JPG - Android Photo
VID_0001-9999.mp4 - Generic Video
Edit: Also created a version for 3D Software Filenames (all of them tested, only a few had some effects)
Autodesk Filmbox (FBX): my_model0001-9999.fbx
Stereolithography (STL): Model0001-9999.stl
3ds Max: 3ds_Scene0001-9999.max
Cinema 4D: Project0001-9999.c4d
Maya (ASCII): Animation0001-9999.ma
SketchUp: SketchUp0001-9999.skp
[1]: https://www.reddit.com/r/StableDiffusion/comments/1fxkt3p/co...[2]: https://www.reddit.com/r/StableDiffusion/comments/1fxdm1n/i_...
https://i.postimg.cc/vT6SV7pq/replicate-prediction-6ap8z1jv5...
https://i.postimg.cc/vZzMTM71/replicate-prediction-7r4b4p6sj...
https://i.postimg.cc/rs6wM5LJ/replicate-prediction-d8s4c93v5...
I DEMAND TO KNOW HOW RUN LOCAL SAAR
Of all the models exibiting this behaviour, has anyone published, what are the training data sources? Like, honest list, not the PR-boilerplate.
dont know why all the critical comments about flux are being downvoted or flag sure is weird
It seems likely that they did heavy calibration of text as well as a lot of tuning efforts to make the model prefer images that are “flux-y”.
Whatever process they’re following, they’ve inadvertently made the model overly sensitive to certain terms to the point at which their mere inclusion is stronger than a Lora.
The photos you’re showing aren’t especially noteworthy in the scheme of things. It doesn’t take a lot of effort to “escape” the basic image formatting and get something hyper realistic. Personally I don’t think they’re trying to hide the hyper realism so much as trying to default to imagery that people want.
The Original model shows the FRONT, the speed version shows the BACK of the corvette. It's a completely different picture. This is not similar but strikingly different.
So let's also set the record straight for FLUX: only one of the models released is open source -- FLUX schnell -- it's a distillation from the proprietary model that's much harder to work with.
Meta's Llama models have ironically much more permissive license for all practical intents and purposes and they are also incredibly easy to fine tune (using Meta's own open source framework, or several third party ones), while FLUX schnell isn't.
I think the open source community should rally behind OpenFLUX or a similar project, which tries to fix the artificial limitations of Schnell: https://huggingface.co/ostris/OpenFLUX.1
ooh why is synchronous fast? i click thru to https://replicate.com/changelog/2024-10-09-synchronous-api
> Our client libraries and API are now much faster at running models, particularly if a file is being returned.
... thanks?
just sharing my frustration as a developer. try to explain things a little better if you'd like it to stick/for us to become your advocates.
however i do have to ask.. ~2x faster for fp16->fp8 is expected right? its still not as good as the "realtime" or "lightning" options that basically have to be 5-10x faster. whats the ideal product usecase for just ~2x faster?
https://flux11pro.com/ (Maybe the same thing? Unclear.)
https://github.com/flux-framework/flux-core
https://github.com/facebookarchive/flux/tree/main (apparently archived now, but this was the first thing I thought of)
When it's that bad I think that the frequency of collisions for this name is an interesting topic in its own right.
> In general I think that's true and agree that minor name collision commentary is uninteresting, but in this case we're talking about 11 collisions (and counting) in tech alone, 3 of those in AI/ML and 1 of those specifically in image generation.
> When it's that bad I think that the frequency of collisions for this name is an interesting topic in its own right.
I'll respect your judgement on this and not push it further, but this is my thought process here.