FLUX 3 Image
bfl.ai
bfl.ai
Ideogram V4, an open-weight model released back in June can also do this [1], but you have to use a relatively cumbersome JSON structure to describe all the different bounding boxes. So it’s definitely a bit of a hassle.
I'll probably be waiting until it goes open-weight (hopefully soon) like they did with Flux.2 / Klein.
[1] - https://docs.ideogram.ai/using-ideogram/getting-started/prom...
It's one of the most uniquely hostile user experiences I've ever had the (dis)pleasure of working with
However, today I don't see a reason to use ComfyUI at all.
For Qwen Image 2.1, I had Opus 5.5 create a backend outside of ComfyUI and it was able to make generation take 20% less time with some optimizations.
The optimizations it implemented were caching the text computation in Qwen Image 2.1 rather than including it in every step, fusing projections into a larger matrix multiplication, and decoding the VAE in horizontal bands or something like that.
If there was anything interesting in ComfyUI nodes, I imagine I could just have the AI adopt the relevant code instead of dealing with ComfyUI or custom nodes.
From where I'm sitting, it's just turning python functions into boxes and instead of write the function yourself, you drag from the output of one box to the input of another. For 2 or 3 boxes, this is cool, but I opened up a professional workflow and was taken into a view with 100s of boxes and wires all over the place. Uhh, ok?
For myself, I'd rather just create my own python environment, write some quick pytorch or mlx calls, wire up some cli to it and share that in a GitHub.
Disclaimer: never used it for actual generation so I don’t know if it’s doing anything special other than being a subgraph with a different name.
That being said: defining composition made me immediately think that someone probably made a gui like this with easy to move bounding boxes, and I'm happy they did.
Generate a sprite in an image editor, then use a video model to make the loop you want; then turn the resulting video back into individual sprite images.
I can still get absolutely insane results with MiniMax H3 - insane in the sense that it would not make sense at all and would make your head spin.
Video Generation Models as World Models: Efficient Paradigms, Architectures and Algorithms
https://bfl.ai/models/flux-3-action
The idea to use predictive models of future observations for decision making has a long history, including early demonstrations of robot control which combined action-conditioned video prediction with model-predictive control.
Clearly I'm the uneducated one here eh
That paper sets up a task where generating a correct video requires correctly modeling physical laws. From the failure to always generate the correct video, they infer that the model has failed to correctly model the physical laws. The whole premise of the experiment is that learning visual statistics is equivalent to learning causal dynamics, such that failure at one implies failure at the other.
The main difference in applications is that the bar for entertainment is lower, so that even a very bad world model may be acceptable.
https://media.discordapp.net/attachments/1401891025970008154...
I don't know what went into making it, but their twitter is @araminta_k if you're curious
* https://alvdansen.github.io/animating-on-twos/
* https://github.com/alvdansen/animating-on-twos
* https://huggingface.co/alvdansen/h3-keyframe-animation
## Quick Start
An (apparently, as I haven't tried it) ready-to-go ComfyUI graph for the above. In theory you should be able to drop these into ComfyUI and have it work:
https://huggingface.co/alvdansen/h3-keyframe-animation#quick...
Imagine online procedural MMO with old gen final fantasy / chrono trigger styles :)
The other examples range from mostly good to okay'ish at least.
Chats can be awful user interfaces.
The website mentions:
> FLUX 3 Image is available under a commercial weights license for companies running image generation at scale. Fine-tune and deploy it on your own infrastructure. Reach out to us to learn more.
I guess the open ones would be non-commercial?
"Open Weights version of FLUX 3 Image is launching in the coming weeks."
What else?
- Bunch of visible shape inconsistencies: building window sizes and alignment, fence railing, rivets on the bench, coins, accordion features, perspective violations etc...
- Metal clips on the bottom and top edges of the carrying case are on different sides. Said clips look inconsistent.
- One of the case straps originates inside (from under the velvet, even), the further one originates outside. The carrying handle is oddly mis-centered (as with most things).
- I'm not sure that accordion, compressed, fits in that case. She can just about stand in it (diagonally, on one foot).
- It has weird folds that disappear partway through and turn into adjacent folds going the other way. Between this, the keyboards and the dots, I'm sure there are other things wrong with accordion anatomy.
- Awkward left hand: not sure what's going on with the metacarpal joints, and index finger appears to begin forking into two tips pressing into two (misshapen) buttons. Her right hand is getting there too.
- Weird hair midline. I haven't seen a "sub-midline" like this in the wild :)
- The burgundy coat clips through the bench under her right leg. It's like the front slat goes through a hole in the side of the coat.
- We don't usually have that side curling bar on each side of the sitting area. It looks like it's clipping through the flats next to the leather handbag
- Ill-placed buckle on that bag.
- The green metal fence on the right begins behind her for one vertical bar, next one disappears into the plants, no further vertical bars, no top horizontal railing. Actually, it appears to only have fencing on one side, that's not something I'd see often. On the side where it does have a top, its top rail has very irregular width.
- Sand looks like... breadcrumbs? Very repetitive coarse shapes, like it was poorly done with a clone stamp.
- Something weird about the leaves.
- The foremost lamp isn't sure whether it has flat sides or not.
- The other lamp emerges out of nowhere from behind one of the blob-people in the alley.
- A bunch of unfortunate tangents and occlusions which make you wonder whether the diffusion hallucinated detail out of bigger shapes, or if it's been trained on photos where the tangents were done on purpose.
ie. center bottom bar of the bench being exactly aligned with the gray edge tiles, side curly railing-not-present-on-our-real-benches fully formed on her right, almost completely foreshortened out of existence where the bench's perspective forbids it...
Most people are fooled by much cruder fakes though. This is a lot better but keeps plenty of tells for observant people (and most people aren't)
Or just put in more than 10 words of effort, e.g. a better prompt and bounding boxed prompt edits, and all that stuff would be fixed now.
When I was a kid, there was a game in the funny pages that asked you to find six things wrong in a picture. Someone should run a contest like that for AI-generated images.
It might be an interesting benchmark for vision models, too.
No signal in digital images anymore. If you didn’t see it with your own eyes, it likely never happened
Nothing in life is ever perfect. Doesn't mean imperfect stuff can't have a lot of impact.
Me, casually scrolling on my phone, don’t
I have no way of knowing that you wanted the image to have a red wetsuit. I would just assume its a real image of a surfer in a red wetsuit
As I'm finding with the best GenAIs, this allows for granular iteration, which is where it becomes useful in an industry-wide manner.
The more things change, the more they stay the same.
That would be to compare e.g. Qwen Image 3.0 with FLUX 3, with Midjourney etc.
I had seen some attempts - but I do not know well how they try to approach objectivity.
Prompt results are graded based a weighted calculation which includes: adherence to the prompt, image fidelity, and steerability.
My comparison benchmark also tends to favor prompt adherence, which a lot of others don’t. Most of ones that I've seen tend towards rather simplistic prompts (e.g. "neon-lit city facing a robotic uprising, with high-tech battles, in anime style"), whereas the prompts I've created try to test high specificity.
I’ve been running them all the way back to SDXL.
You can compare specific models using the "View All Models" so if you want to see the progression of open-weight models, or model X vs model Y, you can do so.
Just a heads up - I haven't added Flux 3 as I'm waiting until BFL drops the open-weights version.
Generative Comparisons:
https://genai-showdown.specr.net
Editing Comparisons:
Tried adjusting exposure; Didn't work as well.
If you do not consider contemplation, for which we have filled our homes with works of art for centuries,
and if you do not consider those job that involve graphics production e.g. in the Madison Avenue "Mad Men" business (advertising - the payments are authentic),
you could consider "authentic" that when we do video production through generative models some start from stills (image generation) and then have other models animate those stills.
What's new from the last post? GA?
> We will open up an early access phase for FLUX 3 Image in the following weeks.
Not sure if there was a separate post for early access or if they just skipped to this.
Exact prompt:
Generate a photorealistic image of M81 urban BDU camouflage cargo trousers, shown by themselves. One trouser leg should be posed with the knee lifted 30° from vertical.
Accurate reproduction of the M81 urban camouflage pattern is critical. Match its colors, shapes, scale, distribution, and overall appearance as faithfully as possible.
No person, other clothing, or props.
Ground truth swatch: https://commons.wikimedia.org/wiki/File:US_City_Camo_(M81_Ur...
Gemini 3 Pro Image (stronger pattern): https://i.postimg.cc/bZNQYYjx/2026-10-02-google-gemini-3-pro... Flux 3 (this run): https://i.postimg.cc/Xr7wNN0H/2026-10-02-black-forest-labs-f...
Flux 3 gets greyscale urban-ish trousers and a lifted knee, but the blotches aren’t real M81 Urban — softer / wrong geometry vs the swatch. Not the worst I’ve seen on this prompt; clearly behind the Gemini 3 Pro Image example above on pattern.
Curious what other models do on the same prompt.
Flux is in the top 9000 of the most common words.