Show HN: Real-time image generation with SDXL Lightning
fastsdxl.ai
fastsdxl.ai
This feels like the future, near real time image and LLM generation (using Mixtral from Groq as my prompt writer) and Fal API for read time generation!
Competitive prompting?
I am imagining the green lush landscape from early parts of the demo to slowly transform into the dry mountainous landscape from later images while new characters appear in the foreground.
(I'd posted this comment incorrectly under the main HN post earlier, instead of as reply here. Too late to delete it apparently.)
There are a few other UIs for it, e.g. https://replicate.com/lucataco/sdxl-lightning-4step
EDIT - ah you have one. You're welcome. Sign up here folks. :)
Couple of questions in that case: a) What is the avg price per 512x512 image? Your pricing is in terms of machine resources, but (for my use case) I want a comparison to pexels. b) What would the equivalent machine setup be to get inference to be as fast as the website demo? c) Is the fast-sdxl api using the exact same stack as the website?
To all your questions, I recommend playing with it in the API playground, you'll be able to test different image sizes, parameters, and have an idea of the cost per inference.
If you have any other questions, say hello on our Discord and I'm happy to help you.
https://huggingface.co/spaces/radames/Real-Time-Text-to-Imag...
I you have GPU/cuda/Docker you can try it locally
docker run -it -p 7860:7860 --platform=linux/amd64 --gpus all -e SFAST_COMPILE="1" -e USE_TAESD="0" registry.hf.space/radames-real-time-text-to-image-sdxl-lightning:latest python app.py
As for the quality, I borrowed the query ([1]) that people used to test Stable Diffusion 3 and other models today: "Photo of a red sphere on top of a blue cube. Behind them is a green triangle, on the right is a dog, on the left is a cat".
Here is what I got: https://imgur.com/a/XrAuqCB
To compare it with Stable Diffusion 3: https://pbs.twimg.com/media/GG8mm5va4AA_5PJ?format=jpg&name=...
Test the example on Stable Cascade as well (latest open-weight stability model), and yeah, even that is not great at it https://fal.ai/models/stable-cascade?share=eab44060-690b-497....
(seed: 3919562)
Btw this is from fal.ai, I first heard of them when they posted a Stable Cascade demo the morning it was released.
They're *really* good, I *highly* recommend them for any inferencing you're doing outside OpenAI. Been in AI for going on 3 years, and on it 24/7 since last year.
Fal is the first service that sweats the details to get it to the point it runs _this_ fast in practice, not just in papers. ex. web socket connection, use short-lived JWTs to avoid having to go through an edge function to sign a request with an API key, etc.
the latent space stuff became popular through it being a visual allegory, which accidentally confused the technical term it originated from. there's nothing visually smooth about it, it's not a linear interpolation in 3D space, it's a chaotic journey through 3 billion dimension space
These low-step approaches probably preserve a lot less of the 'noise' features in the final image so latent space cruising is probably less fun.
Side note: In my opinion, the fast generations also make up a lot for the shortcomings in image generation quality. I find that even if it messes up, a good result is usually just a seed or small prompt change away.
By the way the racoon is very prone to get two tails if you change the prompt a little. :)
https://developer.mozilla.org/en-US/docs/Web/API/URL/createO...
It's a way to turn a file, or blob, into a URL usable with an image element etc.
What I mean is if my first prompt is a girl talking to cat and second prompt is girl playing with that cat, I want the girl and cat to be the same in both pictures.
Is that possible? If so any links or tutorials will be super helpful to learn.
`late 90s movie poster, 24 hour clock movie "2: Electric Boogaloo" dan aykroyd1`
turned out great
what a hero looks like: https://fastsdxl.ai/share/x9jxax4pnljd
what a terrorist looks like: https://fastsdxl.ai/share/ejtyvv9ahpfs
what the person I wish to be looks like: https://fastsdxl.ai/share/8ekkecm5rqsr
This is very interesting as the fast pace can help to quickly evaluate the biases incorporated just changing the seed.
I am imagining the green lush landscape from early parts of the demo to slowly transform into the dry mountainous landscape from later images while new characters appear in the foreground.
The results also look great, but the more I see AI generated images the more I believe it is not going to eat jobs.
Almost none of the results are production ready. They all contain strange parts. And maybe more important: the look and feel is always the same.
It is amazing this works as fast as it does, but I think AI is still in it's 'hype' stage.
It can do excellent photo real, sketch, pixel art, diagrams, painting, digital painting, 3d, all to a high standard.
I think of current ai/ml stuff right now as a very fast intern assistant that will get you 85% of the way there but needs supervision.
But the technology is still so new it will only get better.
This is an important thing, but not a consistent thing between humans. For example, while I can tell half of these are AI (and not just because I typed in the prompt), they have very different looks and feels to me:
• https://fastsdxl.ai/share/6djh0dlat0s6 "Will Smith facing a white plastic robot, close up, side view, renaissance masterpiece oil painting by da Vinci"
• https://fastsdxl.ai/share/ctwqegl5i3xq "a hand stitched embroidery of a cute tiger-racoon playing in woodland"
• https://fastsdxl.ai/share/mkfrx33xc4ee "a selfie shot of a furry in a furry convention"
• https://fastsdxl.ai/share/mphyrzzjsces "Simple sketch of Mordor, Mount Doom, dark and moody, despair, dense fog"
• https://fastsdxl.ai/share/hgjwx6avyx0h "coffee mug stain on paper"
But there are many others like yourself who apparently have higher standards than I do.
(And there are also many who have lower standards than me, who were happy to print huge posters where the left and right eyes of subject didn't match).
you only end there because it's so fast. nice. really a new explorative quality.
Very disturbing results are just a second away!
One bug though: when I increment the seed, it renders two image and it takes a bit of jumping up and down in seed numbers to get to the same image.
of course the quality of what is being generated is not competitive with SOTA, but this is going in a really good direction!
Regardless of how we feel, lawyers and regulators wait at the door. We should expect new legal precedent within the next year re the generation of copyright-infringing, deepfake, and pornographic material.
Simply try to get it to output an image with a female that's not a beauty queen. Even when specifically prompted to produce an image of ugly people, it can only generate beautiful people.
https://fastsdxl.ai/share/dfxtkr3r0w3z
https://fastsdxl.ai/share/e70xgiudn8j0
But when I switched the prompt to include "portrait", yeah it produces women who are too attractive, like an actress wearing makeup to appear ugly for a role.
Needs work.
OP's post to me feels like a marketing post where the output image is a really close representation of the product they hope to sell. We always called these types of things "carefully selected, random examples", in short they are cherry-picked for their adherence to a standard.
In that same vein mine is also a carefully selected, random example of the output you get when the algorithms don't work well, therefore the "Needs work" qualification.
Both are useful since you need to understand the limitations of the tools that you are employing. In my case I stepped thru animals until I found one that it could not render accurately. It did know that a coelacanth is a fish but it couldn't produce an accurate image of one. Then I added modifiers that it could not place in context.
It's a bit like searching the debris field of a tornado for perfectly rounded debris particles and holding those up as a typical result without mentioning that you end up having to ignore all the splintery debris scattered from hell to breakfast around it.
*Edit:* The previous poster specifically said they use an iPad, that's odd. Worked fine on my 12,9" iPad Pro, Display Zoom is set to "more space", i.e. more content shown. In case it has an issue with not enough screen space?
That being said, I don't think the refresh button (actually a seed randomizer) is related. I was definitely already playing around with it on my iPad when you posted your comment. However, I am not sure about the timeline in relation to the first comment (above yours).