https://fastflux.ai/ for instant image gen using Schnell (but its fixed on 4 steps and is mainly a tech show off of inference engine by runware.ai)
https://www.segmind.com/ has API support with lots of options, I am using it to generate and set wallpaper using an AHK script. It's very very slow though.
https://replicate.com/black-forest-labs/flux-schnell/example...
https://huggingface.co/spaces/black-forest-labs/FLUX.1-schne...
https://getimg.ai/text-to-image
There are other tools now if you Google 'Flux image generator online'
I think there is a line somewhere between 4-bit to 8-bit that will hurt performance (for both diffusion models and LLM). But I doubt the line is between 8-bit to 13-bit.
(Another case in point: you can use generic lossless compression to get model weights from 13bit down to 11bit by just zip exponent and mantissa separately, that suggests the effective bit rate is lower than 13bit on full-precision model).
But yes, I do believe that we will find proper lossless quants, and eventually (for real this time) get "only a little bit of loss" quants, but I don't think that the current 8 bits are there yet.
Also, quantized models often have worse GPU utilization which harms tokens/s if you have the hardware capable to run the unquantized types. It seems to depend on the quant. SD models seem to get faster when quantized, but LLMs are often slower. Very weird.
If we start from peering into quantization, we can show it is by definition lossy, unless every term had no significant bits past the quantization amount.
so our lower bound must that 0.03% error mentioned above.
However this seems to be model size dependent, ex. Llama 3.1 405B is reported to degrade much quicker under quantization