I'm excited to see what a Flux 2 can do if it can actually use a modern text encoder.
The image generators used by creatives will not be text-first.
"Dragon with brown leathery scales with an elephant texture and 10% reflectivity positioned three degrees under the mountain, which is approximately 250 meters taller than the next peak, ..." is not how you design.
Creative work is not 100% dice rolling in a crude and inadequate language. Encoding spatial and qualitative details is impossible. "A picture is worth a thousand words" is an understatement.
Controlnet has been the obvious future of image-generation for a while now.
We might find that the entire "studio system" is a gross inefficiency and that individual artists and directors can self-publish like on Steam or YouTube.
Automation tools are always more powerful as a force multiplier for skilled users than a complete replacement. (Which is still a replacement on any given task scope, since it reduces the number of human labor hours — and, given any elapsed time constraints, human laborers — needed.)
"""
That's a fun idea—but generating an image with 999,999 anime waifus in it isn't technically possible due to visual and processing limits. But we can get creative.
Want me to generate:
1. A massive crowd of anime waifus (like a big collage or crowd scene)?
2. A stylized representation of “999999 anime waifus” (maybe with a few in focus and the rest as silhouettes or a sea of colors)?
3. A single waifu with a visual reference to the number 999999 (like a title, emblem, or digital counter in the background)?
Let me know your vibe—epic, funny, serious, chaotic?
"""
Sora is one of the worst video generators. The Chinese have really taken the lead in video with Kling, Hailuo, and the open source Wan and Hunyuan.
Wan with LoRAs will enable real creative work. Motion control, character consistency. There's no place for an OpenAI Sora type product other than as a cheap LLM add-in.