A pro/con of the multimodal image generation approach (with an actually good text encoder) is that it rewards intense prompt engineering moreso than others, and if there is a use case that can generate more than $0.17/image in revenue, that's positive marginal profit.
That said it’ll be 10-20x cheaper in a year at which point I don’t think you care about price for this workflow in 2D games.
- specificity (a diagram that perfectly encapsulates the exact set of concepts you're asking about)
AI companies are still in their "burning money" phase.
Enshittification is not on the horizon yet, but it's inevitable.
I believe, over all, development will go forward and things will get better. A rising tide lifts all ships, even if some of them decide to be shitty leaking vessels. If nothing else we always have open source software to fall back on when the enshittification of the proprietary models start.
For a practical example: The cars we drive today are a lot better than 100 years ago. A bad future isn't always inevitable.
etc
Anyone using image gen for real work not just for fun.
Although you're way better off finding your own workflows with local models at that scale.