Muse: Text-to-Image Generation via Masked Generative Transformers
muse-model.github.io
muse-model.github.io
What this model does more than anything else is demonstrate we're still in the early stages of generative models, and we can expect a lot of progress from architectural improvements over the next decade (in addition to the progress in compute and data that we're already counting on).
But the promise of a big efficieny gain will be an incentive for companies like midjourney to give it a go with their data.
This can go in any number of fantastical directions. But visual media both as a private/personal medium and salve as well as an enterprise-grade tool of mass entertainment and propoganda? Baby we're just getting started!
Am I wrong or is that the same architecture as DALL-E 1?
https://dreambooth.github.io/ https://textual-inversion.github.io/