What is currently the best model (or multi-model process) to go from text-to-3D-asset?
Ideally based on FOSS models.
Ideally based on FOSS models.
If you actually want something consistent you should really generate images one by one and provide extensive description of what you expect to see on each frame
And if you want to make something like animation it's only really possible if you basically generate thousand of "garbage" images and then edit together what fits.