>Does class-conditional mean that it takes a class like "car" or "face" and generate a new random image of that class?
Yup
nope, we don't do Imagen-style super-resolution. we go direct to high resolution with a single-stage model.
The overall architecture diagram does not explicitly show the conditioning mechanism, which is a small separate network. For this paper, we only trained on class-conditional ImageNet and completely unconditional megapixel-scale FFHQ.
Training large-scale text-to-image models with this architecture is something we have not yet attempted, although there's no indication that this shouldn't work with a few tweaks.
Can this architecture be used to distill models that need fewer timesteps like LCMs or SDXL turbo?