Isn't this outdated already. It's 2 models slapped together rather than natively multimodal.
Not to mention it's built to be fine tuned and commercially permissive!
- single interconnected neural network (LLM attention layers break this, autoencoders complicate this)
- single training pass (LLMs have multiple passes, GANs have a single but produce multiple models)
LLMs have multiple passes? wdym?