Isn't "two models slapped together" basically how all of these things work, starting with CLIP? Not sure about GPT4o, obviously, I don't think they released any underlying architecture details?
- single interconnected neural network (LLM attention layers break this, autoencoders complicate this)
- single training pass (LLMs have multiple passes, GANs have a single but produce multiple models)
LLMs have multiple passes? wdym?