This is the same approach as Mini-GPT4 of training a linear layer using frozen image encoder and text decoder. the difference is that they use CLIP instead of BLIP, and LLaMA instead of Vicuna (which was trained on LLaMA. Interesting that they both came out at the same time.