A more honest name would be Visual-Vicuna or Son-of-BLIP.
A more honest name would be Visual-Vicuna or Son-of-BLIP.
The model as a whole is just BLIP-2 with a larger linear layer, and using Vicuna as the LLM. If you look at their code it's literally using the entire BLIP-2 encoder (Salesforce code).
maybe even add an "a" for extra spice: Son-of-a-BLIP
The number of parameters used for GPT-4 is unknown.
[0] https://twitter.com/SebastienBubeck/status/16441515797238251...
So we're back to guessing ...
A couple of years ago Altman claimed that GPT-4 wouldn't be much bigger than GPT-3 although it would use a lot more compute.
https://news.knowledia.com/US/en/articles/sam-altman-q-and-a...
OTOH, given the massive performance gains scaling from GPT-2 to GPT-3, it's hard to imagine them not wanting to increase the parameter count at least by a factor of 2, even if they were expecting most of the performance gain to come from elsewhere (context size, number of training tokens, data quality).
So in 0.5-1T range, perhaps ?
https://www.reddit.com/r/IAmA/comments/12rvede/im_stephen_go...