Open Flamingo – open framework to train multimodal LLMs
laion.ai
laion.ai
Counting: https://imgur.com/KTuQ1Bv
Parse the chess board: https://imgur.com/2zYFK1P
(Result): https://imgur.com/Ei4MAl7
Few-Shot Object Detection (Pascal VOC): https://imgur.com/gZkDMn8
Few-Shot Object Detection (simplified): https://imgur.com/Hk8QGMd
Not quite there yet. I've been more impressed with the other new zero-shot multimodal models like Grounding DINO and Azure Dense Captioning. Really looking forward to putting multimodal GPT-4 through its paces as well.
Could it be that the actual issue has to do with it having trouble with small tokens (letters, numbers)?
Does it give a different result if you ask it to answer in a format like this?
> Please name what kind of piece is on each square of this board > A1: white rook > A2: white pawn > A3: empty > A4: empty > ...
Prompting can be so unintuitive sometimes. Maybe it just has an issue with the output representation or something...
[0] https://i.postimg.cc/GtrGs8mw/Screenshot-2023-03-28-at-5-19-...
You can almost definitely run it on consumer GPU if you swap out the language model for something smaller as well (although the performance would definitely not be as good on the language side).
It's so ugly to have the words break anywhere and for the intentional line breaks to still occur anyway.
All they needed to do was use media queries for at least three screen widths and adjust the font size in there accordingly.