Counting: https://imgur.com/KTuQ1Bv
Parse the chess board: https://imgur.com/2zYFK1P
(Result): https://imgur.com/Ei4MAl7
Few-Shot Object Detection (Pascal VOC): https://imgur.com/gZkDMn8
Few-Shot Object Detection (simplified): https://imgur.com/Hk8QGMd
Not quite there yet. I've been more impressed with the other new zero-shot multimodal models like Grounding DINO and Azure Dense Captioning. Really looking forward to putting multimodal GPT-4 through its paces as well.