I would be more interested in image-to-text models. Does someone know of any decent model? I saw the GPT4 demo, and they showed that they do image-to-text... but then that was actually a fake (i.e., the model was interpreting the image filename).
Image: https://i.imgur.com/husplYZ.png
Output: "a white horse with a sign that says rexel's in space, pixelperfect, inspired by Paul Kelpe, official simpsons movie artwork, alternate album cover, in style of nanospace, by Apelles, pickles, pespective, pop surrealism, ingame, in a space cadet outfit, sifi"
For a more-robust-but-hard-to-run model, you can use BLIP2: https://huggingface.co/Salesforce/blip2-opt-2.7b
I believe the current SOTA test for NLVR is VQAv2[0] or GQA[1].
0: https://visualqa.org/ 1: https://arxiv.org/pdf/1902.09506.pdf