Are any of these multimodal yet? I'd love to try asking a model with calibrated probabilities to answer question like, "do these shapes match?". Sure, you can ask a LLM....
v21 supports vision: https://huggingface.co/StrandsAgents/strands-decider-2B-hobs...
(Disclaimer, I work at Cloudflare, but not on models)
I believe image classification/analysis by deciders (not just OCR, not everything is about text) is still lacking.
Cloudflare's Clef had fair results on my test, but it's larger and slower. Wondering about Strands.