Do you think we’re getting closer to models that actually understand what they’re seeing, or are they just getting really good at recognizing patterns?
"yes" but that's a philosophical question. I think they're getting better at "early fusion" IE, training the model that "apple" and these visual tokens are the the same concept, but LLMs are fundamentally a pattern matching machine so even with perfect fusion I personally wouldn't call it understanding.
All I'm fine with for now is that I can almost exclusively communicate with Sol through collages and my scribblings (all kinds of web page / block screens with all kinds of arrows and text all over the place) This was not practically ppossible in 5.5 and a tragedy in 5.4. Not sure how much weight is codex uploading in higher res carrying here but it's great to work with.