The really wild one is even blind models will do this and they'll try to run stats on the pixels to figure out what it looks like... the even wilder thing is that it kind of works!
I don't think the n by n subgrid fixes this the way most harnesses do, as it'll fail to count things if you have more overlap and fail relatiomships if you have less
And if you are counting things it should be trivial to note the position of your items and not double-count them, no?
Can you point me at one that can understand a detailed block diagram? Frontier is fine, soliciting recommendations
I tried to parse hand-drawn ER diagrams in the past and did not have much success with any model, frontier or otherwise. If you have annotated data, I'd recommend finetuning a recent (dense) VLM, but don't expect 100% accuracy. https://unsloth.ai/docs/basics/vision-fine-tuning
Ideally, look at second diagram and see something similar in the same location and understand it's the same physical object.
LLM issues tool call to read high res image ->
harness sends high res image to server ->
server downsizes it to 800x800 (blurry) ->
LLM issues bash command (e.g. `convert`) to crop a small subimage (e.g. 600x600) from the high res image ->
LLM issues tool call to read subimage ->
harness sends subimage to server ->
server does not resize the subimage because it is small already, so it is not blurry when finally ingested by the LLMMaybe there are some use cases where you need high detail everywhere at once, but for OCR of small text and the like a zoom ability should be sufficient
Can't remember if I stole this idea from some existing public harness though, can't remember. If someone knows of public harnesses that do this already, please share them :)