The LLM’s logits should be translatable into probabilities although I’m not too sure how meaningful those might be as models can sometimes be quite confident in entirely invalid predictions.
I don't think that's the right way to think about LLM logits. Fundamentally the logits represent probability of similarity with the text it's been trained on, given the current prefix. Mixed in with any correspondence with truth is not only tone, phrasing, dialect, language syntax, but also stuff like the likelihood that specific details are related to general concepts. Even if we're talking about a person with three legs, or a horse riding a man, it'll be hard for the LLM to not assign a fairly high probability to sentences that describe two legs, or a man riding the horse and not the other way around.
I haven't done deep reading on LLM architectures and I don't really know if LLMs have logits in the traditional sense of a CNN or something, but I think the problem with this is that the LLM's logits would have absolutely no bearing on it's confidence of the location being correct, only on it's confidence that the tokens making up the answer it provides follow from the tokens that were encoded from the provided image, which isn't the same thing.