- Do you ask the multimodal LLM to return the image with boxes drawn on it (and then somehow extract coordinates), or simply ask it to return the coordinates? (Is the former even possible?)
- Does it better or worse when you ask it for [xmin, xmax, ymin, ymax] or [x, y, width, height] (or various permutations thereof)?
- Do you ask for these coordinates as integer pixels (whose meaning can vary with dimensions of the original image), or normalized between 0.0 and 1.0 (or 0–1000 as in this post)?
- Is it worth doing it in two rounds: send it back its initial response with the boxes drawn on it, to give it another opportunity to "see" its previous answer and adjust its coordinates?
I ought to look at these things, but wondering: as you (or others) work on something like this, how do you keep track of which prompts seem to be working better? Do you log all requests and responses / scores as you go? I didn't do that for my initial attempts, and it feels a bit like shooting in the dark / trying random things until something works.