The obvious use case, especially on HN, is frontend dev of any kind at all. The second most obvious one is OCR of paper documents.
I really missed this feature when I had DeepSeek code a small game for fun. When writing UI and rendering code it could execute the game and get screenshots back, but then had to rely on my feedback on what had gone wrong. Models with vision can do much better here, finding more issues on their own