I like the idea, especially as "content-length" and "content-type" should indicate the likely size and quality of the image. Given that, even the simplest things become problematic when we're downloading millions or billions of objects.
If a roundtrip to get a response from the server was 100 ms and we used this technique for the 720 million potential images they found from processing Common Crawl, that's 833 days if done sequentially. If you perform 100 requests per second, that's still 8.3 days just spent retrieving the HTTP headers, let alone to then download promising images.