Web image size prediction for efficient focused image crawling
blog.commoncrawl.org
blog.commoncrawl.org
If a roundtrip to get a response from the server was 100 ms and we used this technique for the 720 million potential images they found from processing Common Crawl, that's 833 days if done sequentially. If you perform 100 requests per second, that's still 8.3 days just spent retrieving the HTTP headers, let alone to then download promising images.
I've used fastimage with success: https://github.com/sdsykes/fastimage
From testing with 1800+ image URLs, you can almost certainly get the dimensions within the first 64kb.