A well-written scraper would check the image against a CLIP model or other captioning model to see if the text there actually agrees with the image contents.
Or is the internet so full of garbage nowadays that it is necessary to do that on every page?