After all, the web was built on accessibility of information, not on purposeful obfuscation.
If you go so far as to essentially flatten the webpage to the point where you might as well print it out and then do OCR on it then you've thrown out the baby with the bathwater, you had all that information when you started. Or at least, you should have had it.
Otherwise we might as well kiss HTML goodbye and render the web as pdfs, with or without links.
Essentially that's what Diffbot (https://www.diffbot.com/) does, except we don't the render pages as an image nor do OCR.
Diffbot renders the page in a headless browser, and uses computer vision to automatically identify the key page attributes and extract normalized data for specific page types (Articles, Products, Discussions, Profiles, Images, and Videos).
This approach enables us to work in any language and on sites that we've never come across before automatically with better than human level accuracy.
Answered on the parent, but it's somewhat similar.
There's a lot of very bad HTML out there.