I have a strip-tags CLI tool which I can pipe HTML through on its way to an LLM, described here: https://simonwillison.net/2023/May/18/cli-tools-for-llms/
I also do things like this:
shot-scraper javascript news.ycombinator.com 'document.body.innerText' -r \
| llm -s 'General themes, illustrated by emoji'
Output here: https://gist.github.com/simonw/3fbfa44f83e12f9451b58b5954514...That's using https://shot-scraper.datasette.io/ to get just the document.body.innerText as a raw string, then piping that to gpt-3.5-turbo with a system prompt.
In terms of retaining context, I added a feature to my strip-tags tool where you can ask it to NOT strip specific tags - e.g.:
curl -s https://www.theguardian.com/us | \
strip-tags -m -t h1 -t h2 -t h3
That strips all HTML tags except for h1, h2 and h3 - output here: https://gist.github.com/simonw/fefb92c6aba79f247dd4f8d5ecd88...Full documentation here: https://github.com/simonw/strip-tags/blob/main/README.md