My instinct was also to use LLMs for this, but it was way to slow and still expensive if you want to scrape millions of pages.
Edit: There's also the added complexity of running a browser against 1M pages, or more.
This library also supports HTML as input so running a browser is not required.