* those which don't offer a _reasonable_ API, or (I would guess a larger subset) those which don't expose all the same information over their API
* those things which one wishes to preserve (yes, I'm aware that submitting them to the Internet Archive might achieve that goal)
* and then the subset of projects where it's just a fun challenge or the ubiquitous $other
As an example answer to your question, some sites are even offering bounties for scraped data, so one could scratch a technical itch and help data science at the same time:
Scraping works great to get the data.
I don't like node/js but I use it to do the scraping as I view the code as trash and full of edge cases and unreliable data / types and I can't complain, a dynamic scripting language is great for that.
It tells you who is your governor, local/federal representative, senator and municipal president. Each representative lives on a different website so I wrote scrappers for each one.
GoodRX built a scraping system that tapped into all the major providers. Thats what a group of vaccine hunters in my state used to get appointments for folks that had tried but were unable to.