Maintaining scrapers for 18k county websites and PDs is no small task and looking through the docs for PDAP, it seems like this is still a very open question.
Maintaining scrapers for 18k county websites and PDs is no small task and looking through the docs for PDAP, it seems like this is still a very open question.
The 80000 Hours podcast has an interview with the (non-technical) creator of OWID. I seem to recall some interesting stories about them getting emailed PDFs with COVID data and such.
I had the same question as you, and I was hoping to find ideas in the comments. It seems like the kind of thing that's both inherently messy and scrappy yet if you don't get at least somewhat organized it can't scale.
Update: link to the podcast episode page with quotes, transcripts, etc. https://80000hours.org/podcast/episodes/max-roser-our-world-...
It's interesting that even one of the largest still uses manual execution for almost all of their pipelines (at least in the covid data project[1]). This [2] seems like the bulk of their data importers (scrapers) but most are still operating as manual jobs. I guess with open-source data work, hours and minutes don't matter as much, and being a few days behind the latest data is acceptable.
1 - https://docs.owid.io/projects/covid/en/latest/data-pipeline.... 2 - https://github.com/owid/importers