I've been considering trying to launch an OpenStates style scraper project for US laws and admin codes, but haven't had the time to attack 100 more scrapers. Even with AI help, the volume is significant.
I've been considering trying to launch an OpenStates style scraper project for US laws and admin codes, but haven't had the time to attack 100 more scrapers. Even with AI help, the volume is significant.
I was thinking something more along the lines of a git repo per state.
If it was as easy as writing a scraper and dumping it all in a bucket or repo, it'd already be done. It's just the usual thankless hard work over time grind.
Even with openstates, we have an API but don't "just" dump the bills to git for legacy nerd reasons.
The nice thing about laws is that the host websites (or PDFs) don't change templates _that_ often, so generally you can rescrape quarterly (or in some states, annually) without a ton of maintenance. With administrative codes you need to scrape more often, but the websites are still pretty stable.
The downside is that codes in particular are often big, so a single scrape might need to make 20,000 or more requests, so you have to be very careful about rate limiting and proxies, which goes to my original point that it sucks that accessing this stuff is such a mess.
Current information is gated behind a web2.0 view of their live data with severe limits. It wasn't designed to be scraped and is in fact hostile to the attempt. I'd imagine they're seeing rising hosting costs and that they'll keep rising.
I should reach out to them and see what this looks like from their angle. The local commercial real estate community is pretty tech-savvy and I'm wondering if we could all be a bit more proactive around data access.
I'd love to hear your thoughts on county vs state vs national data! I'd be very interested in any bandwidth usage or processing requirement info you might have recorded.
But the person who did the codification has some rights thereto, meaning that while NV can post every act that passed the legislature, they can’t publish someone else’s codification of the statutes.
This matters very little because everyone just has Westlaw and no one uses the state legislature’s website to cite statutes.
IMO, this should also extend to opinions -- if there is precedent that guides what the law is, it needs to be publicly published free of charge so that the public is put under notice what the law is. (someone might mention something like PACER is free in small quantities, I would counter it would cost you a gazillion dollars to be fully informed of all the precedent that forms the full common law meaning of the laws.) This is especially important in mala prohibita crimes since there's no way to even guess through moral/ethical deduction.
I reckon that’s why the sixth amendment exists but if you want to make a free PACER, go for it.
If anybody is worried about the jobs those businesses created, then tell them to pivot into publishing commented editions of the codes (add cross-references, references to relevant court decisions, etc.).
But you could do it too! The Congressional Record is a thing, and it publishes all the acts of Congress, all the way back to the beginning.
The problem is that after you were done, the first thing someone would ask you is to cross-cite everything into the West Annotated code because no one else has your code and no one cares about it, because we all have Westlaw.
(Which publishes commented editions of the codes, with cross references, references to relevant court decisions, etc.)
It's all a little bit antiquated but it works fine. Someday it will change. I too thought it should work the way people are describing upthread when I was a computer guy but it is what it is.
I have no idea whatsoever what is going on in Nevada.