The Unitedstates Project
theunitedstates.io
theunitedstates.io
One thing that my friend who works in Open Data has told me is that it's important for websites like this to exist, to be able to point non-technical people at them and say "SEE. THIS is why you can't just publish everything as a PDF".
This not only needs access to pacer, but a good algorithm and a huge staff to catch up to westlaw/lexis.
But even a completely functional Courtbot-style site is really only a competitor to something like Findlaw.com, not Westlaw or LexisNexis. That's because those companies have features that would be difficult to replicate:
- Very complex and comprehensive search options and parameters. This is possible to do but time-consuming and tricky. - LexisNexis has 15,000 employees, and I suspect a significant number are involved in reviewing cases, summarizing, noting authorities and conflicts, etc. It's not yet possible to replace a trained lawyer reviewing an opinion with a regular expression. :) - A subscription to LN/WL typically gets you, depending on your package and how much you're paying, far more than just court opinions. You get news articles, journal articles, congressional transcripts, and a slew of databases that can be used to look up info on people, locate assets, etc. A lot of this means licensing deals, and LN/WL effectively gives you a one-stop shop for a wealth of data. Some of this is coming online and is becoming searchable, but not enough to make a real dent.
The one thing that challengers have in their favor is that Lexis and Westlaw are expensive. I've had free accounts because of faculty affiliations or a newsroom subscription, which is grand, but it's cost-prohibitive for many people and businesses. The ABA has published a list of alternatives; note the majority are actually still owned by Lexis and Westlaw: http://www.americanbar.org/groups/departments_offices/legal_...
- We've built https://www.courtlistener.com to provide a powerful search system with millions of opinions and scrapers for lots of jurisdictions.
- We have the RECAP project that ingests content from PACER: https://www.recapthelaw.org
- We'll start collecting an archive of oral arguments soon (just got funding for that, if all goes smoothly).
Everything we do is open source and open access, so hopefully if we fail, people will take our code and content and keep it alive.
You're right that it's a big challenge, but I think we're making some headway.
> This is an unusual, and occasionally chaotic, model for an open data project. the /unitedstates project is a neutral space; GitHub's permissions system allows many of us to share the keys, so no one person or institution controls it. What this means is that while we all benefit from each other's work, no one is dependent or "downstream" from anyone else. It's a shared commons in the public domain.
From http://sunlightfoundation.com/blog/2013/08/20/a-modern-appro...
I've done a bunch of technology voters guides for Wired and CNET by crawling House/Senate records (what a pain) and that's one thing I always thought would be useful. Not enough attention is paid to them, and many bills don't get to the floor. There were plenty of SOPA committee votes on amendments, but the legislation never made it to the floor.
The Github @unitedstates Project, is an open, relatively decentralized directory to find tools and data related to the United States. Based on the organizations involved in its birth, I'd say its ethos is, broadly, about civic-minded issues. The tools mentioned vary and have different user experiences.
Enigma is a login-required, commercial offering (with a free option, at least for the time being) providing a web application interface to public data, worldwide. It is, at its core, a search engine that lets you drill down into data rows from a common user interface. Its ethos seems to be "find the data you are looking for, whatever your purpose: academic research, business analysis, civics, etc.
I helped with a tiny tiny piece (the contact-congress repo), and even that was worked on for months before by the folks at Sunlight (in particular Dan Drinkard and Eric Mill).
Edit: All snark aside though, this really is awesome. I can imagine all kinds of useful things that come out of this sort of structured data, including just interesting information (like demographic patterns of various politicians, etc).
"The depopulation of Chagossians from the Chagos Archipelago, that is, the compelled expulsion of the indigenous inhabitants of the island of Diego Garcia and the other islands of the British Indian Ocean Territory (BIOT) by the United Kingdom, at the request of the United States of America, began in 1968 and concluded on 27 April 1973 with the evacuation of Peros Banhos atoll.
...
On April 1, 2010, the British Cabinet announced the creation of the world’s largest Marine Protected Area (MPA) which consists of most of the Chagos Archipelago, homeland of the Chagossians. The MPA will prohibit extractive industry of all kinds, including commercial fishing and oil and gas exploration. Some Chagossians have claimed that this MPA was created to prevent the islanders from returning to the islands.
On December 1, 2010, a leaked US Embassy London diplomatic cable exposed British and US communications in creating the marine nature reserve. The cable relays exchanges between US Political Counselor Richard Mills and British Director of the Foreign and Commonwealth Office Colin Roberts, in which Roberts 'asserted that establishing a marine park would, in effect, put paid to resettlement claims of the archipelago’s former residents'. The cable (reference ID '09LONDON1156')[citation needed] was classified as confidential and 'no foreigners', and leaked as part of the Cablegate cache."
http://en.wikipedia.org/wiki/Depopulation_of_Chagossians_fro...
Yeah, it would be really awesome to just click a few times to see make up of a politician's district.
Hope the project does well.
My Congressional District: https://www.census.gov/mycd/
Other tools from census.gov: https://www.census.gov/data/data-tools.html
Perhaps maybe the disconnect is in that I -- and maybe op as well -- are thinking in the context of people as a whole, and you are thinking in the context of people as members HN.
https://www.google.com/search?q=congressional+district+demog...
For the majority of the (US at least) population? Absolutely yes. Hence the term "circumvention". No way this fight is won, at least at this point, over a battle of logic.
In this day and age, instigating change needs to be as simple as possible.
Maybe what I'm saying doesn't make sense. If so, apologies.
I have the same question for the United States project. Why YAML for congress-legislators? It is certainly better than creating their own custom format, but I still have to do work if I want to import the data into a database or Excel.
1. http://www2.census.gov/census_2000/datasets/
2. http://www.census.gov/data/developers/data-sets/decennial-ce...