In memory of Aaron, bulk XML of every federal and state law and court ruling
webpolicy.org
webpolicy.org
Don't get me wrong, it's better than nothing, but to get any buy-in from existing lawyers these issues need to be addressed. I'm concerned that they haven't on the site.
It's a fantastic system (and a wonderful demonstration of the advantages of immutable graph data structures). You can read cases from a century ago and still find every cited reference. Hyperlinking pales in comparison.
What we need is a technical mechanism for embedding referenced source documents into a document, in a way that is as easy as hyperlinking to add the references and follow them. Probably also a new fair-use provision in copyright law as well.
Agreed on having a way to embed documents. Especially in a way that supported some kind of signing. If I could have a reputable third party (web.archive.org or anybody else) sign an embedded snippet of a page and say "Yes, this was actually posted on X website at Y time" that would be fantastic.
The hyperlink equivalent would be "you can find this book at this library in this town, as of the time this was written".
I imagine that in real life lawyers don't need to worry about that sort of instruction because those books are printed and widely distributed across the country, while on the web documents tend to exist in only a single location, maintained by a single person or organization. USENET might be the closest on the internet we have gotten to that sort of distribution model.
The difference is that a citation is a reference to an immutable object. No matter what happens to Roe v. Wade, say it is overruled, the document at 410 U.S. 113 won't change. A URL on the other hand is a reference to a location. Like with the memory address pointed-to by a C pointer, what is at that location can change.
One can imagine a system of hyperlinks that behaved differently. In this system, a URL would uniquely identify an immutable document. A new version of a document would get a new URL, and servers would be required by the protocol to preserve all previous versions of the document. This is essentially the premise of Git: every blob is stored forever and different revisions are stored as deltas such that older versions always remain accessible.
In effect, the volumes and page numbers from the official reporter are anchors in the text -- and are stored that way in other databases (and sometimes included as textual anchors in secondary printed references.)
Most online versions I've seen do include both the citation and page numbers from the official reporters.
> If a court says (in slightly more words) "pages 205 to 207 are overruled" there are often paragraphs that span from pages 204-205 and from 207-208 that are ambiguous
US courts don't generally do that. They cite prior cases using page references, but they don't say "pages X-Y" are overruled. Its not a matter of more or less words, that's just not how they work at all. If they are reversing a lower court decision, they simply state that the decision is reversed (and if it is reversed in part, they describe which effects are reversed, which may not map to specific separable parts of the text). If they are stating a new legal rule overriding a prior precedent, they simply state the new legal rule.
> and it gets even worse when a different book version is paginated so that the range above covers part of page 621 to part of page 624.
Different books aren't the official reporter. Different books (or online sources) that are intended to be legal references will often include, as anchors in the text, the page numbers from the official reporter at the point in the text where the pages break in the official reporter.
As an example in an online source, consider the Findlaw entry for The Amistad [1]. The heading includes the reference to the official reporter (40 U.S. 518) -- 518 is the page number on which the case starts in the official reporter.
Throughout the text on Findlaw, you'll see blue notes like "[40 U.S. 518, 523]" -- these indicate points of page breaks in the official reporter (they follow the style of standard legal citation, so 40 U.S. 518, 523 marks the beginning point of page 523, in the case beginning at page 518, in volume 40 of the United States Reports.)
[1] http://caselaw.lp.findlaw.com/scripts/getcase.pl?court=US&vo...
This leads to page numbers like 247.1151a-iii. Which is then the canonical page number for a block of text.
It's enough to make the Library of Congress filing system seem simple and rational!!
1) How comprehensive are the court decisions ? For example which Federal Courts are covered and for what time periods ? If there are variations in coverage of state courts what are the high low and typical cases of coverage - both for dates and court levels ?
2) How was the court decision data obtained ? I was under the impression that there were significant obstacles to obtaining much of this data since although statutes are available freely online for many jurisdictions access to court decisions is typically very costly. I once payed several hundred dollars for a months access to NYS court decisions and I believe that service no longer exists, having been replaced by much more expensive long term plans that are out of reach of anyone except law firms or large corporations.
Getting court decisions online for free or at an affordable cost would be of great benefit to anyone needing/wanting to do legal research in the US and would help improve the increasingly dismal state of democracy in this country.
However, I can say that I extracted my state's cases and it came out to 100k+ so it's certainly more comprehensive than any other free data source (for my state)
EDIT: Downloads went down (Dropbox) right as I posted so I reached out to author to see how to get my hands on the rest of the files
FINAL EDIT: See my note below about him being on vacation and taking a look when he gets back near better internet.
Connecting to dl.dropboxusercontent.com
(dl.dropboxusercontent.com)|23.23.88.93|:443... connected.
HTTP request sent, awaiting response... 509 Bandwidth Error
2014-01-08 23:36:42 ERROR 509: Bandwidth Error.
Dropbox limits personal accounts to 20GB per day of public sharing.I'm a little surprised the moderators haven't fixed it yet.
Will update if I hear back
I just heard back and he's on vacation (go figure!).
He says he thought he had a pretty generous amount of traffic but he'll take a look when he can.
I think Akoma Ntoso would make bulk access, maybe even piecemeal API access as with other similar works like this, easier for consumption (think NLTK). The Italian Senate (the Senate in Rome) uses it, the Library of Congress has introduced some "data challenges" using it as well, and I think it is the future. Using a common data format / XML schema has its advantages.
I'm building a product and API that would take advantage of the whole collection but it's not ready for primetime yet.
But that's the kind of thing I'm working on and I'm sure others are as well.
http://www.uscourts.gov/FederalCourts/UnderstandingtheFedera... http://www.fjc.gov/
States are a grab bag but generally statistics poor.
The State Decoded takes a more Jeffersonian approach (appropriately perhaps, given that it's based in VA), allowing citizens to code up their own state statutes. PlainSite in contrast is more Hamiltonian: centralized and standardized. There are advantages and disadvantages to each approach.
1) To make this set usable from a practical point of view you have to know when it starts and finishes. "[E]very federal court ruling" is a bold statement. Federal Reporter Third? All 1000 volumes of F2d? What about the original Federal Reporter? F.Supp.? Not all federal district court decisions are published. Since our federal courts have become criminal courts (starting in the 1980's) most of the written decisions will be at the appellate level. What about "Do Not Publish" opinions? There are thousands of them and they are still useful. Usually only DoJ has copies.
2) Not having everything is not critical to the practical value of the set. In the 1990's a West salesman would tell you that there was no need to buy anything before 500 F2d if you were trying to put together a small federal library. For most states they would try to sell you everything, except perhaps New York, California and few others. The issue is updating. Florida updates (or used to) its appellate decisions on a monthly basis. You could sign up and they would send you a zip file every month. I don't know if all states do this. The problem of recency is a major one. A case could have been decided yesterday but you won't find out about it for a month. You can fix the problem on appeal--theoretically, assuming a client who wants to pay--because judges will not, except in rare cases, revisit older decisions they have made because case law that was not available at the time was dispositive.
3. The issue of citing to a particular page of a decision in addition to the official citation is not a huge problem. In many states, appellate decisions are relatively short and court rules have provided for the use of just the official citation. Cites to new Westlaw and Lexis cases do not have page numbers. When page numbers are unavailable, you can cite them as ( U.S. )(2014) [my Blue Book syntax is probably a little off here). If you cite an unpublished opinion you normally have to provide the judge and your counterparty a copy of the decision.
4. FLITE was the U.S. Air Force's effort to computerize case law in the 1980's. Westlaw and Lexis fought ferociously to prevent this database from being released to the public. They were successful. The same is the case with JURIS, a DoJ caselaw database. Now there are several providers (such as Fastcase) which compete with Westlaw and Juris. Access to PACER, the U.S. courts database of case, is limited. Efforts to mass download the database have been frustrated. The courts use PACER as a revenue tool. Also, criminal cases at the district court level are not on PACER (unless this has changed) supposedly to protect informants. So it would be interesting to know how this database was obtained.
5. Putting aside the practical value of this database, once the extent of the content is established, it could have real value for researchers. Could it be used to spot trends in the law? I wonder what might be shown if tools to measure things like historical market performance were used to analyze the database. You could see all sorts of data points for terms like "Dalkon Shield" or "asbestos" occurring within specific time ranges. There is definitely a "me too" aspect to the law. And while judges make law all the time, they have no control (usually) over the cases brought to them. Do cases involving "terrorists" match the pattern of cases involving "communists"? Or, say in the period 1910-1920, "Germans"? On a practical level, what is the statistical incidence of cases involving the Statute of Frauds? The "ancient document" exception to the hearsay rule? Are criminal conversation causes of action really coming back? If historically the incidence of data points A, then B always led to C can an analysis of such points today of any use in predicting future decisions?
Just a few thoughts.
> it could have real value for researchers.
Yes, historian researchers but not legal researchers, I'm afraid.
Umm, exactly my point. Maybe I wasn't too clear.
> federal and state law
Your entire post is about:
> case law
Suffice to say, those in the biz tend to focus mostly on case law, probably because they know the basics of statutory and common law, case law is not codified, and (all) case law is not offered for free even on horribly designed, practically useless government websites.
But, at least for my purposes, state legislation and statutory codifications (and their regulatory counterparts) are also very important to have bulk access to. Even outdated materials are useful, as once I know a section is relevant (because I can run complicated queries against entire datasets), I can begin research using more arcane methods (such as government websites and printed materials.) My treks through the California Codes and the California Code of Regulations would not have been possible without bulk access (I had to do it myself of course, after the Legislative Counsel got forced by CFAC/FAC and MAPLight.org to release the DB), and there was little case law on the issues I was doing research on (that I knew or know of) to guide me.
Outdated material may not be so useful for practicing lawyers, but its extremely useful for the 99%.