Three hundred and sixty years of United States caselaw
case.law
case.law
This is a first public release and the Harvard Library Innovation Lab is a small team, so please look at the site as just the beginning! In particular we're starting by targeting developers with an API and bulk data downloads. More user friendly features like a front-end browser, PDF scans, ngram browser, etc. are on the roadmap, but we're hoping other developers will step in and build some great tools as well.
I and a single Rails dev modified the current Blacklight codebase to handle PDFs (and other office docs) with minimal effort a few months ago, I think doing so when you get to PDF handling and a web UI would be a completely valid starting point.
I'll save you the trouble of figuring out which is the only existing feasible open source library to sift through hundreds of thousands of PDF pages... QPDF is it.
Not at all the easiest system to work with, but it does allow you to handle a lot of structured data such that you need the minimum amount of customisation possible - which is still a lot!
Benefits of our data set: we're more complete, being a census of all known volumes of official caselaw back to the beginning; we're easier to work with for data processing, since all of our data, across centuries and states, is in one consistently structured format; we have the page images, meaning if there's any question of accuracy, we can check the final authority.
Benefits of FLP (and there may be others I don't know): they're updated in realtime from scanning court websites, so they'll stay up to date in a way we won't; their scraped text for modern cases doesn't have OCR errors; their site is much more featureful.
At this point I see our strength as being a complete/consistently-formatted/authoritative data set of printed cases, which leaves lots of room for other caselaw databases with complementary goals.
It looks like the link from the map is using "wa" for Washington, when it should be "wash".
I'm curious now. Some states do use a 2 letter code, some use 3, and some use 4. Why didn't you use the same naming format for all of them?
For the jurisdiction slugs we use the standard legal citation abbreviations for each state:
https://law.resource.org/pub/us/code/blue/IndigoBook.html#T1...
This has the advantage of matching the citations to cases, like "123 Wash. 456".
It took me 12 minutes to post because I tossed in some explanations for people not familiar with legal citations. As soon as I posted, I saw yours, and deleted mine as redundant.
There is one thing, though, that I discovered working on that post and am curious about now. I picked a case at random to use as an example, Peterson v. City of Seattle, 316 P.2d 904, 51 Wash. 2d 187 (1957).
Here's a link to the Washington Supreme Court's opinion:
https://law.justia.com/cases/washington/supreme-court/1957/3...
Within the opinion they cite the case as 51 Wn.2d 187 (1957) at the top. Inside the opinion they cite some Washington cases as Wn and some as Wash. I could see no obvious patter as to which they pick.
Anyone happen to know offhand what determines which form they use?
<clicks "New Rhyme!" button>
...
> It is expressly laid down in Bull
> I refer to the receipt in full.
> The company first.
> Decision reversed.
> Did they state they would replace your wool?
> Brown with his papers in his trunk.
> Kilburn laid down on the top bunk.
> The court reconvened.
> The State intervened.
> She thought perhaps Kristin was drunk.
Developers of this project: Please make the information freely available in a way that doesn't require agreements with a giant for-profit company.
For now, I'm convinced that this project is nothing more than a veiled advertisement for lexis nexis.
> Access limitations on full text and bulk data are a component of Harvard’s collaboration agreement with Ravel Law, Inc. (now part of Lexis-Nexis). These limitations will end, at the latest, in March of 2024.
Hopefully this means no logins, also, but that's less clear.
The blame on this should really fall on the courts that allow private companies to paywall access to the rule of law. At least some states have started to do it right:
> Once a jurisdiction transitions from print-first publishing to digital-first publishing, these limitations cease. Thus far, Illinois and Arkansas have made this important and positive shift and, as a result, all historical cases from these jurisdictions are freely available to the public without restriction. We hope many other jurisdictions will follow their example soon.
The only things we have behind logins are what we are contractually required to, yes. This gets pretty fine-grained -- if you do a logged-out search across jurisdictions, requesting full text, the json contains error fields for the specific fields we aren't allowed to share without a login yet.
(I mean, Harvard isn't a bad home for this -- I work in a building with books that predate the printing press, and I work on stuff like Creative Commons-licensed forkable textbooks. Libraries are cool places. But Harvard definitely shouldn't be the only place that preserves this data set.)
As far as preservation-friendly formats, our bulk data download format is xzipped jsonlines, which is tuned for NLP (highly compressed, parseable in a few lines of python with low memory requirements) rather than preservation:
https://case.law/bulk/download/
Internally we have a preservation format where each volume is stored as a bagit bag containing METS XML for the OCR and case-level data, plus color and black and white images of each case. These are much harder to work with, so it's not a focus to share them right now, but we can definitely share if someone makes a case for it.
Why isn't this data available for free to any American Citizen who pays taxes? In the form of a torrent, the distribution cost is negligible. A reasonable duplication fee seemed reasonable back when replicating vast amounts of information involved massive amounts of paper and toner. But today.. I don't understand.
I'm curious enough to have asked this same question on Quora [0].
YMMV.
[0] https://www.quora.com/unanswered/How-can-LexisNexis-own-the-...
But that doesn't mean that private entities who via their own labor, digitize it, are required to give your their resulting digital datasets.
All the more reason to be thankful that these folk are doing exactly that, and giving us tools to get at the data.
For this project we had to scan 40,000 volumes of caselaw. We used a high speed scanner at the Harvard Law Library, and went through about 40 million pages at a rate of 500,000 pages a week over a couple of years. The pages then had to be redacted of copyrighted material like headnotes inserted by private publishers, since courts typically don't publish the cases themselves, and those redactions had to be checked by humans.
That work was funded by a startup, Ravel, which is why we ended up with temporary limits on commercial use of the data. No later than March 2024, however, it will all be fully available for bulk download by anyone in the world. If necessary we'll set up a torrent. :)
(Hopefully earlier! For any state that starts officially publishing its caselaw in digital form, we can immediately release their caselaw back to the beginning, as we have already for Illinois and Arkansas.)
To answer your question re: its origins being publicly funded, no. When there was no internet database to connect to people bought the books from a print publisher (either at huge maintenance cost or sparingly at the expense of not knowing what was current). The print publishers bore the cost and reaped the profits of consolidating all of this data.
To answer the next obvious question: yes, there are probably people in prison or not depending on whether a small town law library bought the updates from West Law in a timely fashion 20 years ago.
Yes and no. It’s important to realize that the US courts (1) are a distributed system comprising hundreds of autonomous courts; and (2) predate the internet, photocopiers, telephones, telegraph, a large centralized federal government, and indeed the federal government itself.
Today court opinions are published as PDFs on courts’ websites. But back in the day, they were published as slip opinions stored in the clerk’ office of each individual court. Private companies like West undertook to collect cases from all these hundreds of courts and publish them in books called “reporters.” Back then (and even today) that meant sending someone out to hundreds of courts to collect and copy the decisions. They not only published the opinions, they organized everything within a comprehensive ontology of their own creation, and added their own annotations.
When computers were invented a century later, these publishers were well placed to digitize their collections and offer access to them over pre-internet electronic systems. Then, of course, those moved to the Internet.
Even collecting these cases together on a going forward basis is no easy task. As noted above, the courts are decentralized, by design, even within the federal system. Just getting the decisions from hundreds of courts and uploading them would be an expensive endeavor. Nothing stops someone from undertaking this—court decisions themselves cannot be copyrighted and you’re free to go to and court and ask to copy published decisions.
TLDR: 500 cases per day, but looks like you can buy access from Ravel [1].
API appears powerful, but lots of free alternatives that don’t require a comfort level with shell scripting or REST calls.
I also really like casemine.com.
I have no idea how incomplete those archives are but I doubt they’re missing much I’d be looking for.
EDIT: casemine appears to be neither free, nor to have all published US decisions as far as I can tell.
If you’re just looking for a nice case browsing interface, in the interim, you should check out Ravel’s site.
German and French law is already there.
https://api.case.law/v1/cases/?jurisdiction=ill&full_case=tr...
Full text for cases other than Illinois and Arkansas is limited to 500 cases per person per day, which is why we don't include it in API results by default.
https://api.case.law/v1/cases/
State courts didn't come into existence with the US Constitution -- the Massachusetts Supreme Judicial Court, for example, dates back to 1692, and those precedents are still "good" in some sense though unlikely to be cited.
We don't have English precedents, unfortunately, as guessed by some sibling comments.
https://api.case.law/v1/cases/?search=coleslaw&full_case=tru...
Apologies that the link requires a login to view the full text, and for various other shortcomings in the current browsing experience. Also apologies to anyone who actually reads the coleslaw case -- caselaw is a scary place.