They scanned, OCR'd, and applied metadata to 40k volumes, and (digitally) redacted by hand all commercial material (eg head notes, key citations) in all in-copyright volumes, so what's left is entirely in the public domain.
Disclosure: worked on that project for several years.
Unfortunately Caselaw limits access to the full text bulk data of most jurisdictions without a research account and I’m trying to find an alternative.
Please let me know how you find using the data and if you'd like to see any additions or changes!
I only just started yesterday but so far so good!
Having it in Parquet files on Git LFS makes a huge difference. It only took a few lines to add the entire dataset to our CI/CD cache which is an improvement over the ingestion scripts we have to normally write with change detection and all that. It took less than an hour to start running the cases through our pipeline - I wish all of the GovInfo bulk data were available this way!
That's only true until Feb of 2024! It should be totally unrestricted in a few months.
Any idea why they're time limited? I assumed that was a license restriction from reporters or FastCase, et. al. which would have been permanent.
This should have more up-to-date and accurate information than I do: https://case.law/about
is it complete? for example it says it is 144k cases in California, I would expect more..
In the about page there's more detailed information about the scope and process.
> Applications for access to the Cambridge Law Corpus (CLC) can only be made by researchers who are employed full-time by a recognised university or other research institution. The applicant must hold a permanent position at the level of Assistant Professor (or higher) or equivalent.