The Cambridge Law Corpus: A corpus for legal AI research
arxiv.org
arxiv.org
They scanned, OCR'd, and applied metadata to 40k volumes, and (digitally) redacted by hand all commercial material (eg head notes, key citations) in all in-copyright volumes, so what's left is entirely in the public domain.
Disclosure: worked on that project for several years.
is it complete? for example it says it is 144k cases in California, I would expect more..
In the about page there's more detailed information about the scope and process.
Unfortunately Caselaw limits access to the full text bulk data of most jurisdictions without a research account and I’m trying to find an alternative.
That's only true until Feb of 2024! It should be totally unrestricted in a few months.
Any idea why they're time limited? I assumed that was a license restriction from reporters or FastCase, et. al. which would have been permanent.
This should have more up-to-date and accurate information than I do: https://case.law/about
Please let me know how you find using the data and if you'd like to see any additions or changes!
I only just started yesterday but so far so good!
Having it in Parquet files on Git LFS makes a huge difference. It only took a few lines to add the entire dataset to our CI/CD cache which is an improvement over the ingestion scripts we have to normally write with change detection and all that. It took less than an hour to start running the cases through our pipeline - I wish all of the GovInfo bulk data were available this way!
> Applications for access to the Cambridge Law Corpus (CLC) can only be made by researchers who are employed full-time by a recognised university or other research institution. The applicant must hold a permanent position at the level of Assistant Professor (or higher) or equivalent.
So no, you will never see this unless you are a researcher. Academic gatekeeper strikes again
Assuming its consistent with fundamental justice and precedant (which shouldn't really necessarily require an advanced legal education to competently analyse and resynthesize into a case that integrates the facts of the case in question), we should value the abillity of the citizenry and all litigants in general to quickly and easily present an informed meritorious case and not rely on obfuscation or access issues rooted in technology and deficiencies of any given education system to limit participants' capabillities unduly.
Once you feed an LLM all case law, answers to questions like "...so how can I kill my family, collect life insurance on them and get away with it" become trivial for the plebes to access.
As cool as tech advances are, you really don't want everyone unlocking the secrets of estate law or nuclear fusion. Some gatekeeping is necessary.
Sometimes unlocking the secrets of something isn't enough - you need the underlying capital or power to actually pull it off. Normal individuals with normal incomes can't build reliably fusion reactors or pull off insurance fraud schemes no matter how much compute or intelligence they have. If they could, we wouldn't be living in a statist society, we'd be living in 2b2t.
“potentially sensitive nature of this material”.
It’s one thing for someone to have to dig up records through obscure legal websites. It’s another for it to show up on the front page of your Google search listing.
Okay this isn't conlaw 101 but my point is privacy and right to privacy are technical terms, possibly unintuitive to you. Now someone's rights may have been arrogate here - I don't know - but that would be again argued in a court of law.