Harvard Law's Library Innovation Lab has the Caselaw Access Project— a complete collection of precedential US caselaw with structured metadata. It's at
https://case.law. It's readable online (including the original pdfs for most,) accessible via a rest API in fully structured documents, and through bulk downloads. There are OCR mistakes here and there, but the accuracy is over 99% even with weird things like the "long s" that looked like an f that was common before the 20th century. While updates are too slow to replace the commercial tools, it's perfect for uses like this... Rather, it will be in a little over 3 months. In Feb of 2018, it was released under agreement with a funder to limit access to 500 cases per day per user for 6 years except for a few jurisdictions available now— so the entire corpus will be completely open then.
They scanned, OCR'd, and applied metadata to 40k volumes, and (digitally) redacted by hand all commercial material (eg head notes, key citations) in all in-copyright volumes, so what's left is entirely in the public domain.
Disclosure: worked on that project for several years.