Harvard Law Library sacrifices trove of legal volumes to digitize them
nytimes.com
nytimes.com
Here's how big this is: we don't even know yet how many cases we'll end up with, to within the nearest million.
PSA: we're hiring a devops engineer[1]. In addition to building amazing tools to access all this data, we're running a distributed linkrot preservation service[2] that after just two years is in use by 40% of American law schools and 10% of state supreme courts; an open-textbook-as-forkable-playlist[3] tool in use at Harvard Law and a half dozen other law schools; and a research project on distributed encrypted library archives[4] for preserving high-value cultural records. We're basically the alien in the brain of a 200-year-old library -- it's a fun place to work.
[1] http://librarylab.law.harvard.edu/blog/2015/10/20/hiring-dev... [2] http://perma.cc [3] http://librarylab.law.harvard.edu/projects/h2o [4] http://librarylab.law.harvard.edu/projects/time-capsule-encr...
I am an English teacher, and about 80% of my students are professional lawyers, so I wonder how free I will be to use this material in my classes. I already use Harvard Law School's free case studies, they're great, and under CC license if I remember correctly.
For academic researchers, before the eight years are up, we can also provide a full data dump -- you just have to sign an agreement not to redistribute bulk data.
There's a whole separate problem of turning those scans into a high-quality data set. The first pass will be decent-but-not-perfect OCR of the full text (with page-location data, like Google Books), plus human-checked metadata for stuff like case name, judge, and date. Since we have the original scans as well, there's lots of room to iteratively improve the data conversion from there via ReCAPTCHA and the like.
This isn't there yet, but I wonder if that's where we're heading...
There will be some loss, true. Even where everything is properly photoed, the programs will make some mismatches. Potentially, the error rate can be less than a few words per million volumes, far better than even hardcopy republishing with manual copyediting. -- Sharif, Rainbow's End p 129
But of course this is fiction. The parallels to what Harvard are doing are obvious, but I don't think anyone would seriously suggest building the Libreanome project in the real world.
What's the hurry?
Please don't editorialise submission titles [1]. The title of the NYT article is "Harvard Law Library Readies Trove of Decisions for Digital Age".
[1] FWIW, I agree with the sentiment of loss implied by "sacrifices". On a worse day than this I might even think of these books as being mutilated.
Moreover, it used to be common to create bound volumes by binding multiple issues of a serial together. (In some cases the bindery would crop pages to fit!) The binding is often so tight that the volume cannot open flat enough for a full scan. Separating the pages from the spine allows for the entire page to be imaged without distortion.
Edit: Incorporated correction. I had accidentally stated the claim more strongly than the FAQ supported; however, my point does not depend on the claim being as strong as the form in which I had stated it.
Prototype 1 could scan the majority of books without
damage, but may tear one or two pages in some books. Out
of 50 books tested, 45% had one or two of their pages
either torn or folded. This is a very early prototype and
there are many areas for improvement in the design.
http://linearbookscanner.org/faq/That being said, they are probably throwing away an opportunity to sell the books as novelty or decorative items. But who knows, maybe the cost-benefit is just not worth it.
The rule of thumb apparently is, best case, books you can't cut the binding off cost 5-8x as much to scan as books you can.
It also makes cases easier for to argue by mapping out the corner cases of the law and legal concepts and keeps every lawyer from having to be an absolute expert in every facet of every case they're litigating. The same applies to judges.
Precedents do have a limited lifetime but it's not measured in years, if someone is being tried under the same law as existed 100 years ago old rulings can be just as relevant as one made last year.
This isn't always true. Many of the fundamentals of civil engineering, mechanical engineering, mathematics, power engineering, chemistry, physics, even computer science, are not supplanted by more recent texts. In fact, when I was a computer science and math student, I lamented many of the modern texts (math and physics ones being the most egregious) for their overly fluffy delivery, compared to my father's and grandfather's and mother's math textbooks from (at the time) 20 and 60 years earlier. I would use theirs to study because they were much better at conveying the topics than the recent calculus or statistics or whatever book we were given.
Now, in industry, and in the software industry in particular, this is more applicable. So much is changing so fast, but the basic concepts (functions, types, network theory, latency versus throughput, even multi-process computing) are the same, but the technologies we use and how they present or make use of these fundamentals change rapidly so my 199x book on javascript is only barely applicable to what we have today.