Thousands of early English books released online to public by Bodleian Library
bodleian.ox.ac.uk
bodleian.ox.ac.uk
That first site is remarkably poorly designed for actually finding the information, a common theme I find for websites created by librarians. There should just be a box at the top listing the download links!
Web developers working for libraries: Far too often I visit a site and and am confronted by acres of text explaining what the project is, who is involved, how to enter data in forms and all sorts of hand holding, BUT NOT THE ACTUAL DATA! Usually I find there's some link hidden in the least visible part of the page, like the lower right hand side, that actually lets me get started! Does Facebook have paragraphs of text with a welcome message from Mark, explanations of what Facebook is, how you use it, or does it just let you dive in and get started?
edit: And despite all that text it doesn't explain what phase I and II are.
I should say this is an amazing bit of work and it's really important that it's being released public domain, and a good sign of the direction libraries are taking. It's just a little frustrating that the final delivery step is so obfuscated.
There is a project called DREaM, at McGill to standardize for "distance reading" (macro analysis).[1] It uses a program called VARD (a text preprocessor trained to correct spelling).[2]
Strangely, this application is licensed with the creative commons. I think this means that it is closed source. Does anyone know of any open source alternatives?
It cannot handle such an immense amount of data,[3]
[1] http://earlymodernconversions.com/introducing-dream/
[2] http://ucrel.lancs.ac.uk/vard/about/
[3] http://www.matthewmilner.name/2014/11/18/VARD-and-EEBO-TCP/
Licenses that don't allow modifications are closed-source. At least, they are inconsistent with the Open Source Definition (specifically, with criteria #3: "The license must allow modifications and derived works, and must allow them to be distributed under the same terms as the license of the original software.")
That's why I think it's strange it is licensed CC.
Files are hosted by github and box. Will Internet Archive be included?
For marketing this accomplishment, a few choice examples of long-inaccessible text may attract new readers.
!!! This is not silicon valley. I wonder how they ensure accuracy.
Link to the books http://ota.ox.ac.uk/tcp/
> TO THE RIGHT VVORSHIPFVLL MAISTER RObert Clarke,
The mistakes look like typical OCR errors.
[1]: http://tei.it.ox.ac.uk/tcp/Texts-HTML/free/A01/A01716.html
It's kind of interesting that they look like the same errors as those generated by OCR.
The difficulty of deciphering the text makes this huge task even more impressive!
Perhaps there is a specialist antiquarian OCR package which can deal with long s, interchangeable u and v, non-standardised spelling, etc, but I have yet to come across one.