There is a project called DREaM, at McGill to standardize for "distance reading" (macro analysis).[1] It uses a program called VARD (a text preprocessor trained to correct spelling).[2]
Strangely, this application is licensed with the creative commons. I think this means that it is closed source. Does anyone know of any open source alternatives?
It cannot handle such an immense amount of data,[3]
[1] http://earlymodernconversions.com/introducing-dream/
[2] http://ucrel.lancs.ac.uk/vard/about/
[3] http://www.matthewmilner.name/2014/11/18/VARD-and-EEBO-TCP/