How we found a million style and grammar errors in the English Wikipedia
fosdem.org
fosdem.org
> So we manually looked at 200 of the errors, finding that 29 of the 200 errors were real errors. Projected to the whole Wikipedia (currently at 4.3 million articles), that's about 1.1 million real errors
Projections are fine, but maybe the title should be something like "How we found (probably) a million style and grammar errors in the English Wikipedia".
Would you have clicked if that were the title?
Sad part is that I suspect many people upvoted based on the title alone, without reading the article
If you work for a magazine publisher, it's likely you have 2 lines, a web line and a published line. The web line might not go through the copyeditors/production/subeditors before it goes online whereas the published line will likely go through those people/departments. By adding this into the web line, it can help enfore the rules of the publication so that web and magazine both match up nicely.
Imagine a pre-commit hook for writing quality test ;)
Here are some other scripts (much more basic) for "automated language tests": https://github.com/ivanistheone/writing_scripts
And, for an article on finding grammatical errors, there were a surprising number of grammatical errors.
I'll take a wikipedia that has these kinds of errors over a time before wikipedia. As a boy going to the library and immersing myself in encyclopedias and world maps I am gobsmacked at how wonderful great resources like wikipedia, freebase and google maps are.