I would be furious if my essay was graded by a computer. Is the model really good enough to account for all the variances in human language? Is the whole thing graded automatically or just the grammar?
I don't know about the specific project SoMisanthrope is talking about, but these types of tools are often used in conjunction with human graders. e.g. Instead of having 2 human graders, you automate the grading and have 1 human grader, and if the grades differ by some amount, only then do you bring in a secondary grader.
Good point DamVigilante. We trained the model using hundreds of human-scored essays. They were all double or triple scored, to validate IRR. I think that's why the model is performing so well. But, there is always room for improvement! :-)
I hear you WoodenChair. For this particular task, the stakes are very low (really, zero). The model agrees with our human graders, quite well. In fact, statistically speaking, the machine scoring is actually tightly correlated to the highest inter-rater reliability that we could achieve, among our human graders. You would be surprised.