Flawed Algorithms Are Grading Millions of Students’ Essays
vice.com
vice.com
Yes, education takes time and costs money. Yes, not educating is both cheaper and faster. Note how the rationalizing ignores the needs of the students and the quality of the education.
I live in Utah and my children have been subjected to this automated essay scoring here. One night I came home from work and my son and wife were both in tears, frustrated with each other and frustrated with the essay scoring which refused to give a high enough score to meet what the teacher said was required, no matter how good the essay was. My wife wrote versions herself from scratch and couldn’t get the required score. When I got involved, I did the same with the same results.
Turns out the instructions said the essay would be scored on verbal efficiency; getting the point across clearly with the fewest words. I started playing around and realized that the more words I added, the higher the score, whether they were relevant or grammatical or not. Random unrelated sentences pasted in the middle would increase the score. We found a letter of petition online for banning automated scoring for the purposes of grades or student evaluation of any kind. It was very long, so it got a perfect score. I encouraged my son to submit it, and he did. Later I visited his teacher to explain and to urge her to not use automated scoring. She listened and then told me about how much time it saves and how fast students get feedback. :/
I could perhaps see value in having unlimited tries, as a teaching aid, if the result wasn’t being used for grading. That would at least leave room for curiosity and exploration. And, more importantly, I could see value if the software wasn’t essentially a scam that fundamentally is not able do what is advertised. If the software really could grade essays reliably, and provide meaningful suggestions for improvement, then maybe it could be used to help educate students, in conjunction with the teacher’s guidance. But the software does not grade reliably, and it absolutely does not offer meaningful constructive feedback, and the teachers were using it to avoid reading essays, not to supplement their own expertise.
One of the several amusing ironies here is how the software company has convinced the state and teachers to willingly replace themselves with bots, despite obvious evidence that the humans can do the job better.
Teachers are extremely overworked, underpaid, and underappreciated. I'm not surprised that it was easy to convince them to offload the difficult and time-consuming work of manually grading essays. This also means they don't have to deal with complaints about unfair grading. A machine did it and it's out of their hands.
Note: This is not hyperbole, I have seen this exact scenario more than once.
There may be a place for a well-designed one, but if it exists, I've never seen it.
if only the last one is not required of the scoring system, then it's been available for ages, and its just a really poor implementation. That's not programming but plumbing data-pipes.
Having an ignorant teacher is almost as bad as a flawed, black-box algorithm.
Math is only not commutative in very rare areas. A specific area is concatenation of 2 strings, or 2 rotations of a rubiks cube.
> said it was impossible to calculate the exact area under a curve
And this is why a lot of people are against unions and similar protections for teachers. How do you get rid of someone who blatantly lies and informs falsely to students? How do you get rid of horrible teachers?
The state chose fast and cheap. Well, its cheaper than more teachers.
But even with these examples, the path of appeal and rectification of mistakes is much easier with all humans involved. I fear soon people will side with the machine out of ignorance or to be justified in an incorrect stance.
The idea that we could be so poorly taught by broken automated systems, that we become incapable of detecting the system is broken seems like a possibility with AI that is much less likely in pure human systems of education (though not impossible).
It would mark you as incorrect for using too many decimal places, even though it wouldn't tell you how many significant figures was required. I often remember it marking my answer as incorrect, even though it was identical to the answer they gave. Sometimes you'd have to show your working, but it couldn't handle brackets. Once I put the answer as "1+x=y" but they wanted the answer "y-1=x", and they marked it as incorrect.
I'm sure academic software design is leaps and bounds above what it was in the early 2000's, but to have a pupils futures hinge on what generally seems to be poorly tested code is dangerous.
It's also a bundle with the book sold by unis, because the code 'allows' you to submit homework required for the class. So they're doing both resale prevention, AND horrible grading.
People I know who have to interact with it call it "My Meth Lab" - because you have to be high to like it.
edit: ...just like how previous generations have to learn how to game social systems.
When such changes occur, the gamer will be docked until they can reverse-engineer the new algorithm. There's also the risk that all their previous inputs "gaming" the system might be reconsumed to terrible results as well, effectively rewriting their historical performance disasterously.
As always, those with the social standing and power to have insider knowledge or guidance will be in the best position to profit off such systems.
Every essay I wrote, I'd always force myself to reach the max "12.0" grade level. While writing I'd struggle over word choice, sentence structure, rearranging paragraphs, working on my tone etc, all in pursuit of the 12th grade way to phrase things. All my revisions were subject to the approval of the Grade Level checker.
Whenever I could I would check the grade levels of my friend's writing - usually by showing them a "neat feature" they could enable. Then, I'd smugly applaud myself for being the better writer whenever their grade level was below 12.0.
The Grade Level feature fascinated me and to try and master it, I found a book about Microsoft Word and looked through it in a bookstore. I was absolutely gobsmacked at how simple the formulas was. I had childishly been expecting something, like perhaps Utah educators imagine they have. I genuinely expected the method to be complex beyond my understanding.
Instead, Word used a variant of Flesch-Kincaid. There was a direct relationship between sentence length and grade score, and polysyllabic words and grade score. Meaning, the longer your sentences and words, the higher your grade score.
As soon as I got home from the bookstore I loaded a draft of something I had written. It was "pre-12.0" writing from me. I simply deleted all the periods but one and checked again. 12.0.
Automatic grading is a wonderful lure. It's nice to imagine that there's some objective writing quality easy to tap into. At the moment, I think we're far from that ability.
Personally, I feel the solution to insufficient teacher time is to use peer grading much more, and spot checks. Get kids to read and revise each other's works frequently, and teachers should aim to grade at least N papers per student where N is much less than the number of papers a student writes.
Revising is a really vital part of writing. Getting more chances to do revision, plus having to write something good enough to show your peers, plus having the risk of any paper count for your grade should compensate for incomplete teacher grading.
The point I am trying to make in agreement with the parent is: there are qualities that are very hard to score with algorithms. The difficulty of solving this problem equals if not exceeds that of automated translation, which still only works properly for specialized and limited domains, e.g. weather forecasts.
That's how it's done in creative writing courses. I've always found it infinitely more helpful than only having feedback from the instructor, even if the instructor's feedback was generally more helpful/useful than peer feedback.
Not a reflection of physical reality like sensor data or even accounting information, but the method of communication explicitly invented for production and consumption by humans.
Of course feedback from humans is more valuable than feedback computers, it would be irrational/miraculous if anything was better at giving feedback than a human.
It is a shame it isn't self evident to instructors how poor of a solution this is, and how much better the results are when using critique by peers and instructors -- the classic way of doing things.
Who the hell came out with such an idea. I would even hesitate to use "AI" for automatic spell checking as it is sufficient to give some character unusual name and it will be marked as error.
My guess is that soon or later people will learn how to game that AI. I wouldn't be surprised if there were some software that will generate essay that Utah "AI" likes.
Already been done. http://lesperelman.com/writing-assessment-robo-grading/babel...
Here's a sample essay that is complete nonsense and got a perfect score on the GRE.
http://lesperelman.com/wp-content/uploads/2015/12/6-6_ScoreI...
"Calling has not, and undoubtedly never will be aggravating in the way we encounter mortification but delineate the reprimand that should be inclination. Nonetheless, armed with the knowledge that the analysis augurs stealth with propagandists, almost all of the utterances on my authorization journey. Since sanctions are performed at knowledge, a quantity of vocation can be more gaudily inspected. Knowledge will always be a part of society.Vocation is the most presumptuously perilous assassination of mankind."
Yet the robo-scoring acclaims it as:
* articulates a clear and insightful position on the issue in accordance with the assigned task * develops the position fully with compelling reasons and/or persuasive examples * sustains a well-focused, well-organized analysis, connecting ideas logically * conveys ideas fluently and precisely, using effective vocabulary and sentence variety * demonstrates superior facility with the conventions of standard written English (i.e., grammar, usage, and mechanics) but may have minor errors
Any teacher faced with the requirement to use such tools would be better placed instructing their class on civil disobedience.
There's 2 ways of finding out these artifacts of AI essay grading: pure luck, and being able to afford extensive test-prep (rich).
The luck one can't be accounted for. So I im lead to believe that the purpose of these essays and their AI grading is to find and escalate rich people.
Well, of course. How many poor people are allowed to decide what is good for children's education?
'There's a reason why they're poor. Better pull themselves up by the bootstraps."
Mixed alongside with poverty stricken neighborhoods are the primary funding, resulting in poor school systems. And those students obviously wont have the money or the access to get the test-prep needed to "succeed".
It's all too laid out to be accidental.
I suspect that, in addition to the scoring rules being written for speed and frugality, they were shaped by a poorly thought out attempt to make the scoring 'objective', and independent of the scorer's beliefs, attitudes and unconscious biases.
In one sense, this software (I will not call it 'AI') is an extension of all those bad ideas, only greatly amplified in a way that only software can.
In evolutionary biology, there is the concept of 'honest signalling'[2], a true and unfakable indicator, to a potential mate or predator, of an animal's fitness. That is what we are missing here.
[1] https://www.bostonglobe.com/opinion/2014/03/13/the-man-who-k...
At the time the student takes the test, he should be prompted with the informed choice asking him to grant either 1) no license to keep a copy for training purpouses 2) a non-exclusive license, and the website where he can get a copy of his own essay 3) a public domain license, again with the relevant domain linked so he can find his own and other's essays. 4) any of the above or other as a function of the resulting grade!
At the same time he should also specify his desire for or against attribution, again probably best as a function of the resulting grade. And under what moniker he wishes this contribution to exist.
These options to be filled out during exam time should have no default options (no opt-out), and preferably should be standardized by the community and lobbied for to be mandatorily enforced at state or federal level (forcing examinations to present the student with an informed choice)
A public dataset of legally obtained essays (without scores or names) would already be a very important first step to invite others to make actual performant ML grading systems.
I don't believe the current datasets in these "non-profit" organizations actually comply with copyright law, organizations who don't profit from the grading service towards the state, but do provide a stable ML job on the income from charging the people with financial means to test submissions, enabling a stealth class based society.
I'd guess this is a product of dwindling state finances and contempt for any form of real education. AI's are orders of magnitude cheaper than real teachers. They also don't form unions and wouldn't voice any opposition against changes in the curiculum.
They are also pretty useless, as you have pointed out. The consequences of this policy will be postponed until the students reach a certain age -- that'll be like 10-15 years in the future.
Granted I already have a very fatalistic view on my future already so take my opinion with a grain of salt.
There is a compelling case for immigration for that sake alone.
I do not know if it will be sustainable, but I doubt it will be static.
[0] https://www.nytimes.com/2012/04/23/education/robo-readers-us...
To be fair, the GP here is specifically describing that he gamed the AI via a copy-paste of a critique of the AI, his kid submitted it on their own accord, it was graded without comment, and then when the GP went into comment on the gaming of the AI, the teacher not only did not care that the AI was gamed, but expressed gratitude for the AI saving hours of work, still ignoring that the AI fundamentally made things worse, all at the expense of the entire point of being a teacher in the first place.
The issue, for the teacher, is that in 'the system' in which they collect a pay-check, the AI works flawlessly. The point, for the teacher, is not to educate children. It is to have assignments that children pass with some sort of distribution that can be sent in and calculated by some person in a beige suit, wide tie, and hair troubles. The difference is subtle at first, but when you get further along to the point where the GP is sitting, then the difference is comical.
The AI allows the teacher to increase their effiency in processing assignments, ones that never really mattered to the teacher in the first place. In valley-speak: the incentives are not aligned.
It is frankly a sign of a diseased culture to use it in any capacity except an exercise to improve AI.
Frankly, this does not change anything from my experience in school decades ago. The teachers always said that the length does not matter and we should not pad the papers. However students who wrote more pages got better scores every single time.
It certainly could be a problem if the prompt was too narrow, or time constraints, or some other factor.
Do you think the correlation by itself indicates something negative/inefficient I'm missing?
Apart from the fact that your story is straight up frightening, isn't this part completely backwards, too? I mean, clearly using more words to convey the same message is /less/ efficient, not more so?
That is just mad.
Kids, forget everything you know because crime does indeed pay off. Best grades will be reserved for those that try to cheat this system however it is implemented. Botting your essays is the way to go in the 21st century.
idiocracy will happen not because people get any stupider, but because the bots reward the stupid ones first.
My best Utah education anecdote - In the first day of British literature class the teacher came in and asked "Does anyone here know what A.D means?" someone said After Death - she said no. I figured this was my time to shine so I raised my eager hand and said "Anno Domini, in the year of the lord" - she said no.
Then she announced: "A.D means after the Deluge, and B.C means before Christ".
She also totally lied to me one time about whether she would be considering a particular textbook question as applying to Rosencrantz or Guildenstern.
Anyway I think that was one of the many classes I got an F in after stopped going and would walk past it every day on my way to play chess with my German teacher.
What’s even cheaper than AI? Tell the students to write some pages, have the teacher glance at the number of pages written, give full credit if the mark was met, and throw the papers out without reading them. It sounds like it would be similarly effective and less aggravating. Unfortunately, this would require humility on the part of the educators.
One question she had to grade was essentially, "What's something you want your teacher to know about you?"
It was an essay answer, and she was supposed to grade it on grammar, etc. Just the mechanical aspects of writing. (The real question explained the details more, but that was the core of the question.)
She saw answers that would make you weep.
"My daddy touches me."
"I haven't eaten today. I don't know when I'm going to eat again."
Stuff like that.
And my mother was going to be the only human who ever saw their responses. Their teacher had no chance to see their responses, just my mom.
So she goes to her supervisor and asks, "What can we do to help these kids?"
The supervisor said there was nothing you can do. Just grade the answers.
https://www.childwelfare.gov/topics/systemwide/laws-policies...
So do what? Contact her local police?
With a written accusation from a child? Is that enough to get a warrant to force the company to release the demographic information?
And people don't work at a job like that because they want to. They work there because they need the money.
Everything she took in and out of there was monitored, too. So it's not like she can go to the Xerox, and walk out of there with a copy.
It's beyond dehumanizing. For everyone. The kid, the people who work there.
> With a written accusation from a child? Is that enough to get a warrant to force the company to release the demographic information?
Absolutely! My girlfriend works as a counselor at a school and she is required by law to report all serious abuses by parents.
Collect or photograph all the evidence, record every conversation with supervisors, escalate as much as possible internally, then contact local police, and at the same time go to the media. Don't quit, but if necessary let them fire you and then sue. None of this is easy.
> With a written accusation from a child? Is that enough to get a warrant to force the company to release the demographic information?
Yes, she should have absolutely went to the local police. A child's first hand account in writing of child abuse and neglect is slam dunk evidence to secure a warrant to link the essay ID to the individual child.
> Everything she took in and out of there was monitored, too. So it's not like she can go to the Xerox, and walk out of there with a copy.
Doesn't matter. She could have went to the police herself as a witness. That alone would be enough for a probable cause warrant to retrieve the essays.
It is very sad she saw these signs of abuse and did not report it.
I don't find it adds much here.
So how many are true and how many false? I have no clue. Literally none. And no it doesn't make me feel any better about the screams of existential agony even if that were a low percentage. Could be high too.
In the US, school funding is based upon standardized test results, and bad results can shut a poorly performing school down.
It's drilled into every kid's head that these tests are very important, super strict and if they accidentally mess up, it can ruin their academics, because retesting and regrading are expensive.
I thought the state was holding the school hostage, threatening to cut funds or shut them down if they ever stopped. We never learned anything about civics or American history. Until I was out of highschool, content regarding atrocities like slavery and the trail of tears was not on the test and that was enough to whitewash the whole curriculum.
Standardized testing is to the U.S. what lead waterpipes were to the roman empire.
Food Security Status of U.S. Households with Children in 2017 Among U.S. households with children under age 18:
84.3 percent were food secure in 2017. In 8.0 percent of households with children, only adults were food insecure. Both children and adults were food insecure in 7.7 percent of households with children (2.9 million households). Although children are usually protected from substantial reductions in food intake even in households with very low food security, nevertheless, in about 0.7 percent of households with children (250,000 households), one or more child also experienced reduced food intake and disrupted eating patterns at some time during the year.
..which means students wrote whatever the hell we wanted. I was assigned a Captain Morgan (rum) ad. I wrote that the ad was glorifying maritime piracy and was likely responsible for pirate activity in Somalia.
As a child you're really not prepared for the concept that your parents are treating you badly. So that realization doesn't come until much later.
Faculty, administrators, athletics staff, or other employees and volunteers at institutions of higher learning, including public and private colleges and universities and vocational and technical schools (11 States).
https://www.childwelfare.gov/topics/systemwide/laws-policies...
https://www.childwelfare.gov/pubPDFs/manda.pdf
This includes penalties for failure to report in multiple states:
https://www.childwelfare.gov/topics/systemwide/laws-policies...
I'd believe that ML could spot abuse that humans miss pretty well from signals like non-overt references in homework and school records, if one could come up with an adequate training set.
Much more likely than teaching ML to score reasoned and creative activity in any reasonable way.
But this type of thing seems like the exact kind of spooky correlation that ML is good at spotting compared to humans.
"Spooky" machine learning results happen when a correlation is abundant in a dataset [2]. Otherwise, machine learning techniques will probably miss it altogether.
______________
[1] Quick online search: https://www.inquirer.com/philly/blogs/healthy_kids/What-is-t...
[2] The archetypal spooky machine learning story is surely the one about Target sending baby item coupos to a girl in high school before her father knowing she was pregnant:
https://www.forbes.com/sites/kashmirhill/2012/02/16/how-targ...
The total incidence of child abuse of all types from infancy to adulthood is on the order of 1 in 3. This is not terrifically rare-- it's of higher prevalence than pregnancy and of positive screening events.
A much bigger concern is non-causative correlations. It'd be pretty easy to train ML to be racist or look for e.g. indicators of class, which are correlates of abuse.
As to false positive rates-- you can pick your false positive rate to be whatever you want it to be, by twiddling the threshold for a positive result. I'm not sure false positives are of that great of a concern, if the output from a system is a notification to school administrators that they may want to keep an eye out for this student.
Child sexual abuse isn't extremely rare and familial abuse is a very large minority of child sex abuse.
How? Particularly, where do you get training data at the required scale?
You survey those kids in adulthood about whether and how they were subject to abuse and other types of relevant adversity.
You attempt to control the data so that you don't just latch onto other correlates of abuse (e.g. social class).
The people whose lives are ruined by being mis-identified by the system.
The positives must be evaluated by a human anyway.
Those same people whose lack of competence people are bemoaning throughout these comments.
When a child writes "daddy touches me between the legs" in an essay, it doesn't matter if a human spots it or an AI that forwards it to a human, this needs to be investigated either way.
> Those same people whose lack of competence people are bemoaning throughout these comments.
It's not a lack of competence that's bemoaned, it's a massive amount of understaffing (and resulting overwork) in teachers and other school resources, as well as a drastic lack of financing because it's easy to cut budgets for schools for politicians as the effects only show up two decades afterwards.
There were some cases in the UK about a decade ago where bugs in software the Royal Mail was using led to incorrect accusations of fraud. People actually went to jail over this, it took years to resolve.
When a child writes a set of things that individually are not very concerning, they may have cues that could say "hey, this kid, you should maybe keep an eye out for evidence of abuse."
Particularly attuned, experienced individuals might spot these cumulative cues, but we all know that this is not all people dealing with children.
It's an interesting problem.
Society's bigotry is going to flood that bad boy so quick you might as well name it Gobbels.
I love ML. I want children to be safe. This is not the place for ML or AI or Quantum or any tech.
What needs to exist is better resources for those children, that mother grading the tests, the teachers of those children, and social services that are meant to support them. If you want to make a difference about this, look there.
Don't go building a automaton King Solomon who decides why this kid should be taken from these parents because speaking Spanish was worth -0.1 on some goddamn weight trained on data generated from a racist society.
This isn't a "spooky" correlation a cool algorithm can detect, it's a serious, layered social problem.
Totally what I advocated for and not a strawman attack /s. Indeed, the chance that such an algorithm could be racist or classist and there being needs to avoid bad correlations and have appropriate controls is important.
I think there are opportunities here. Ideally ed-tech doesn't take humans out of the loop, but asks schoolteachers and administrators questions like, "Hey, are you sure students A, B, and C are being supported correctly for subject Z? Are you sure students D, E doesn't have some kind of abuse or other significant home problem? It sure looks like student F is in this subpopulation that research shows benefits from educational intervention Y. You might want to keep your eye out for that."
And then the teacher goes "Oh, crap. Now that I think about D, there were always these little things 1, 2, and 3 that seemed off... maybe this is worth a referral to social services to check on what's up."
Or "Oh, ... maybe F's struggles in reading really are a speech problem and we should handle that"
It would be interesting to know if a child psychiatrist could be held liable if incompetence prevented them from seeing obvious sign of abuse, but I doubt that is covered under the cited law above.
A thing you quickly find if you try to download off-the-shelf NLP tools and apply them to anything is how little is reliable at all, unless you can constrain the domain. Even basic topic identification only works with low error rates when constrained to something like NYT stories, or PubMed abstracts, not arbitrary text by arbitrary writers. And I would bet ETS is using worse tech than research state-of-the-art.
I still agree that automatic essay grading is beyond the reach of SOTA NLP models today, but youmake it sound like virtually nothing can be done in a production-grade manner that solves real world unconstrained NLP problems. This is manifestly false.
I'm not aware of any meta-analyses myself. I have been keeping up with the ASAP competition and various attempts to improve on the initial systems for a number of years. The two papers I believe are having the most success are [1] and [2]. [3] seems promising for balancing the opposing forces of high accuracy for true positives and the risk of false positives via adversarially crafted inputs.
I'm also vaguely aware of research happening around extracting features from neural nets. I'd love to be able to help students understand why the system is predicting a particular score.
[1] https://www.aclweb.org/anthology/D16-1193 [2] https://arxiv.org/pdf/1606.04289.pdf [3] https://arxiv.org/pdf/1804.06898.pdf
People making big decisions with a lot of money around computing know nothing about it and are marks for con-artists. Think big consulting firms selling to senior public servants in washington. "For a successful technology reality must take precedence of public relations." But reality just gets in the way when conning a mark for a successful snake oil sale, right?
Call it out, publically, cite your credentials. Encourage colleagues, your competition and everyone with a clue to pour scorn on whoever is selling this evil, toxic waste as drinkable.
Like this kind of thing should be cool, not insane. I mean wasn't it cool in your AI class when you learned that DFS could play Mario if you structured the search space right?
It took me much longer to pass my English language O level (exam taken at 16).
It is fine to play with "cool" techniques when you are doing consequence free stuff like playing Mario. When you are creating systems that have significant and long term effects of people's lives a different standard applies.
* 5% - cool, we could make a company that grades essays
* 15% - cool, we could make a company that grades essays and sell our source code to the test-prep industry
* 80% - fascinating, it sounds like the exam designers need to reevaluate what they are trying to measure with essay questions
And if you succeed you will simply be measuring an uninteresting but manageable subset of the problem which will then become in some people's eyes the definition of the problem.
Education is supposed to be about teaching people to think, to give them the tools with which to do it, to be able to evaluate, criticise, invent, etc.
I'd find it highly interesting to see what kind of result I'd get using an automated system.
Why?
Because, I once asked a teacher (also an examiner) why I got good grades above the others, and the answer surprised me: my answers were generally unique /refreshingly different, to the point/ not too long and easy to read.
I suspect with this new system, I'd be an average student. It'd also be interesting to find out, several years down the road, if the automated system could be gamed at all -- I suspect it could, and teachers would help students 'maximise' their scores as a result of that.
You know, those average writers like Hemingway /s
In fact, throughout the article I kept being surprised by the idea that long is good. When writing, I tend to prefer being brief.
That only really shows that the humans they're training on are terrible at grading essays.
GIGO is our God now.
Yeah, it's cool, but what about your savings account?
Narrowly for grammar however - is even that a good thing? It probably helps scale grammar help to more students, but if those tools became ubiquitous in grading and editing then unique voices would just disappear and a lot of potentially “great writers” might choose different careers because the machines don’t like them
The fact that it's being implemented in society is insane because anyone who is paying attention to the state of AI today already knows how it will go wrong: without reading the article I already guessed that it systematically discriminated against certain demographics. Which was in fact what the article claimed.
It's interesting that it's possible to predict what the scorer would decide, but the moment you actually implement it is when all of the known problems become relevant, and the intellectual wonder must take a backseat to the human problems.
To emphasize, this was 16 years ago.
Maybe nobody really makes a big deal about it because it is pretty much irrelevant anywah. Applicants provide a letter of intent that the grad dept people can, y'know, actually read for themselves, so I think unless you totally bombed the writing section nobody cared.
One major problem with algorithmic approaches, whether automated or not, is that they become the definition of good in the context and therefore become something that cannot be argued against. And of course it makes 'teaching to the test' an even more likely outcome.
If I were a conspiracy theorist I'd attribute this to wanting a dumbed down population. Unfortunately I think it is probably the other way round, the population is already dumbed down and a belief in AI unicorns is the result.
As Aristotle said to Alexander: 'There is no royal road to geometry', and so it is with education; it's hard work for both the student and the educator and no amount of AI/ML/algorithmic snake oil will change that without also changing the meaning of the word education.
So the reason this isn't the case, is because there are very simple metrics that tend to highly correlate with essay quality. It doesn't mean the grading-bot is actually evaluating essay quality. It's just looking for properties that are statistically associated with good essays. Remember, at the end of the day as long as the bot's ranking is close enough to the human grader's ranking, nobody really cares about the internal logic.
A very straightforward example is spelling mistakes. People who make spelling mistakes aren't necessarily bad writers. And vice versa, there may be great speller who can't write for shit. But by and large the people who spell poorly also tend to write poorly. Easily detectable grammatical issues, like misplaced modifiers, subject verb disagreement, or inconsistent tense, are also correlated indicators.
A very simple metric is essay length. Especially if its a timed exam. Good writers tend to have verbal fluidity, with words easily flowing to paper. They don't struggle converting thoughts too sentences. So they tend to end up with the most words written down within a fixed time period. By and large the longer a timed essay is, the more likely that its actual quality is high.
Grading bots basically rely on these statistical relationships. They're not measuring anything intrinsic to good writing. But at the end of the day, their student rankings are usually pretty close to that of a typical human grader. In some cases the bot will have a closer ranking to a random human grader, than two random human graders will have to each other.
The biggest flaw here is Goodwin's law. When the test takers become aware of the kludges that the bots use, they can exploit it. For example just dump a bunch of verbal diarrhea with as many correctly spelled words as possible. But even then it doesn't really hurt the bot's ranking accuracy too much. Because the kids who do the most test-prep and learn all the tips and tricks, are usually high-achievers who do well on essays anyway.
This is related to current fairness-in-AI discussions. In many cases the basic problem is ML systems leverage correlations for making causal decisions. Here, there is a huge ethical difference between scoring a person based on "is this a good essay" and "do the features of this essay correlate with features of good essays". Just like there is a huge fairness and discrimination difference between "is this person qualified for a loan" and "do the features of this person correlate with features of people who qualify for loans" (algorithmic redlining). Your last sentence has a big discrimination/fairness issue also, since you are testing even more for parental income and parental free time.
Thanks for catching the mistake.
Anyway, yeah, not really a correlation.
>Remember, at the end of the day as long as the bot's ranking is close enough to the human grader's ranking, nobody really cares about the internal logic.
This isn't true at all. Imagine you got a B or C on an essay that a human would have given an A to because you wrote it concisely and in plain language, or because you used language that's statistically correlated with being black. Does the fact that this is rare console you? "Sorry, but it's usually very close to the human grader's ranking." Close enough isn't good enough when you get the short end of the stick. "Sorry, you aren't going to get to go to the college you wanted because you use language statistically correlated with poor writing." Or just because you're different, so the statistical correlation doesn't apply to you, you filthy outlier. Just because it's a rare event doesn't make it okay.
In adulthood, this is like hiring or firing for work statistically correlated with good work. Remember when amazon rolled out the resume scorer? [0] Sure it was biased towards women, but it was close enough to human scores, so who cares about the internal logic?
>Grading bots basically rely on these statistical relationships. They're not measuring anything intrinsic to good writing.
At the end of the day, our goal here is to measure good writing. If the bots aren't measuring anything intrinsic to good writing, we shouldn't use them.
https://www.reuters.com/article/us-amazon-com-jobs-automatio...
You could ask a student to write an essay taking a firm opinion on some subject, and they could change standpoint every paragraph and there's no way these systems would know.
If I was a student I would be extremely offended at people wasting my time like this.
If it is cost-prohibitive for every essay to be graded by humans, then they should be dropped from the tests. Otherwise, we are missing the whole point of essays which is to communicate effectively with another human, not just match certain text patterns.
Apparently it is. But everyone still wants writing to be assessed…
If you want to grade on form to test the ability to write correct rather than coherent sentences, make those separate questions, and mark them so.
I agree, this is traditionally the purpose of an essay. But to play devil's advocate, consider the rising number of people who are writing SEO or ASO content which is actually targeted at machines.
And “between 5 to 20 percent” of essays are randomly selected for human review.
So the takeaway is that if you’re one of the 80-95% of (typically black or female) people who the machine scored dramatically lower, but are not selected for human review, your education future is systematically fucked and you have no knowledge of why or how to change it.
Absolutely reprehensible. Anyone involved in the creation or adoption of these systems should be ashamed.
At least the machines offer the following hope: even if unbiased humans are rare among paper-grading teachers, those humans can be used to train the machines, so then bias-free or lower-bias grading becomes more ubiquitous.
Basically, the system has the potential for systematically identifying and reducing systematic bias. A computer program can be retrained much more readily than nation-wide army of humans. Humans can be given a lecture on bias, and then they will just return to their ways.
AI certainly has a lot of potential for bias, but claims that AI bias is somehow worse than good old human bias always seem shoddily supported (Note I'm not claiming it's untrue. Just that it's never been shown to my satisfaction, which is not surprising given how quickly AI is changing. It may well be true.)
Well, AI bias can combine sampling bias with human bias. Like say we train the AI with the output of only 10 human paper graders, all chosen from the same school district.
Due to the sampling bias, that data could create markedly more (or less!) bias than the entire population of human paper-graders.
The resulting AI will ideally mimic those 10 humans, though; it shouldn't show more bias than that group. If those 10 are flaring racists, and grade accordingly, the AI will be the same. (In fact, we hope that it will be the same, if the algorithm actually works in mimicking human grading.)
The bias comes from the human-generated training data in the first place; the machine isn't introducing its own. For instance, the machine has no inherent concept of disparaging someone's language because it's from an identifiable inner city dialect. If it picks up that bias, at least it will apply it consistently. When we investigate the machine, the machine will not know that it's being investigated and will not try to conceal its bias from us.
On the other hand, eliminating bias from humans basically means this: producing a new litter of small humans and teaching them better than their predecessors.
That's the problem - there is seemingly no shame these days. People involved "saved time and money", got paid and that's it. "If I didn't do it someone else would" and all of that.
I remember taking a standardized test, can't remember if it was SAT or CSAT (Colorado pre-SAT test). This was at a time when I'm confident that humans were the graders.
I started with an intro that would be appropriate for a standard 5 paragraph essay; i.e. the thing you write when you don't know what you're talking about and you're just following a format.
In the third paragraph I took a leaf from family guy, and just interjected "WAFFLES, NICE CRISPY WAFFLES, WITH LOTS OF SYRUP." for the next page and a half, I berated the very foundation of the essay prompt, insulting it the way only an angst ridden early teen can.
... I got a 98% on the essay.
Fast forward several years. I write an essay for for an introductory college course final. My paper is returned to me with a coffee stain and a "94% - good work!" note scribbled on the top. That note was scribbled by a TA that would turn out to be my girlfriend for 2 years. One night in bed, she tilts her laptop to me, showing an article that I used as the central theme to the above essay; "can you believe this?"
"Are you joking? Of course I can believe this, it was the subject of the essay you gave me an A on 2 years ago"
She admits she didn't read past the first paragraph of anything she grades, and just bases grades on intuition based on how articulate the essays are at the outset.
...
The point I'm making:
Does AI suck at judging the amount of informative content in a student essay? YES
Do humans suck at judging the amount of informative content in a student essay? ALSO YES
"People worry that computers will get too smart and take over the world, but the real problem is that they're too stupid and they've already taken over the world."
Any essay writing test which could be adequately graded by a machine is not testing anything of value.
Edit: I’ll further add that as soon as people’s careers depend on a metric, the metric becomes useless as a metric, because it will be gamed and manipulated by everyone involved. Almost nobody involved is incentivized to accurately measure student’s writing ability.
A lot of what students write is actually garbage from that point of view. Even if they happen to have a good basic idea about what they want to say, the point of essay writing is to master the mechanics of expression so that you get the idea across effectively.
Whether the student has a brilliant idea isn't even so important, and it wouldn't even be fair; imagine if high school computer science expected students to turn in a best-selling app for a term project. Not everyone can come up with something brilliant to say; and even relatively mundane lines of reasoning can be given a good treatment in writing to develop the skill.
I remember when I had essays graded in school, a lot of the comments were low-grade fluff like "run on sentence", "wrong word", "faulty parallelism", "missing colon before 'for example'" and such points having nothing to do with the content being original, well-considered and well-argued. That sort of thing might as well be done by machine, at least as a preprocessing step to improve a student's rough draft.
It's the same reason you see keyword posters in math education. "Together" means "plus", that kind of thing. It's completely worthless, except for one-step problems, and even then it doesn't always work. What is happening is collusion between teachers and testmakers. You can't teach understanding, but you can teach test-passing techniques because the way the test is set permits this.
You see the same thing here, in English you can get away with not teaching quality writing if you teach techniques to score well.
Hint: together does not mean plus.
When your essays are graded they're marked down for mechanical and wording problems. There's really no point in trying or grade 'good ideas' on a subject piece you had maybe 10 minutes to skim.
That's a travesty, and you know it because when the kids are in college and they have as much time as they like to write their assignments they all use the wrong words and then misapply them.
However, it is rather important that students know their essays are not judged as essays, but only judged on the content. Otherwise you teach students that form trumps content in essays.
When judging an essay as an essay correct English barely matters. What matters is how convincing you are, and how interesting of a read the essay is. This is a great skill to have, and testing it also makes sense. Really though, we should separate these two forms of testing.
This always gets made into some kind of techluminati conspiracy for the machines to ingrain structural racism whereas it's pretty clear all the algorithms fail to do is improve an already bad situation stemming from a flawed premise.
It would make sense to use the AI as a first pass, and then not randomly grade the essays with a human, but specifically choose all the essays that are on the cusp of the pass fail line. Then use all those human generated scores to update the model, especially if someone moves from pass to fail or fail to pass. Then maybe throw in a few of the really high and really low outliers to make sure those are right, and throw away your entire model if the human scores are drastically different (and obviously don't tell the humans what the computer score was so they have no idea if they're reading a "cusp" essay or an outlier essay).
But putting the educational fate (and therefore future earnings) in the hands of an AI is unconscionable.
I think the right approach for machine learned systems is to automatically "whitelist" essays rather than "blacklisting" them. Students in the middle of the distribution of essays aren't really interesting, so whitelist them, give them a pass. Those at the extremes can be either exceptional or terrible, but usually terrible. The judgement of those at the extremes should be decided by a human, not a machine. You wouldn't want to blacklist the Einstein of essays because he did something genius that is indistinguishable from insanity.
However, I think there are some essays that can automatically be blacklisted. For example, those with:
1. Plagiarism (perhaps human moderated)
2. Extremely low word count
3. Extremely high count of fake words
And at the end of the day, these essay assignments aren't there to judge whether a student is the next writing sensation; they are given to judge whether the student can write legible sentences and words, to ensure they are prepared for the future. So perhaps it is at least possible to automatically blacklist on sentence structure and spelling (you should just lose points for invalid structure or invalid words, you shouldn't gain points for big words or complicated sentences). To make this fair, the student should be informed of this requirement. If they are informed and still fail, then they need to be remediated. If we discover that a disproportionate number of minorities are getting blacklisted, then we should investigate why the school is failing to teach them proper sentence structure and spelling, not pretend we can change the world to make AAVE an acceptable dialect of english in the workplace.
Grading students' code using a machine is not such a bad idea in contrast, because in that case there is [1] no exceptions possible in a programming language, [2] the machine (compiler) has to understand it anyway, and [3] it does save time verifying correctness. But communication in a human language really needs to be assessed by humans. Anyone who thinks "AI" can accurately assess human language is either severely delusional, or trying to make $$$ from it.
The problem that I encountered was actually how to apply it in a useful way, since the problems mentioned in the article are quite obvious when you design the model.
Options that I saw:
1. Use it as autonomous grading with optional review by the teacher, see the linked article for the problems with this.
2. Use it as a sanity check on the teachers manual scoring, but it would not reduce the work load and probably just undermine the teacher.
Do you have any suggestions for how such a model could be applied in a practical and ethical way?
Had some thoughts on how to measure actual knowledge about a subject, but that would require a massive knowledge graph which would introduce a huge amount of complexity just to see if it would be a feasible approach.
One trick might be to write an independent AI to summarize the essay back and see how closely it matches the essay title. This might weed out gibberish essays with sound English sentences.
First of all, complaining about minorities getting lower grades because their English is not as sophisticated as that of others is the inversion of the idea of teaching. That feedback is actually great. We have machines that can give that feedback (e.g., grammarly)? Then use it to make everyone's writing better. Grades are just a measure of the success of learning, after all. I never got why one would not allow a student to repeat a particular test as often as they like, tbh.
Second, grading essays this way is a clear violation of the idea of teaching. What do you want the students to learn? Structure? Knowledge transfer? Grammar? Writing an essay is such a complex task it is a really too broad goal. And then naturally grading becomes quite difficult.
It was a few weeks ago that someone shared “The Dark Age of AI” on HN [1]. I think we are promising way over what Drew McDermott thought we would not going to promise. This is to the extend that we are applying AI on assessing Art, Creativity and even quality and novelty of Science, something that in a way we don’t even understand (or trying to understand) ourselves at the time that we are publishing it.
Evaluation of how well someone follows arbitrary language conventions is worse than useless.
I only got to university English 101 outside of some technical writing in the engineering department, but I have to say none of my education in writing was worth anything past elementary school. It is perhaps one of the most difficult things to teach and evaluate, to be fair, but I feel like I am missing a huge chunk of my education and general ability because of it. I can't write or form an argument particularly well, rambling on HN and the like is the closest thing to education I have had.
Prescriptive language rules are not entirely useless. That is the best you can say about them.
I would support any student who refuses to consent to their work being used in this fashion.
If you don't talk about the actual problem, how can you possibly expect to solve it?
The point then isn't that the algorithm hates black people or the programmers are racist (even if they were they would likely find it hard to train it to specifically exclude accurately without major side effects), but that their training set or analysis is flawed - potentially even if it is representative. Even if the results are problematic trying to address hate which isn't there won't be helpful. Call for better vetting procedures instead or that it clearly isn't ready for the proposed application.
Say it gets the highest accuracy on a set of photos of non-people images and a set of people which represents the exact ethnic make up all tagged just as "has people" or "doesn't have people". Going with something biased towards the most populous groups would get the best results quickest.
Talking about say image tagging to ensure automated performance checking across various characteristics could be productive on the other hand.
> Of those 21 states, three said every essay is also graded by a human. But in the remaining 18 states, only a small percentage of students’ essays—it varies between 5 to 20 percent—will be randomly selected for a human grader to double check the machine’s work.
So that applies only in a minority of cases.
This reminds me of a wonderful essay/speech by Stephen Fry on the harm done by pedantry. I also feel that schools focus so much on a single structure of essay writing and similarly take the joy out of language.
There's no necessary connection there, especially if one of the reasons that teachers are being overwhelmed is that the teacher/student ratio is increasing.
> You don't need a 5 page essay to determine whether a kid has read a book.
No, you need it to determine whether a student has (1) read and understood a book well enough to apply structured thought to the contents and (2) has developed the writing skills to write a 5-page essay.
Determining whether a student read a book is rarely, on its own, of significant interest in school.
Don't think anyone passed that class.
I am absolutely apalled.
Not even at the idea of grading by algorithm, but by the fact that many, many people had to cooperate to make this happen.
Wouldn't it be better and less biased if we each wrote our own AI systems and had them discuss with each other instead?
(And we should publish our training data as well, of course)
This seems like a terrible idea.
It's not a stretch to imagine the opportunity for nefarious behavior this allows - think of the recent college admission scandals, and how happy they'd be to have a guise of algorithmic indifference'.
If used long-term, it could offer a big advantage to the wealthy in other avenues. Another hypothetical, probably not far from reality: the algorithm becomes solved (almost or completely) by some premier 'tutoring' company. Said company can charge a pretty penny given its stellar track record, offering yet another hidden advantage to the wealthy/elite.