U-M finds students with alphabetically lower-ranked names receive lower grades
record.umich.edu
record.umich.edu
However, I also grade weekly exercise sheets during the semester, and these are committed into a repository, where each student has a folder that... begins with the first letter of their first name. Everyone I have ever worked with acknowledges that you have to shuffle the order in which you grade these submissions each week, for fairness. Several effects come into play: (1) your are usually less tired at the beginning, (2) your mood gets better during the last 2 sheets because you know you are done soon, (3, and crucially) at the beginning, you have not yet seen all the common errors / developed a "feeling" for them, and you might thus miss them in early submissions, but spot them immediately in later submissions.
Another alphabetic effect: In elementary school, my name was on top of the list of students in my class. I remember that I often had to do some special job simply because I was the first name on this list (for example, carry a group ticket when we visited some museum, keep track of something, be the first at something where nobody wanted to be the first, with everyone watching, be the first to be graded in PE, again with everyone watching, etc.). As a fairly shy kid, this already annoyed me in first grade.
Crucially, this is not quite what the poster said. It’s not about stack ranking students against each other.
Say every paper makes the same subtle mistake, and you only notice it halfway through the pile. Unless you go back through them all, you’ll unfairly grade the later entries more harshly.
It's not, but it sort of has that effect, albeit indirectly, and definitely unfairly.
If you weigh the severity of students mistakes (or successes for that matter) in relation to each other rather than to an objective rubric, you’re effectively stack ranking them whether you mean to or not.
This ensures that everyone who made the same mistake(s) gets the same grade. It also tends to shuffle the order of the exams after every problem.
Obviously you don’t need this strategy for simple multiple choice questions, and it’s probably also not a great fit for long-form essays. But it worked great for technical short answer problems in CS and security.
Also: if you're sorting into "mistakes piles" for single exercises, how can you parallelise marking of separate and independent questions?
[Edit: not disagreeing with your point.]
No grading is perfect, but there’s also some undercurrent of an attitude that students have paid to be there and are entitled to a certain grade.
Given that students have taken on hundreds of thousands of dollars in debt that they'll have to repay no matter what and on top of that a lot of jobs being completely out of reach these days without an academic degree (that for fucks sake isn't remotely required by virtually all jobs requiring it!), that's completely understandable.
Want to fix higher education? Bring the hammer down on companies abusing it as a proxy for legally discriminating against classes of society that are closely correlated with poor academic outcomes. Academic education should be reserved for the best of the best of our youth, and it should be fully paid for by the government, not simply another hurdle to pass to get a job that pays barely more than flipping burgers.
I also think that the vast majority of poorly paid, non-tenured professors and other teaching staff don't love being the targets of this harassment, since it's not their fault and largely out of their control, and it's not like they're getting the bulk of the tuition money. (That mostly goes to administrative expenses and sports programs.)
Heck, most adjunct faculty are often paid below minimum wage and qualify for food stamps.
Would that my students were this engaged before the exam. Guess which students show up the most often for office hours? ... yeah, the ones that are getting the best grades.
If my students spent half as much time learning the subject as arguing with me about grades, they would be getting a higher grade than the one they are arguing for.
Often, you still get big problems, but the set of solutions is small. It's always three options plus a fourth option (none / all). If you make a mistake you score negative points. It's not perfect, sometimes wording is ambiguous and it's unclear whether you need to tick the fourth catch-all option, but I found it better than the alternatives as it removes most arbitrariness from the process, but obviously has other issues.
Regular exams often had wildly different grading standards for the same course depending on the class, and thus on the professor who was correcting exams. This was really annoying.
Then, each problem was assigned to a TA. Either there's a predefined rubric, or you create it as you go (-1 point for mistake X, half credit for mistake Y, etc.). There's a pretty slick interface where you just read the answer, and use keyboard shortcuts to apply the relevant deductions.
It still has the issue that every time you change the rubric, you'd need to go back and re-do previously-graded instances of that problem. But it was way faster and (equally important) less tiring.
(disclaimer: I briefly worked on the software for my bachelor’s thesis)
When I was grading homework, it took about 5 hours a week per class per run through. They didn't pay me enough to make sense for it to be 10 hours.
Graders know that wrong homework takes much longer than correct homework to grade. It's correct? Full marks, move on. Is it wrong? Well, how wrong is it? Did they make a bad assumption, but followed it through to its conclusion? Did they forget a minus sign? Or is it complete hogwash?
So it might not be 10 hours, but still would be around 8 hours. And that's still too much.
Get all the profs and TA's together, break in to groups taking one problem or set of problems. Then you random sample (each group takes a stack) to get a feel for the 'typical' errors, once that's done - you are a machine going through the stacks.
Every once in a while (not that often) you run into a novel error or approach, and the group discusses.
If you use Canvas or Gradescope with the default settings, it’s almost impossible to avoid this sort of bias.
Worse yet, in Gradescooe you’re strongly steered towards grading with a fixed “rubric” with specific points off for each of N pre-defined errors, allowing grading to be done by TAs with little more knowledge than the students themselves, resulting in scores which have little relationship to the quality of the student answer.
You could probably cover most of this with an LLM, and access to a large body of graded material for a given course, provided said material was graded fairly. Generating that data would be time consuming, as, any given assignment would need to be graded by as many people as possible in order to find a fair average.
From there, it's simple comparison between your sample work and the presented work. We're probably a decade from this really being viable en masse, but, it no doubt will happen, and for better or worse we'll likely end up with EDUAAS (education as a service).
And not everything has an objective solution. Even those that do often have a process associated with them and factoring in that work/process is an important part of grading. Reducing that subjective grading process to only objective solutions being right is grossly reductive and disproportionately punishes students who have the process right and understand the material but make small errors. That's exactly what you don't want to do.
---
Instead the solution is to make sure each assignment gets multiple eyes on it and in a random order. Then to document biases and trends in biases so that the TAs and professors can be aware of them and mitigate them.
It's a process problem that can only be solved by a process solution. Replacing the graders with technology or reducing problems to a binary right/wrong will never ever solve this and in many cases will end up being more harmful than the biases they claim to solve would be.
Have a group of N graders. And a parity of k. Let's say N is 6 and k is 2. Randomly shuffle the assignments and partition the assignments into N groups.
Each grader gets assigned k of the N groups such that they share at most 1 overlap with any other grader and each group is assigned to k people. The assignment orders are shuffled for each grader. They mark up and then grade the assignments.
Then for each of the N groups, randomly shuffle the group and equally distribute the assignments to the N-k graders.
Now each grader reviews the assignment grades/markups (in random order) and assigns a grade based on the k grades/markups from the previous rounds along with a rationale for the grade assigned.
From there the student receives the final assigned grade, the rationale for the grade, and the k markups. If they have a complaint they can go to the professor (who then can also see the k initial grades along with everything else) to dispute the grade for the assignment.
---
This way each TA only has to mark up (class size * k / N) assignments, and review (class size / N) assignments to assign a final grade (which should take far less time to do than the initial markups). On top of that every assignment has a guaranteed (k + 1) separate eyes on it. And then the professors can serve as an unbiased arbiter while retaining all the context from the process.
To take it an additional step further, the professors could sample a random subset of the assignments to verify the markup and grading is going properly.
And those reviews/grade adjustments can then be recorded (along with the final grade/rationales) to document how a given TA's grading deviates from the final reviewed grade or the grade the professor assigns. Likewise for a TA's final assigned grade deviating from the professor's. This would allow deviations to be mitigated over time and major deviations to be identified.
I received an 80% with no notes or markup.
I have been left wondering for the last 25 years how much student work is actually even reviewed.
I work in EdTech and every time we add a feature that requires manual teacher review of student work you will see that some teachers are VERY diligent while others never touch it.
I don't think this is fair. It's just a more randomly distributed unfairness, rather than by a deterministic factor (like the student's name)
'Fair' would be each student is assessed independently for the work they did, rather than their mark being impacted by how early or late they were marked.
In this example, I think it's kind of fair to give everyone an equal chance of being advantaged. You're not hurting anyone specifically.
Maybe this is a case that AI could actually do quite well.
Manually grade the answers and identify the classes of mistakes. Then hand the classes of mistakes to the AI and ask for it to determine which answers have which types of mistakes.
Once you've done that, you just need to associate a deduction for each type of mistake and do some simple math.
Some students might mix up the algorithms, some might give the incorrect computation complexity, some might describe them incorrectly in some way.
Manually grade some (or all of) the answers by noting the kinds of things students got wrong (e.g. the above criteria). Then, feed in to ChatGPT (or your favorite alternative) the answer + the categories of mistakes to expect.
Here's a simplified example: https://chat.openai.com/share/bf801e12-51d5-4255-9968-bbf91b...
Or just having a Kafkaesque pass fail grade with no feedback for each student relative to their own performance over time with an expected growth rate applied?
Obvs this takes an administration that is OK with that, which most aren't.
(I suppose the cons outweighted the cons.)
Did you perceive any pros?
I suppose one way to do grades is first read through all papers to get an idea of the levels of the students. Though you still have bias/nepotism and such then. Perhaps a teamwork or commitee would work, or teachers swapping classes/schools?
I had a French teacher on high school who dropped a pen on list of students and then where it landed that person would get rehearsal. People in mid (waves) were fried.
Plus, there is also the issue of certain last names being common in certain cultures, leading to skewed statistics.
Finally, I guess I’ll admit that I’m probably very biased because my initials as A.B. and I’ve always gotten excellent grades, so… maybe maybe maybe
It seems like you're agreeing with me, but jumping to their defense with "people are fallible." People are fallible, that's why we build systems to take human elements out of it. Recognizing where humanity has soured something is key to that.
It's also worth noting that randomization in a context like this is inherently an imperfect solution to a problem that generally can't be solved perfectly. If we find out that weird ordering biases exist, I think randomization is done on the assumption that many we don't know about could also exist, that there's no clear way to mitigate them completely, and then randomizing the order per-instance is just the best we can do to ensure it's fair (Which, again, won't be perfect. Perfect isn't available)
They are claiming that randomness is not sufficient for fairness.
And I agree. Adding some determinacy to the shuffling design would reduce the variance of unfair scores that students receive. Ie, it could reduce the likelihood that a student's scores will be biased by a run of advantaged or disadvantaged scoring orders.
It wouldn't be very hard to design such a scheme. It would produce fairer results in the short run. And in the long run, it would converge to the same fair averages as the fully random process.
In fact it would converge faster. This short run statistical efficiency matters. Students don't have access to the long run results - since they're evaluated in the short run.
It is important to get a feel for the collective level of writing before grading essays individually, and it is important to avoid over-grading or under-grading essays at the beginning or end of a stream of papers. Therefore I did a three-stage grading process, with three colors of pens:
The first pass, with a red pen, is marking up single-point problems like misspellings and glaring usage errors. Also of course the general level of writing begins to seep in. This pass of course includes all papers, and it can pass fairly quickly.
The second pass, with a green pen, mostly just marks in the margin where a (good) point is made or a conclusion is reached. This is to prepare for the next pass. Again, all papers are done in this pass.
The third pass, in blue pen, is where the quality of the writing is assessed and critiqued. Maybe some short notes in the margins, maybe just comments at the end of the essay.
When students get their papers back, there are some chuckles (or whatevs) when students see all the pretty colors. But after I explain the method and its rationale, the method is clear and understood (and also appreciated?).
My parents couldn't figure out for the life of them why I was suddenly struggling and thought I was having adjustment issues. I had taught myself to read when I was 3; how could I suddenly be having trouble keeping up?
It took longer to figure out because I was only nearsighted in one eye. I was tall for my grade, so as long as the person in front of me to the left was shorter than me or the teacher was writing high enough on the board, I was fine, because my left eye was fine. But when everything aligned just wrong, I was suddenly helpless, because my right eye could barely see clearly an arm's length from my face! It's a hard thing to notice when only one of your eyes isn't working very well, especially when you're 9.
I haven't noticed a grading/ranking difference, but far more frequently I'll hear that "oh, we ran out of item/time/etc. before we got to you", which has made me much more sensitive to issues of planning/organization.
It was 40ish years ago and I still don't think she's forgiven my dad.
It's a remote job so, being frequently visible in that list can be an advantage.
We should ask people with later letters if they remember this more.
I can just say I do remember being last in a lot of arbitrary and official things and seeing other friends just get done with it faster and have to waste less time sitting and waiting.
I remember when working on joint tasks, by the time it got to me, most of the people that worked with me had already given their updates and details. So when it was my turn, I'd say "same as A, B, C", cause they'd given all the juicy details.
Other than that, it's pretty straightforward and boring. The world doesn't magically function differently for us.
I wonder if this was the kiln of my patience and acceptance, or if people who road rage and get frustrated with waits are more likely to have earlier lettered names?
I was last in the alphabet; this was already an issue with books we had to read; you could choose which book to read, but it was always in alphabetical order. When it was my turn there were just a few left, and certainly all the popular high-demand ones were gone.
Anyway, when it finally was my turn to get my marbles he was all out. When I asked "where's my marbles?" he just shrugged and said "all out". I must've been about 7. Lots of crying ensued and I think I got some marbles from other kids, but it wasn't about the marbles – not really.
I still don't understand how anyone can expect any different result...
Kinda surprised, my last name starts with C and I was hyper-aware of this and how random it was probably all the way from kindergarten. Being a child and therefore an asshole, I was grateful for my advantage rather than thinking the system was unjust.
We didn’t put our names on any of our work other than our dissertation (and a few trivial assignments that didn’t impact overall marks). It wasn’t that hard to de-anonymise, but it meant that the system had a bit more integrity.
It’s a really straightforward system to implement and I don’t know why it isn’t done more frequently.
I also think that our VLE sorted assignments by time of submission rather than any identifier.
Maybe you could make it even more impartial by allowing conditional scores in the first two rounds. Like "Jim is a 6, but a 8 if his talk is about molecular biology" or "this Lessons Learnt talk is a 5, but if it's by X, Y or Z it's a 9"
Certainly a talk by X that's totally unconnected from anything they're directly involved with has less value.
But computer science courses tend to have very objective rubrics for grading, so I am not sure the anonymity mattered much.
A random identifier coupled with a random sort order seem like the way to go here.
Fixed in the sense that the bias will be random. Presumably students graded last will still receive lower grades.
Of course it’s better than a fixed order, and if it’s easy to switch then might as well. But we should keep thinking about how we can make it even better.
Still, yes, you could flip the order from midterm to final instead of randomizing both and the effect goes to more like +/- 0.1 out of 100 for the luckiest and unluckiest.
(For example, if you have two, you can simply swap: but what about other biases? like if it's broken in half to assign to 2 grades. Or what about if there are three exams? And what about balance across other courses? if you want to do variance-reduction and tricks like antithetic sampling, you need to know all this in order to structure it properly - get it wrong, and you may make things worse.)
So that's why simple random shuffling would be preferred. It allows total ignorance of all other uses (past present and future), handles all ordering biases, and can be done independent in parallel across arbitrary sets of courses/exams/grades/students.
From my experience as a tutor, yes, this bias exists. But it won't turn a horribly wrong or an excellently correct solution into anything else.
I eventually knew my strugglers and my excellers. I'd skim the excellers first, because if they messed up, something bad was going on. Then I'd go through the strugglers to see problems. And then I'd grade the rest first in whatever order I got the sheets, then the strugglers and then the excellers. I needed the baseline to see how bad the worst ones actually do. Some exercise sheets were an accidental adventure, I can tell you.
And writing it like that, it sounds totally callous and cold. But focusing on the lower third in the exercises and communicating their struggles to the TA and prof was very appreciated by everyone, especially those students. It makes sure to get the important fundamentals right.
There was also the factor that the ones I graded initially did not make certain mistakes or answered in expected ways, such that when I did encounter unexpected answers/mistakes, I had to go back and rethink the grading on the papers I had graded previously. Eg if someone answered in a way that made me think an answer I considered incorrect was actually less wrong.
I only had to deal with a small class, so backtracking was doable and I graded the papers in whatever shuffled up order they were turned in, otherwise there would have definitely been a bias.
I’d either find that:
A bug was really common, got to re-evaluate after the first couple times I see it, apparently it is an easy mistake to make.
Or, I’d find a new bug that was pretty common, but which I didn’t know about at first. Got to update my tests and re-run everybody.
I tended to be really thorough and re-do the whole stack eventually, but it was a real pain. Could have half-assed it of course, but they spend weeks on these things, feel like I owe them honest feedback.
It would tend to lead me to “softer” grading as well, if you are lazy and only check for a couple bugs, you might take a large number of points off for each problem. Finding some problems and punishing them harshly is not very fair for those students that randomly hit the bugs you expect. If you find every bug, you can only take a couple points off per bug without tanking everybody’s score.
Grading papers in submission order just introduces a different bias, though.
(For what it's worth, I'm in the same boat and I do the same, because I don't trust my ability to give the papers any true random sorting by hand, so I take the very weak randomization that the submission order gives me.)
I was usually (almost always) the last person to turn in assignments, I like to be one of the last people out of a door or the last person in a line (I don't like crowds).
Grading by order-turned-in would almost always mean my assignment would be one of the first or last one's graded.
If I were to guess that if you did a frequency analysis of people to order, you'd find there were always a certain group who turned it in first, and another group that turned it in last.
That's what I'm saying—it's reasonable to believe that the submission time is correlated with other factors, such as ability or confidence (though the effect can cut both ways, with extremely able students submitting early because they finish early or late because they are extra careful, and similarly for other factors). Thus, this isn't really randomization, just correlation with another factor than the name.
My last name starts with E and my wife's with Y. Bucking tradition, she didn't change her name when we got married, so when we had kids we had to decide what name to give them. We opted to hyphenate.
Historically, hyphenated last names were [Woman's last name]-[Man's last name]. However, my wife hated that her last name was near the end of the alphabet growing up.
We bucked tradition once again and put my name first, so that when sorted alphabetically they would be at the front of the list. Incidentally their first names start with A and B so that they show up at the front when sorted by first name too.
Unless you were married earlier than the 90s, I wouldn't really call that "bucking tradition" any time from, say, the mid-90s onwards.
If you really want to buck tradition, then don't get married - just live together, and have kids :-)
(After all, there's nothing more traditional than marriage, is there?)
But you hit on an important point — a lot of couples are just skipping marriage now.
We went halfway there — we bought the house together years before we got married.
> we should dismiss this finding, simply because it is impossible. When we interpret how impossibly large the effect size is, anyone with even a modest understanding of psychology should be able to conclude that it is impossible that this data pattern is caused by a psychological mechanism. As psychologists, we shouldn’t teach or cite this finding, nor use it in policy decisions as an example of psychological bias in decision making.
But the article is written specifically to make the point that it should be enough to observe that it isn't possible for the effect to be real. You aren't making a good point when you cite an effect that is obviously nonsense.
More worrying is when e.g. job candidates are discussed (often in alphabetical order) and people simply tire out near the end of the meeting. When this happens, be sure to suggest taking a break!
[1] https://electionlab.mit.edu/research/ballot-order-effects
[2] https://www.cbc.ca/news/canada/british-columbia/vancouver-do...
1) Rubrics are often defined, but the application of the rubric is by a human. Application will shift as the grader gets a sense of the classes understanding.
2) As you get fatigued while grading, you'll make mistakes, and be less tolerant of others. Especially if you're an overworked adjunct or graduate student.
3) There are probably a lot more last names early in the alphabet so weighting is important.
My policy on this when I was a grad student was to publish the rubric, and ask all students to check their grades too.
there is a similar effect found here https://en.wikipedia.org/wiki/Hungry_judge_effect
Also, it is never your professor who grades you - the answer sheets are collected and lecturers/professors will correct them at the state level across all the engineering colleges in my state.
I do not know how it is now as there has been an explosion of colleges in the state. But expect the standardized tests are similarly conducted.
It's quite different from the way it was when I studied physics in the 1970s when only the final counted. Annual exams only determined whether one was allowed to continue but had no effect on the class of degree that was awarded.
The database could tag every assignment with a UUID4, and present them for grading top-to-bottom in UUID lexical order, without exposing who is being graded in any way.
You can't fix fatigue bias, but this would distribute it randomly. It also removes the opportunity for favoritism and hostility, subconscious or otherwise, which is probably more important.
Once grading is completed, the assignments are reconnected with students. Give the profs a way to mark assignments with metadata, sometimes they need to talk to a student personally about something, this should be made easy.
Grades can't be immutable, professors need discretion in that, but it would leave an audit trail if professors maliciously modified grades (or the opposite). That should be uncommon to begin with, but both professors and students benefit from an audit trail here.
A system like this should be used whenever it's practical, and always for high-stakes tests like midterms and finals. Not making a case against oral exams here, just that when it's possible to blind the grading process, it should be.
As in, order matters
So I guess "alphabetically lower ranked" means the last letters of the alphabet, not first? Confused.
The programmer's perspective and the user's perspective aren't always the same, and both need consideration. A user is going to see a list: it starts at the top, and it ends at the bottom. The first fields are higher, the later fields are lower.
Of course, if this is a sorted list, the first field will be the "lowest" value, for whatever comparison is used to sort it.
I can actually believe the effect going in either direction and it's small.
I searched the authors in google scholar but I couldn't find it.
Not on sci-hub, but downloadable in pdf for me without any issues
As someone whose first and last names are both very early in the alphabet, I was always called on first or second when I was in elementary school and middle school. I always had to be there early.
My friend whose name was very late in the alphabet learned he did not have to be ready for the first minute or two of class.
He would be standing near the door talking as I was quickly pulling out last night's homework, and I would be marked down for not being ready while he would later be commended for being ready when the teacher called his name.
As a teacher, I see that the kids who stand outside the door talking do not do as well as the kids who are there early.
That would be my best guess for a rationale behind that result.
Even now, for example, if you go to Play Store and want to know the apps that you had but are not installed, the default sorting is by name.
I've found that gradescope is helpful in this regard, because it at least forces every point assignment to be matched to a rubric item. I don't have data, but I believe it makes our grading a lot more uniform compared to the pre-gradescope days. (This might be easier in grading computer science exams than in more subjective areas, though.)
> The researchers collected available historical data of all programs, students and assignments on Canvas from the fall 2014 semester to the summer 2022 semester.
Thousands of students X 8 years X lots of assignments per year and you get a sample size so big that it would be hard not to find statistically significant effects.
1. Keep the present system of grading by alphabetical order
2. Record the order in which the papers are actually graded
When the grading is done, the teacher assigns a point scale (A = 90, B = 80 or whatever) but the computer does a regression fit and removes the bias.
The evidence is a strong negative correlation between bias and achievement: Extraordinary achievements so disproportionately achieved by people in groups that are not the target of bias. Look at top government officials, SV leaders, Nobel Prize winners, etc etc - mostly white males.
The biggest targets of bias in the US, for example - probably women and black people - genrerally get the worst results (in areas where there is discrimination). By contrast, as an example wherever black people aren't subject to bias, such as certain forms of music and certain sports, achievement is extraordinary. Imagine all that talent and drive in other fields.
Anyone know whether this is happening, and if not why not?
Interesting category of problems...
The hand-wringing over such a small effect size seems unwarranted. I suspect you would find similar effect sizes for other small interventions, like whether the grading took place during the week or the weekend, or in the morning vs. the evening.
Almost 20 years ago I worked for a standardized test essay grading service. We graded against all sorts of secondary-level rubrics (not AP, who do their own). These would usually be from 9 - 12 grade, from every US state, and evaluating everything from reading comprehension to subject matter-specific assessment. We'd do weeks long jobs of a single test (e.g. Alabama 9th grade reading proficiency). These usually had at least 3 dimensions, and at least 4 points per dimension. We would go through a week or more of training on a rubric, then another week of 'leveling', where a manager would occasionally bring you aside and talk through why that '3' you gave on a dimension should have been a '2'.
By the end of the training, we usually had had enough discussions and encountered enough edge cases to understand the weaknesses/inconsistencies in the rubric (which we had to abide by anyway). Once we were running at full-speed, everything was still double-graded and inconsistent scores were reviewed. Sometimes graders were pulled if they still didn't get the rubric.
It was a simultaneously stimulating and very boring job, and most readers were educators themselves. I wonder how long before it disappears completely.
We were always woken up by my daughter screaming as here friends called her. No such luck for the post-pandemic kids.
This is critical. Otherwise we could not discount some group (e.g. some ethnicity) disproportionately occupying one end of the alphabet or another.
Super interesting and important finding. I hope this gets wide visibility and universities take a break from politicking to fix the problem - presumably through enforced randomizing.
Based on these results, it would mean that the graders are just getting tired/lazy/inattentive the further they get in their stack of papers to grade. That's the problem the needs to be fixed, not the order they get graded in. Enforced randomization is simply a short term alleviation so no student(s) get disproportionately affected by this phenomenon.
Or maybe they are getting better / more picky.
I know in code reviews I often pass a few and then notice something that I realize was also wrong in previous reviews I allowed, but later reviews that day (week?) will not allow that.
It'd seemingly be more work but would result in averages that are more reasonable to the changes in stress.
Maybe decision fatigue is supposed to bias humans toward the optimal solution for the fiancee problem [1].
Additionally, universities (and, by extension departments) want grades to approximately follow a normal distribution (and yes, you in the back, their actions show they do actually want that, even if they say otherwise).
When you start grading a problem you have some idea what a "good" solution looks like, what an "ok" solution looks like, and same for "bad" solutions... If you award points based on that, the result will be a normal-ish distribution. But your idea of a good/ok/bad solution evolves as you see more papers.
There's two reasons for that:
First, you can't (ahead of time) imagine all the ways that students will invent to fuck up a problem set, and find edge cases in your grading rubric that result in unfairly-high or -low scores. As you gain experience teaching, you will anticipate more of the ways, but you will never anticipate every way.
Second, the TA/grader wants to be able to stack-rank the papers and have the scores be monotonic. The grader wants this because non-monotonic scoring triggers far more complaining than harsh scoring or picky scoring. When you come across papers that are worse than ones you've already recently graded, you assign even lower scores.
This results in a ratcheting effect with more extreme scores as you get closer to the bottom of the pile. But, since the mean score is usually a B/B-/C+ (~75-85), and since scores are usually limited to the range 0-100, this means that papers closer to the bottom will receive statistically lower scores.
Now, you could go back a re-grade ones you've already done, but:
1. The university is officially only paying you for 20hrs/week (and requires a signed end-of-semester statement attesting to the same).
2. The assigned workload of teaching and grading doesn't permit a two-pass grading scheme while keeping within 20 hours.
3. If you complain to the graduate ombudsman about the workload needing more than 20 hours, you won't have funding next semester (so you have a prisoner's dilemma among TAs who might want to grade more fairly).
4. If you're grading (say) a final exam for a frosh/soph class, you're probably in a room with 4-8 other graders late into the night. One effective way to make your coworkers hate you is to be that guy who always finishes grading his stack last, when everyone is worried about catching the last train/bus.
Basically, all the incentives are aligned to make this happen.
Essentially, unless it's an old exam where the universe of bad answers is already known, you need two passes - a discovery pass followed by the grading pass.
I will admit to this. Initially, my patience and tolerance for errors is significantly higher than towards the end of the grading. By the second hour grading, I am not only mentally exhausted my tolerance is significantly lower.
I try to prevent this by creating very explicit grading rubric and I stick to it as much as possible.
Evenly distributing the problem does fix the problem. Proportionality is what matters. Grading being arbitrary is fine if everyone is graded equally.
It’s certainly better than fixed order.
No one's getting hurt by this system if it's randomized. It's a matter of graders giving out partial credit for wrong answers which is discretionary. Rarely students are granted a small mercy. Seems OK.
What do you think is the cause of this? Do you become more cynical (and less generous) because you’ve seen so many BS answers previously? Is it just that getting fatigued makes you less generous?
The fascinating thing was that the distribution of grades was about the same every year.
And I had a math prof for analysis who would give negative points for BS answers. You could say “I need X but don’t know how to prove it” in the middle of a proof, but if you made up something that was incorrect, you’d get negative points.
I feel like only the most obsessive compulsive humans would have this issue (without computer "help"), as the last thing I wanted to do as a TA was to add another step of ordering all the papers before grading them. I also always reviewed the first few papers I graded after grading the rest to make sure I was being fair, because it was obvious to me that until I saw a representative distribution of answers I couldn't do fair grading/marking.
If there's more than one assignment you can basically erase it by randomizing each separately.
If you really care beyond that then randomize for one assignment, flip it for the next, then randomize again for the next etc.
It smell just like every other interesting psychology result that at best is a fluke.
Classic confirmation bias.
I can very easily notice my own over strictness from early in the stack.
For instance the Journal of Personality and Social Psychology [1] is a terrible journal, with a replication success rate in the 20% range. Yet it's ironically well regarded. Both can probably be explained by the exact same phenomena - go read their articles and reads like a stream of bias confirmations for those of a certain ideological orientation -- the same orientation that's clearly widely shared amongst social science researchers.
[1] - https://psycnet.apa.org/PsycARTICLES/journal/psp/126/2
1. Grade problem by problem. This actually makes grading sooo much easier on your own mind
2. Take a second pass to look for outliers in consistency
3. When possible, craft problems that can be automatically graded for correctness. This leaves more time for commentary on the quality of the solution
(I taught computer science, which lends itself to some of this)
The harder bias to handle is the one you develop for students one way or another through the course of a semester or course. Perceived effort shifts grades
Most importantly, none of the researchers is a psychologist or behavioral economist or any kind of "social scientist."