Right on point. I find particularly striking how little is said about whether the best students achieve the best grades. Authors are even candid that different LLMs asses differently, but seem to conclude that LLMs converging after a few rounds of cross reviews indicate they are plausible so who cares. The apparences are safe.
The objections to SFH exist and are strikingly similar to objections to WFH, but the economics are different. Some universities already see value in offering that option, and they (of course) leave it to the faculty to deal with the consequences.
Even for distance education though, proctored testing centers have been around longer than the internet.
It is about a third of the students I teach, which amounts to several hundreds per term. It may be niche, but it is not insignificant, and definitely a problem for some of us.
> Even for distance education though, proctored testing centers have been around longer than the internet.
I don't know how much experience you have with those. Mine is extensive enough that I have a personal opinion that they are not scalable (which is the focus of the comment I was replying to). If you have hundreds of students disseminated around the world, organising a proctored exam is a logistical challenge.
It is not a problem at many universities yet, because they haven't jumped on the bandwagon. However domestic markets are becoming saturated, visas are harder to get for international students, and there is a demand for online education. I would be surprised that it doesn't develop more in the near future.
I think the end result though is that schools either limit their students to a smaller number of locations where they can have proctored exams, or they don’t and they effectively lose their credentialing value.
In 1789 there were 1,000 enrolled college students total, in a country of 2.8M. In 2025, it is 19M students in a country of 340M. https://educationalpolicy.org/wp-content/uploads/2025/11/251...
In 1950, 5.5% of adults ages 25-34 had completed a 4 year college degree. In 2018, it was 39%. https://www.highereddatastories.com/2019/08/changes-in-educa...
With attendance increasing at this rate (not to mention the exploding costs of tuition), it seems possible that the methods need to change as well.
Some people dream that technology (preferably duly packaged by for-profit SV concerns) can and will eventually solve each and every problem in the world; unfortunately what education boils down to is good, old-fashioned teaching. By teachers. Nothing whatsoever replaces a good, talented, and attentive teacher, all the technologies in the world, from planetariums to manim, can only augment a good teacher.
Grading students with LLMs is already tone-deaf, but presenting this trainwreck of a result and framing it as any sort of success... Let's just say it reeks of 2025.
If a student is willing and desire to learn, an LLM is better than a bad teacher.
If a student doesn't want to learn, and is instead being forced to (either as a minor, or via certification required to obtain work & money), then they have every incentive to cheat. An LLM is insufficient in this case - a teacher is both the enforcer and the tutor in this case.
There's also nothing wrong with a teacher using an LLM to help with the grading imho.
At least in Germany, if there are only 36 students in a class, usually oral exams are used because in this case oral exams are typically more efficient. For written exams, more like 200-600 students in a class is the common situation.
If you are going to set an exam that can be graded in 5-10 min, you are not getting a lot of signal out of it.
I wanted to do oral exams, but they are much more exhausting for the prof. Nominally, each student is with you for 30 min, but (1) you need to think of slightly different question for each student (2) you need to squeeze all the exams in only a couple of days to avoid giving later students too much extra time to prepare.
That's entirely false; this is why we have multiple-choice tests.
Thinking deeper, though, multiple choice tests require SIGNIFICANTLY more preparation. I would go so far as to say almost all individual professors are completely unqualified to write valid multiple choice tests.
The time investment in multiple choice comes at the start - 12 hours writing it instead of 12 hours grading it - but it’s still a lot of time and frankly there is only very general feedback on student misunderstandings.
I don't believe that your argument is more than an ad-hoc value judgment lacking justification. And it's obvious that if you think so little of your colleagues, that they would also struggle to implement AI tests.
- They have a base marks of 20-25% (by random guessing) instead of 0.
- You never see the working. So you can't check if students are thinking correctly. Slightly wrong thinking can get you right answers.
- They don't even remotely reflect real life at all. Written worked through problems on the other hand - I still do those in my professional life as a scientist all the time. It's just that I am setting the questions for myself.
- The format doesn't allow for extended thought questions.
In my undergrad, I had some excellent profs who would set long work through exam question in such a way that you learned something even in the exams. Simply a joy taking those exams that gave a comprehensive walk through of the course. As a prof, I have always tried to replicate that.
One of these is not like the others.
Work study and TA jobs were abundant when I was in college. It wasn't a problem in the past and shouldn't be a problem now.
> In our new "AI/ML Product Management" class, the "pre-case" submissions (short assignments meant to prepare students for class discussion) were looking suspiciously good. Not "strong student" good. More like "this reads like a McKinsey memo that went through three rounds of editing," good...Many students who had submitted thoughtful, well-structured work could not explain basic choices in their own submission after two follow-up questions. Some could not participate at all...Oral exams are a natural response. They force real-time reasoning, application to novel prompts, and defense of actual decisions. The problem? Oral exams are a logistical nightmare. You cannot run them for a large class without turning the final exam period into a month-long hostage situation.
Written exams do not do the same thing. You can't say 'just do a written exam'. So sure, the students may prefer them, but so what? That's apples and oranges.