Analyzing student votes across AI models for college essay help
studyarena.com
studyarena.com
The length of the response is a huge factor:
> Students tended to prefer longer responses. The selected answer was 37% longer on average than the alternatives. The longest response won 47.7% of decisive writing comparisons. The shortest still won 25.0%.
So the score is partially a proxy for longest responses.
Makes me wonder how much the reviewers actually read the text. Were lazy evaluators picking the text that looked the longest or most structured without reading it all?
Interestingly we trained a simple LORA layer on top of Inkling and it turns out a model can smell which model wrote a response 70% of the time. I wonder if the smell of gemini is just more preferred by students.
How many of the college students voting have years of ai use under their belt, where ai voice itself has shaped and influenced their stylistic choices of language preference?
Given that this is increasingly the go-to for a college degree, college needs to rethink its cirricula and place in the world. Or at least get rid of the essay.
Lest it become a place where student and teacher ais go to play pay-to-win social deduction video games.
They're adding steps like having the students discuss and defend their essay, which immediately reveals the people who had AI write something and thought they could bluff. This triggers complaints about social anxiety and such, which are unfortunately becoming the go-to defense when unable to discuss the work.
They're also moving toward more in-person writing. Instead of long essays, shorter writing segments as part of the test. Submitting a written essay earns you feedback from the professor and a better understanding of the topic, but that's it.
Why did this ever go away? That was a completely regular part of college a decade plus ago.
I suspect the best schools into the future will integrate lots of socratic defenses of theses, building real things in real time or solving problems with a professor in a case study manner. My last company was a good hint to the future.
When prompted to use simple English without jargon, it's still filled with load bearing honest caveats in every footgun seam it talks about — what I should have led with. <Insert whatever other Claude cliche you prefer>
And I'm not the only one to notice this. Next time it's up for renewal, my team is abandoning it for GH Copilot in order to use literally any other frontier model.
The odd phrases are one thing, what numbs my brain is how everything is evenly bombastic and lacks any sense of rhythm. Technical documents should not read like catchphrases strung together.
I'm curious to know how "Claudish" emerged during the training process. Why each AI has a particular voice if much of the training material is the same across AIs?
That, and my conspiracy hat says it's a way of watermarking the output.
One could argue that Google Colab would supply the training data for better coding performance. I would argue that Colab is mostly used for non-complex (e.g., small number of variables) and self-contained (i.e., runnable in one page) code that can’t train a model for multi-folder and multi-page projects that rely on global connections, which real-life coding would often require.
You may well be right about Google not including the said data but your support for this point is weak.
Malicious acts from large corporations are certainly well known, but they're also grossly over-represented in the news and mindshare. Mostly things just work as they say they work.
I don't really know what proof you're looking for, it's well known that it's hard to prove a negative. I'm not suggesting taking Google (or anyone else) at face-value, but rather saying that when you think about T&Cs, laws, PR, etc, it's unlikely.
Given all that, I can definitely see how Gemini would be preferred. And good for Google, because I would rather my offering be the top choice for the most people rather than better serving a small subset of users.
Fable & Claude Opus 4.x or 5 are terrible to talk to about anything. I gave up on that entirely and just use Anthropic's models for work.
Any less serious technical work I'll use GPT for, as the usage limits are quite fantastic.
I do go back to ChatGPT for things that need to be calculated, but _Gemini simply hallucinates less_. IMO that's whats important for students.
I did get suckered into extending my Grok subscription which is really fun for images/videos. Grok in car is a must too (my jaw dropped when I asked what is pantsula music and it added multiple albums into playlist for me).
Anyone else scratching their head at these descriptions? Uh, your writing is shapeless. You’re 5% too mechanical. Who talks like this? Fiction writers and editors??
at most, the result could be useful for fellow students
I have no doubt other cohorts would rate differently
it is known (on HN at least) e.g. that SWEs tend to prefer brevity, contrary to these students apparently
The data we have is just there to compare for the student base. I'd love to try it out on other cohorts - but the acquisition of such users and the product to make them happy is hard to achieve!
Lots of research shows being able to see different answers improves the quality of a student's output - especially as the good ones are able to pick out what's good from the other work and then add their own insights or takes on top.
...and of course this article itself has a bit of AI-ish tone to it.
E.g. Contrast my human written paragraph vs the AI written first paragraph
Human written:
"The three most popular AI models used by college students are ChatGPT, Gemini, and Claude. As of August 2026 the data on StudyArena shows they prefer Gemini. "
vs
AI Written "Most AI comparsions are written like wine reviews. Claude is subtle. ChatGPT is dependable.."