We all know the "AI-tics" that give away a sloppily AI-written piece, but even if you steer them, they still struggle to write consistently high-quality prose.
Somehow I feel that the work of a good copywriter has never been more noticeable.
We all know the "AI-tics" that give away a sloppily AI-written piece, but even if you steer them, they still struggle to write consistently high-quality prose.
Somehow I feel that the work of a good copywriter has never been more noticeable.
(I may be projecting real life experience onto the LLM.)
I am genuinely fascinated as to how Claude acquired its utterly aggravating way of writing. It’s so much more irritating than ChatGPT, which is already not good.
I have education/experience in both literature and coding, I have a pragmatic starting point when approaching text while also being able to recognize stylistic oddities, so I get to be the guy editing out AIsms sometimes.
But I've helped other people copywrite where their environment was all org-speak and academic writing, and AIsms don't really stand out in that case. AI is effectively "doing the right thing" writing the way it does for those tasks. Even tho the right thing is often a bad thing.
Diversity of writing styles was part of that, but I’d point to vernacular exposure as the larger component. We’re going to go through a period where we try to adapt to a form of “Universal English” for those of us who read primarily English writing.
Other languages’ readers may be experiencing the same dissonance when they come across AI-generated prose in their native language (but I’ll let others validate /reject my hypothesis).
I think there's also a lot of recency bias in it. The same thing that makes you suddenly notice how many people are driving the same model car you just looked at, or how many ads there are for Turbo Encabulators after you read an article about them. A lot of the "tells" people picked out in early AI are tells because they're also really common in the material that the AI was trained on and the styles it was made to emulate. But until everyone wanted "one quick trick" to pick out AI writings, people didn't have any particular reason to need to notice those tells and so they slipped under the radar.
But when I submit a novel with em-dashes, will sloppy agents and sloppy editors be able to tell that em-dash was deliberately put there by me?
All these HN posts with "stop with the LLM generated garbage" will slowly fade away because both people and LLMs will talk in a similar format and will be inured to the the distinguishing feature.
Certainly "humans" will keep carving out distinguishing characteristics, but just like "corporate speak" is a thing, AIsm will be a thing.
That's how language spreads.
Though this might be me as a British reader, simply preferring a rather less American turn of phrase.
I reckon the more transatlantic, english-as-international language DeepMind team have had a subliminal (or maybe deliberate) impact on the way it chooses to write.
Or perhaps small open weights models simply aren't under the same commercial pressure to be engaging and sycophantic and are therefore less likely to adopt the samey overly casual, upbeat, Californian sales assistant manner. (Don't get me wrong, I like this from real human Californians just fine!)
Either way, the default tone is much less showy. I would be interested to find out if you agree.
I am very much an LLM cynic. I am engaging because I must, and trying to learn fundamentals, but I would not say I am overly excited by any of this, just glad that small open weights models exist as a counterpoint.
I loathe the way ChatGPT writes, and the Claude-isms that are everywhere; it is actually quite enraging, especially when you start seeing it in internet comments from people who used to try to write out their own thoughts.
But in my experiments with open weights models I have found I am much less aggravated by summaries and outlines written by Gemma 4, so much that I am happy enough to read them, because they have fewer irritants that take me out of the reading flow.
Though this evening it told me very kindly that my photography is a bit "safe". How very dare it… understand me that well.
This seems to be a hard problem for LLMs, as passing would probably require good self-perception ("oh no, I am writing like an AI!") and fine-grained control over its own output ("let's write like a human instead!").
There's a lot of work in the humanities about different aspects of good writing, but that's not quite the same thing. And anyway they tend to assume a pre-existing level of writing ability. Students are supposed to learn good writing through practice; there are rules and exercises but they're incomplete.
It’s much easier to understand this once you think about other generative forms. MidJourney never just sits down and draws for fun, so fun never informs its art (only the outward appearance of others’ fun, separate from the fun itself). Suno doesn’t waste hours trying to find riffs on a guitar, so its output is never informed by the direct joy of getting it right. Its music is never optimised for playability on a particular guitar with a scratchy seventh fret and a too-high action. Neither Midjourney nor Suno have evolved their styles due to short-sightedness or carpal tunnel.
If you had a human writer who over a long career only ever wrote articles from an outline given to them by someone else, and you had all the outlines and all the resulting articles from those outlines, and you could train an LLM to generate an article from an outline, it still would not be kicking itself frustrated by an inelegant phrase in a prior article, it would not avoid certain phrases out of a passive aggressive reaction to some editor’s note, it would not ever just rush an article because everyone is gathering at the pub, and it would not choose an analogy just to rub the author of a bitchy critical letter to the editor the wrong way. An LLM could not “subtweet”. It could not write a series of articles hoping one important person will spot that they are auditioning for a job.
Creators have unseen, undocumented influences and motivations that inform their work over a long period. I don’t mean to say that these individual influences can be reliably detected in individual pieces of work. I do mean to say that I think their broad absence tends to be felt in LLM writing. As readers we develop an affinity for writers as much as for their writing, and we do this in part because we deduce things about them.
However, general purpose LLMs like Fable have been trained on huge amounts of all kinds of data, and therefore find it exceedingly hard to break out of the grooves carved by that data. They can’t avoid defaulting to centroids and averages, even when they are trying not to. This makes it possible for classifiers like Pangram to discriminate their writing.
A plausible way to work around this limitation would be to train a LLM on a limited and cohesive subset of writing materials, so it would absorb their specific writing style.
One example might be Talkie, a LLM trained on pre-1930’s English text. Talkie is a far smaller and less powerful model than Fable.
And yet, Talkie’s writing is so distinctive that it is often classified as human by Pangram.
Of course, the amount of self published work has also helped make the lack of a good copy editor noticeable. I can excuse self published blogs though. But the stuff released "professionally" has really become farcical.