It’s strange how AI style is so easy to spot. If LLMs just follow the style that they encountered most frequently during training, wouldn’t that mean that their style would be especially hard to spot?
It’s strange how AI style is so easy to spot. If LLMs just follow the style that they encountered most frequently during training, wouldn’t that mean that their style would be especially hard to spot?
Or someone made a call that emoji-infested text is "friendlier" and tuned the model to be "friendlier."
(I once got that feedback from someone in management when writing a proposal...)
First, LLM style did not even exist, it's a match of several different styles, choice of words and phrases.
Second, LLM has turned a slight plurality into a 100% exclusivity.
Say, there are 20 different choices to say the same thing. They are more or less evenly distributed, one of them is a slightly more common. LLM chooses the most common one. This means that
situation before : 20 options, 5% frequency each
situation now : 1 option, 100% frequency
LLM text is both reducing the variety and increases the absolute frequency drastically.I think these 2 theories explain how can LLM both sound bad, and "be the most common stye, how humans have always talked" (it isn't).
Also, if the second theory is true, that is, LLM style is not very frequent among humans, that means that if you see someone on the internet that talks like an LLM, he probably is one.
Later the cutest of the emojis paved their way into templates used by bots and tools, and it exploded like colorful vomit confetti all over the internets.
When I see this emojiful text, my first association is not with an LLM, but with a lumberjack-bearded hipster wearing thick-framed fake glasses and tight garish clothes, rolling on a segway or an equivalent machine while sipping a soy latte.
Jk, your comments don't seem at all to me like AI. I don't see how that could even be suggested
Your comment however is just an ad hominem.
Glasses: check (I'm old)
Garish clothes: check
Segway: nope
So there's a 75% chance I am a Millenial hipster. Soy latte: sounds kinda nice
It's not because they can't write PRs indistinguishable from humans, or can't write code without Emojis. It's because they don't want to freak out the general public so they have essentially poisoned the models to stave off regulation a little bit longer.
That's a lot of expensive work they're doing, and ignoring, if they're just later poisoning the models!
I'm like "Sure buddy, sure. And the nanobots are in all vaccines, right?"
Of course not! They would use it to trade and would keep it concealed while throwing the public a bone with a less advanced version. Same thing applies as AGI or even as code gen gets better.
How likely is it that you saw a plane made by humans that you misclassified due to lack of data vs how likely is it that it was actually an Alien made aircraft, coincidentally working in our atmospheric conditions, gravity and drag related physics?
Guess what, it's about likeliness and you are extrapolating the wrong assumptions. Extraordinary claims require extraordinary evidence, not the other way around.
But maybe it originated somewhere else.. In Javascript libraries..?
Why do you think that? I try to stay involved in accessibility community (if that's what you mean by inclusive?) and I've not heard anyone advocate for emojis over text?
I say "meme" because I believe this is how the information spreads — I think people in that particular clique suggest it to each other and it becomes a form of in-group signalling rather than an earnest attempt to improve the accessibility of information.
I'm wary now of straying into argumentum ad ignorantiam territory, but I think my observation is consistent with yours insofar as the "inclusivity" community I'm referring to doesn't have much overlap with the accessibility community; the latter being more an applied science project, and the former being more about humanities and social theory.
But agree excessive emoji's, tables of things, and just being overly verbose are tells for me anymore.
The „average“ style, from the Unix manpages from the 1960s through the Linux Documentation Project all the way to the latest super-hip JavaScript isEven emoji vomit README must still have been relatively tame I assume.
God Bless.
So it's entirely possible that training in one area (eg: Reddit discourse) might influence other areas (such as PRs)
They didn't learn how to write PRs. They "learned" how to write text.
Just like generic images coming out of OpenAI have the same style and yellow tint, so does text. It averages down to a basic tiktok/threads/whatever comment.
Plus whatever bias training sets and methodology introduced