LLM text just plain doesn't do this: they're very good at writing perfectly formed English, but it just winds up saying nothing (and models like ChatGPT have been optimized so they end up having a particular voice they speak in as well).
LLM text just plain doesn't do this: they're very good at writing perfectly formed English, but it just winds up saying nothing (and models like ChatGPT have been optimized so they end up having a particular voice they speak in as well).
This. My partner always speaks frenglish (french english) after talking to her parents. You have to know a little French to understand her sentences. They’re all English words, but the phraseology is all French.
I do the same with Slovenian. The words are all English, but the shape is Slovenian. It adds a lot of soul to your words.
It can also be topic dependent. When I describe memories from home in English, the language sounds more Slovenian. Likewise when I talk about American stuff to my parents, my Slovenian sounds more English.
ChatGPT would lose all that color.
Read Man In The High Castle to see this for yourself. Whole book is English but you can tell the different nationalities of each character because the shape of their English changes. Philip K Dick used this masterfully.
Amusingly, I think this phrase illustrates your point. To the best of my knowledge, a native speaker (which I'm not) would always say "The whole book is (in?) English", leaving off articles seems to be very common for Slavic people (since I believe you don't really have them in your languages).
And if what it does now is unimpressive, it might be a good thing to use to monitor the rapid progress of LLMs.
Whenever I come across text that has a lot of missing articles, the voice inside my head automatically changes to a Russian accent; and in the instances where I've bothered to find out the author, it was always someone from Russia or some other ex-USSR country, so it seems I've already ingrained this characteristic at a subconscious level.
LLM output to me usually sounds very sanitised style-wise (not just content-wise), some sort of lowest-common-denominator language, which is probably why it sounds so corporate-y. I guess you can influence the style by clever prompt engineering, but I doubt you'd get a very unique style this way.
LLM's output the statistically most average sequence of tokens (there's no intelligence there, "artificial" or otherwise), so yeah, that's by design.
You fundamentally misunderstand how this works.
The LLMs learn the various grammars and "accents" implicitly. They automatically differentiate these grammars.
Sounds like you still have this idea that LLMs are a giant Markov chain. They are not Markov chains.
They are deep neural networks with hundreds of layers and they automatically model relations at extremely deep levels of abstraction.
> they automatically model relations
No, they do not model anything at all. If you follow the tech bubble turtles all the way down you find a maximum likelihood logistic approximation.
I know, I know - then you'll do a sleight of hand and claim that all intelligence and modeling is also just maximum likelihood, even thought it's patently and obviously untrue.
Large Language Model (LLM).
Hundreds of layers with a trillion weights and you think "nothing is modelled" there. The comments on this site are ridiculous.
Studies have traced individual "neurons" in LLMs that represent specific concepts. It's not even debatable at this point.
But realistically, how many people are going to actually do that? Communication fed through an LLM represents a rather bleak linguistic convergence.
Those would be _really_ rough translations. Yes, I've seen "It's an achieve my dream's place" written, but that was in an essay written for high school.
And of course you could build a corpus of text written by Chinese English speakers for more authenticity.