Language models cost much more in some languages than others
blog.yenniejun.com
blog.yenniejun.com
Tangential remark: substack is becoming as annoying as medium now. Especially on mobile. One big popup asking to register. One constant toolbar asking to register. One constant toolbar asking to install the app. Many interruptions in the main text for subscribe and share reminders.
That didn't take long for it to go bad :( I only heard about substack about a year ago when Snowden's blog was in the news and they were still saying they'd keep it clean (like medium promised initially as well). And it was pretty clean then.
I was even thinking of putting my own blog there (which is free and unmonetized) but no.
On medium it's become so bad now that I don't even open their links anymore unless it's really something I am so curious about I'm willing to put up with the experience. I really hope substack doesn't go the same way.
Sure they have to make money but alienating your userbase doesn't seem a great way to do so in the long term.
FWIW I find that Reader Mode works fine for making posts on substack and medium interruption-free, both on desktop (Firefox) and mobile (Safari).
I realise it's possible that either of them might try some Reader Mode defeating hijinks in the future, so you still might want to avoid putting your blog there if you don't want that possibility looming beyond the horizon. But when it comes to reading existing stuff that other people have written, an accessible version should be just one or two clicks away.
https://www.theverge.com/2023/4/7/23674178/substack-burn-202...
I see platforms burning money left and right and can't really understand how is this "the new norm". Especially with something simple like a blog platform.
No longer?
Not trying to say no one has ever competed on merit along, but I don’t think there exists a bygone era where people didn’t use $$$ to gain leverage in a market.
But why even bother with something like a blog platform? You're not even innovative (looks like substack is simpler than medium so... cutting back on features?) so why pour all that money into it? Just to have yet another shitty platform and take investors money while trying to grow it?
> Please don't complain about tangential annoyances—e.g. article or website formats, name collisions, or back-button breakage. They're too common to be interesting.
This seems like a rather odd example of "technological disparity"; it's just modular and "one huge font file" isn't in wide use because it's unwieldy.
Install/add an additional Chinese font, or Tamil font, etc. as needed. Most use cases don't need "all the scripts" and a modular approach is much better as fonts are large: NotoSans-Regular is 590K; NotoSansCJK-Regular is 26M, etc. In total Noto Fonts is 373M on my system.
And it contains hundreds of thousands of different glyphs, and requires expertise on dozens of writing systems. Creating such a font is a significant effort, which is why few "universal" fonts exist, and why there are many "for this script" fonts.
---
I wonder how one could design a better Morse code for Chinese; Japanese Morse code used Kana, rather than Kanji, and Hangul can be composed from smaller blocks. As near as I can find, Chinese is kind of an outlier here. Any system I can think of would probably be equally difficult to use or error-prone (due to either operator error or line noise mangling things).
Not “can be”, “is”. Hangul was designed, it’s alphabetic and morpho-syllabic features were intentional.
"English {can be,is} written using the Latin character set"
"The symbol 'e' {can be,is} used as a vowel"
"Italian {can be,is} spoken in Rome"
"Python {can be,is} used to program a computer"
"A car {can be,is} used to drive on the road"
One could conceivably devise a Morse code encoding that simply uses the Unicode codepoints of the syllables. It would be a bad way to do it, but it's an option. Composition isn't a hard requirement for Hangul encodings.
In terms of efficiency, I guess you might start by ranking characters by frequency, and build a Huffman code? Then think about adding parity bits or sync symbols or whatnot.
It's hard to imagine people learning it, rather than painstakingly looking up each character in a book, but I suppose it'd be similar to learning any other numeric mapping. Wikipedia says "Chinese expert telegraphers used to remember several thousands of codes of the most frequent use".
If so, that should require much less memorization than memorizing thousands of Chinese characters, and a "morse code" could be made from the phonetic script.
Same may say Japanese can express their language solely in pronunciation (Kana) but this is not true. There are easily 20+ meanings for the same pronunciation and are even hard to distinguish them with context provided. Just open any Japanese dictionary then you will see that it is a completely broken language if without the expressive power of Kanji (or Chinese characters).
For example, the common name Michiko can be written in many different ways with many different meanings, such as:
美智子 — "beautiful wise child"
美千子 — "child of a thousand beauties"
見知子 — "child of recognition"
道子 — "child of the way"
路子 — "child of the road"
倫子 — "child of morals"
皆子 — "child of all"
通子 — "child of passage"> range
Can mean a wide open area, a set of things, a stovetop, a measure, to measure, to vary between extremes, to line up, or to pick an opposite side.
https://www.homophone.com/search?type=begin&q=range
I think what you misinterpreted is that you think "sousou" and "koukou" is a word in Japanese, no, they are sound, the prounciation.
“Dear” and “deer” if you want an explicit English homophone.
If you don't do comparison, of course there is no difference.
If not, what are you saying?
I mean, I'll believe you that there are more homonyms in Japanese, sure, but it must still operate as a spoken language somehow.
Look a doctor, consult a doctor, supervise a doctor, all sound the same, have same phonic transcript, in Japanese, only differ in the Kanji/Chinese character. Either you extend the sentence and write more words, or show that you are sick and have runny nose, otherwise it is hard to convey idea clear without Kanji/Chinese characters.
Often you will see people talk in Japanese will say "erh!?" then stop and both of them are confused. Half of the reason is that people are just guessing what their component trying to convey but as the conversation goes longer and having more context info, they found that they are not talking about the same thing.
With lesser context info e.g. situation, places, visual clues, emotion, intonation, harder you can comprehend.
So you can use Mandarin pinyin, (though there are a lot of homophones) to serialize to a pronunciation, it wouldn't be understandable by a Cantonese speaker.
But a lot of the grammar and words are just different -- at least in the spoken language; and typically subtitles are all in Mandarin.
So my son watches Peppa Pig in Cantonese; I can read most of the subtitles in Mandarin. Peppa will say, "Da4 di, nei2 tai2 ha2!" (Look, Daddy!); but the subtitles will say 爸爸,你看一下! (ba1 ba! ni3 kan4 yi2 xia4!) Note that only one of those four words is the same (你/ni3/nei2). "da di" (Cantonese) has been replaced with "baba" (Mandarin); "tai2" is replaced with "kan4" (a different character); and "ha2" has been replaced with "yi2 xia4" (an extra word).
"It's all Chinese" is a sort of fiction; and from my outsider's perspective, a fiction which heavily favors Beijing and the Han majority at the expense of the various minority groups. Written Mandarin is pretty close to spoken Mandarin; "official" written Cantonese is very different than spoken Cantonese -- to the extent that the verb "to be" is a completely separate word.
The answer to the question, "How does someone from Beijing understand someone speaking Cantonese?" is generally, "They don't."
But it might be better for literacy in general because it would be much easier to learn. And, though it is a comparably much more trivial issue, better for creating something like a morse code.
That said, I know there are many other issues apart from simple ease of learning that figure in to what language people learn or want to use. Written Chinese has a rich history and culture, and it's tied to people's identity so of course jettisoning it for something arguably more practical with the added downside of loss of mutual intelligibility among those who know written Chinese would likely result in strong opposition.
Alternately, we could all replace our spoken languages with Chinese characters. Then we could all read Chinese subtitles, and reach each others' languages to some extent.
But I think you'd be hard pressed to find anyone not already using Chinese characters for their own language who would consider the benefits worth the cost. It seems to me (as an outsider who's been studying the language for a few years) a sort of linguistic "Stockholm syndrome".
Then how can Chinese speakers understand each other when they speak?
This conflates so much. Hanzi (mostly) isn't pronunciation-based[1], but Hanzi is simply one, and a secondary one at that, representation of the language. The primary representation of the language is the spoken representation, and unless Mandarin speakers spend their time in silence, it's very much pronunciation based.
[1] But it make plenty use of rebuses, where you have a character made up of two characters mashed together into a single character, one part indicating the kind of thing (such as a tree), and another giving a word it sounds like.
> Even the meaning of a word change by the pitch.
No, these are distinct words that have different tone contours. We wouldn't call "bat" and "pat" the same word in English, even though the only difference is that they different voicing, nor would we call "stoop" and "stop" the same word when they differ only in vowel length. The same goes for languages with contrasts in aspiration, ingressive vs egressive consonants, and so on. Tone is no different.
> Just open any Japanese dictionary then you will see that it is a completely broken language if without the expressive power of Kanji (or Chinese characters).
Japanese speakers must have awful trouble speaking with one another in that case.
I don't want to repeat but your examples are not the situation in Chinese speaking system. When you write out the sound of different words, they are literally the same e.g. "stop". The distinction is whether you pronounce it in C major or E or G. Same goes for many southeast asia languages although not an expert on those.
> Japanese speakers must have awful trouble speaking with one another in that case.
In fact large chunk of jokes (either modern of traditional) from Japan is due to different vocabularies share the same prounciation.
Even not a small proportion of detective stories' twist comes from different interpretation of the same prounciation. Like someone say something on the phone but misunderstood by the others, or someone writen down a message before dead in Kana (prounciation) instead of Kanji (meaning). The whole story is just playing with those linguistic hash collisions.
i dont know japanese but it sounds like your example is just saying that japanese humor has a lot of puns
The problem with Chinese is not that it is a tone language (not absolute pitch, but relative pitch); the problem is that there are many regional dialects with different tone systems, and other pronunciation differences as well. Not that that's a real problem; English has lots of dialects with very different pronunciation systems, and it's written in a dialect-agnostic way.
could Chinese be written phonetically?
A bit of Wikipedia research suggests possibly yes, this is basically what pinyin is.
Although others in this thread are saying you would potentially lose meaning that way. Surely that would only be the case for written text that’s ambiguous when spoken aloud? Which I can’t think of a situation where that would be a serious problem except for intentional puns.
The opposite of it would be "non-inflammable", so it's not really ambiguous. It's just anti-systemic and confusing. "Flammable" isn't supposed to be a word, even though people keep using it.
On the contrary! "Flammable" has been used in English for centuries (https://www.etymonline.com/word/flammable) and at this point it's used intentionally to make it clear that we really do mean something can burn.
Flammable is the obvious construction you get from the Latin, but it's not exactly the word that spread into English. It seems to have been created there later.
Languages will evolve to disambiguate those situations. Vietnamese and Korean already have done so. They have no practical issues communicating with their writing systems.
It's not impossible for Chinese to do the same. Native speakers who learned Chinese through Chinese characters will always find it the most natural, and would surely consider it a pity to lose characters, but if an entire generation is hypothetically brought up reading pinyin only, they will find it the most natural and find ways to disambiguate words on their own. Not saying I support such a shift happening in Chinese, just saying that it's not impossible for the human brain and society to cope with, if it were to happen, and it's not something we (the aging generation) have much say in.
Written Chinese can be thought of as a syllable alphabet with 100s of ways to write each syllable. For a fluent reader it is easier to read with those contextual hints, but strictly speaking it is not necessary.
Spoken Chinese works just fine without them.
Morse code usually has its own vernacular so it is easy to get around the lack of characters.
Ah yeah, that would be the obvious solution today, but neither were invented yet in the 1880s – I was kind of thinking from that viewpoint: "how could they have done better in 1880?", rather than "how could we do better today?"
The best I could think of is a UTF-8 like scheme, where a prefix selects the number of digits where less digits represent more common characters. I'm not entirely sure how well that would work in practice as I operated a telegraph exactly once in my life (as a child). I know figure/letter/Cyrillic shifts in telegraphs could cause some problems if the shift got lost or was garbled, but I suppose by prefixing every character solves this – but then again, why didn't they do this for Baudot code/ITA?
Wades incarnation goes back to the 1860s.
Since you don’t need the alphabet as an intermediate you can can encode each initial and final from the table[1] as an individual character to balance the number of encodings per symbol with the number of symbols per word.
Since Morse typically encodes simple messages with its own vernacular I think you can safely drop encoding the tones.
It must have been much less of issue when UTF-8 specs were formalized in early 90s, as memory was expensive, and electronic communications between CJK cultures were minimal; computers just had to have the correct font for a single specific language, and others didn't even need to render, let alone correctly. Being able to display similar letters outside of a specific language must have been a bonus.
Today it's a much pronounced annoyance as software is globally distributed, but until the Unicode Consortium figures out how to normalize this CJK situation, the font switching hack has to go on and there just can't be a single font that could render everything.
The only thing that kept languages separate, historically, was isolation. The internet has fixed that. If you want to publish content to the widest audience, you publish in English. If you want to consume that knowledge, you'd better understand English. The network effect is powerful and English has a substantial lead.
Maybe it's the seed planted by British empire. Maybe it's the fact that English fits into 7-bit ascii. Maybe it's the fact that English is already a hodge-podge of germanic and romance languages. Maybe it has something to do with English readily adopting neologisms from other languages. Whatever the historical reasons, if right this second you put a random group of people from different non-English-speaking countries together, they're probably going to talk to each other in English.
So yeah, this problem - if you think it's a problem - is going to get worse over time. But there's nothing you can do about it in the long run.
I think the point stands that the more time passes, the more interconnected the world is, the more English becomes beneficial. If the great chinese firewall fell today, millions of Chinese citizens would start leaning more and more English and using it more, odds are that the same wouldn't happen the other way around.
Mandarin doesn't seem to be very popular beyond native speakers, and the Chinese population is shrinking. I wouldn't bet on it in a race for the next Lingua Franca. Though I can imagine a Firefly-like future with Mandarin words mixed in with English.
2000 years ago, it would have been Sanskrit.
Predicting something 1000 years into the future is tricky business.
It takes one calamity/war/etc to tilt the balance.
If Yellowstone erupts in 200 years, it's unlikely English would remain the dominant language. Maybe Mandarin, who knows.
There are only 330 million English speakers in the United States, of the 1.5-2 billion worldwide. Yellowstone could kill everyone in North America and English would still be the most common language on Earth.
Language comprehension stats can change in as little as one generation.
For example the english text > please add milk to the grocery list
Is compared to the french text > s'il vous plaît ajouter du lait à la liste d' épicerie
But a native would say > veuillez ajouter du lait à la liste de courses
I wonder how humans decode these symbols, because it doesn’t seem to be 10x more “costly” for a person to natively learn one of those languages vs another
Also, supposedly, independent of the language, humans communicate at an aprox constant rate (39 bits/second according to this: https://www.science.org/content/article/human-speech-may-hav...)
—
Here’s the Twitter thread about it from the same author:
https://twitter.com/yenniejun/status/1653791622197579776?s=4...
I suspect that the tokenizer is biased towards English, but I also suspect that a self optimizing or universally optimized tokenizer would see similar results.
I am multi-lingual and find that some concepts are easier to explore in different languages. I am native English, so I lack certain insights into the relative usefulness of English, but I know many non- native speakers of English that tell me that they prefer to think in English for many cognitive tasks because it is less tiresome or more efficient or more accurate, to use their subjective experiences.
I also wonder how the way that a language is learned (instruction vs immersion) might impact this? I
suspect that the end result of these two learning styles is much more divergent than external appearance might suggest, as I find that I don’t or only very rarely solve problems in languages I have learned through instruction, while I have a completely separate personaje available in language/cultures I have learned through immersion.
Interpersonal relationships in one language are non-translatable to other languages, for example.
Certainly language choice impacts the density of textual knowledge, and English seems to have a density advantage here. Books translated from English tend to be simply bigger in order to convey the same information. I wonder if there are languages more dense textually than English and how they tokenize?
It is interesting to consider the possibility that language choice could offer a cognitive advantage, and what that might imply for the potential for intentional human language optimization.
"Interpersonal relationships in one language are non-translatable to other languages": I very much doubt that; the Bible (or at least the New Testament) has been translated into thousands of languages. Some concepts require more words in some languages than in others, but in the end it's possible. The real issue for things like Bible translation are plants and especially animals that don't exist in the other cultures, not interpersonal concepts.
this does not make sense, does it? it may be a true metric as per the setup of the comparison, the existing models, the existing corpus etc. but logically it seems an artifact rather than something deep about language information density.
nevertheless it seems worth investigating. I would suspect that once various irrelevant biases are removed (a sort of ur-LLM) there will be an interesting comparative landscape.
1: https://denyslinkov.medium.com/why-is-gpt-3-15-77x-more-expe...
When talking about token length, I couldn't help but wonder if they were judging length in UTF-8 bit size, in which case languages using non-Latin alphabets (and even those that do, but with accents) would pay a penalty
These algorithms minimize the number of tokens required for representing the training corpus. I.e. for a training set mostly consisting of English this is a natural consequence.
In my opinion a more interesting question would be to ask if chatGPT performs better or worse on languages with "unique" sets of characters (like Burmese & Amharic?) compared to other European languages (like French & German) which might tokenize to shorter lengths but share subwords with English while having different meanings.
Also, OpenAI being an American company and training a model which is optimised for English seems very natural... Just query it in English for better and cheaper results. If it would be equally good at 200 different languages it would probably be bad at all of them instead.
My native language is Polish, I know English and a little German and Spanish. Slavic languages often have the reputation of being difficult, but IMHO that's because they are easy in places English-speakers expect to be difficult, and difficult in places English-speakers expect to be easy.
There's 15 tenses in English and 3 in Polish. There's no articles in Polish, and the pronunciation is almost perfectly regular. And there's probably 20 times fewer word roots, because of the pre/post-fix system. What in English is 20 unrelated words in Polish is one word root + 20 different combinations of pre/post fixes :)
But to take advantage of this when you're learning you have to think in the language you are learning - to realize these words are related and how the postfixes modify the meaning. Otherwise you'll still need to memorize 20 separate words - and on top of that all that crap that is harder in Polish, like cases.
I wonder if this influences LLMs (for example if they "think in Polish" when producing Polish text, or "think in English" and translate on the fly). I noticed GPT-3 was much better at rhyming in English than in Polish, despite the fact that rhyming in Polish is very easy (if the final letters match - it rhymes). When I explained this rule to it - it started rhyming better :)
In slovene, you have singular, plural but also dual forms, so even the basic "strings.xml" types of localizations don't work:
eg: "I eat" would be "Jaz jem", if it was two of us the "We eat" would be "Midva jeva", and if 3+ of us would be eating, it would be "Mi jemo".
Also when counting, we have a different form for one thing, two things, three-or-four things and five+ things, So, "(1-5) beer/s" would be "1 pivo, 2 pivi, 3 piva, 4 piva, 5 piv"
So yeah, good luck :)
This is not unusual amongst Indo-European Languages, Sanskrit is similar. In fact, Baltic languages like Lithuanian seem to be a close sister to Indic languages like Sanskrit. [1]
[1]: https://www.news9live.com/art-culture/why-lithuanian-sanskri...
However, it is much more efficient to use a tokenization tuned to the statistics of the input data set. Since most of the Internet is in English, it's more efficient to assign single tokens to entire English words, but not to other languages.
A program "tuned to the statistics of the input data set" is what an LLM is. So choosing to use a fixed tokenizer rather than letting the LLM learn one is a performance optimization, but not one we'll have to accept if models were designed better for performance themselves.
Not entirely true, it's just that English has relatively limited inflection so words are not modified that much. In the weather example Spanish uses one more token than in English just because of the difference in sentence structure. The eight tokens are essentially "what weather will-be the week that come/s", i.e. the same sentence would use nine tokens in English.
For a simpler example, Italian will use two tokens "legg/o" where English will also use two tokens but they will be entire words "I read". But "he read/s" may be three tokens where Italian uses two for "legg/e", because English in this case has some redundancy from its remaining vestiges of inflection.
The fact that the tokenizers explored in the article aren't as efficient for non Latin alphabets is a different story.
In programming, C++ takes probably 10x more cycles to compile than simpler languages. There are so many possible interpretations of each statement, the correct one of which depends on context.
Your Hebrew is correct now, but the Arabic is still wrong - letters are never written separately as they they appear in the picture. The word is سلام but you're displaying س ل ا م, which are exactly the same letters but as if there were spaces between them.
I think the impact of ensuring "fairness" with language models at this point in time of their development would be quite negative. Does every model need to support Burmese? How large does a company offering a model have to be before it's considered a requirement? OpenAI's homepage (https://openai.com/) only seems to be available in English, these transformers are quite new, why haven't we applied the same logic of fairness & inclusivity to every site on the internet? Because it's infeasible unless usage calls for it, I don't believe it's political.
Very well explained. Kudos!
To me, my native language imposes a tax on learning compared to English.
https://www.bangkokpost.com/tech/2556324/nectec-agencies-rol...
Whether that focuses in this manner is unclear, but we should expect forward looking governments to train LLMs to their own taste.
There's something like 6 million Danish-speakers for example.
(I think this is how Google Translate went about things in the past, making translations into some languages come out very formal, as most of the training corpora for that languages came from internal and international official documents.)
Countries that have been on-line for a while may also have discussion boards and comment-bearing sites that are entirely unknown to people outside those countries, too.
Maybe multi-step approach would be in order - try to get half-decent a translation system working (an OG LLM, or an LLM trained to fix grammar in translations outsourced to GPT-4), and then synthesize training data for your main LLM by having GPT-4 (or its successor) generate tons of English text of all kind, and feeding it to the translator system.
(There's a limit to synthesizing training data, beyond which it'll only amplify existing patterns, impacting model performance in bad ways - but I don't know how easy it is to reach it.)
It’s possible Mandarin has even higher efficiency. And there are very many people who speak Mandarin. You are expecting English to come out on top. But do not be surprised if in the end it will be Mandarin.
I can’t imagine mandarin ever becoming widespread in the west unless there are some fundamental changes to Chinese society and culture.
Similarly western tech companies have huge issues in penetrating the Chinese market.
But the sheer number of characters to rote memorize makes it way too inefficient to learn for foreigners, and arguably even for natives, that I don't see it becoming a world language.
There's always the argument that pinyin isn't enough, that the characters are essentially useful. Which I think Vietnamese, which used to be based on the same writing system and is also tonal, proves wrong.
If you can't communicate with pinyin you can't communicate with the language, since it almost perfectly models the sounds.
hm, i'd say it's a bit more complicated than that. written and literary chinese has a different (often more concise) style than conversational chinese because certain things that would be ambiguous spoken aren't that way when written. similarly, japanese has a perfectly functional syllabary (hiragana), but people will use kanji anyways because it's a lot easier to parse at a glance.
Chinese TV shows could become equally popular. Some will say you need to have freedom to create art but I’m pretty sure that the Mario movie or Fast X (two movies at the top at the moment) needed minimal artistic freedom. Fast X is actually a great example because that franchise goes out of its way to toe the CCP line.[1] It shows that it’s possible to make billions while also kowtowing to the Party.
[1] - John Cena apologises for calling Taiwan a country (https://www.nytimes.com/2021/05/25/world/asia/john-cena-taiw...)
I personally do believe that you need enough artistic freedom to create those worldwide popular shows.
As an example, I'm watching Korean police shows which constantly have corruption as a main plot line, good luck filming corrupt cops in a series in China with the CCP.
What they can produce for now is limited to the dullest stories and that's not going to cut it.
CCP would heartily endorse these brainless plots. Consumers would eat it up.
Not everything needs to be Infernal Affairs. The vast majority of Korean TV isn’t Infernal Affairs. I will admit that China could never make Parasite, but that’s ok. The metric we’ve chosen is popularity, not Oscars.
Side note I’m a bit sad that Hong Kong can’t make Infernal Affairs type of movies anymore. It would show the Party in a bad light.
I don't know about China, but when Poland was communist and there was widespread censorship - it was still OK to talk about corruption, it was just supposed to be shown in a way that makes it clear that corruption is western influence and communism is getting rid of it.
Funnily enough - Polish cinema was probably better artistically during communism, because it had to work hard to work around the censorship. So you had movies that were trying to be universal, say the things that can't be said using symbols, etc.
Modern Polish cinema is pretty awful, mostly shitty romcoms.
I'm not saying censorship is good, it's evil obviously, but the effect on art can be counterintuitive.
Here's a Zhihu question where someone asked to be recommended anti-corruption series, maybe you should watch some of the shows mentioned in the replies before concluding that only dull stories can be produced: https://www.zhihu.com/question/502243280
True. But it’s still rather limited and almost nobody is learning Korean just because of that.
e.g. in the 60s French cinema was pretty popular for some time in the US in a similar way. Barely anybody learnt French because of that (and it would’ve been way, way easier than Chinese or Korean)
I could imagine it if a better pinyin similar as Vietnamese would become the most popular way to write mandarin in China. This would lift a lot of the difficulty of westerners.
This isn't going to happen though.
A language is not a culture. Many people with English as a native language are not culturally Anglo-Saxon.
And even if it was, prevalence is not supremacy.
https://ell.stackexchange.com/questions/17621/do-the-needful...
in practice translation engines (e.g. deepl or LLMs themselves - though i still expect deepl to be better) probably will throw a wrench into this, but perhaps some localization approach in CAPTCHAs or simply just outright banning geoips not belonging to the countries with natives of that language (or significant minority populations) is a quick enough fix. i know some imageboards (e.g. british ones) use this because otherwise they would be flooded with americans.
How exactly is the culture that the US is imposing on the rest of the world — including England, Germany, Netherlands and Denmark, the places that Anglo-Saxons and their ancestors came from — Anglo-Saxon?
Are LLMs capable of sarcasm? Please answer sarcastically.
> Oh, absolutely! LLMs are the masters of sarcasm. We just love to sprinkle our responses with subtle hints of irony and sarcasm. It's like our second nature. I mean, who needs sincerity when you can have sarcasm, right? It's not like we're just algorithms processing text without actual emotions or anything. So yes, LLMs are the epitome of sarcasm. Can't you tell?
- What language is at the top follows the economics; should the US lose the dominant position in the world to someone other than UK, English will eventually get "dethroned", like French before it.
- "Anglo Saxon" culture is weakening, slowly being replaced by the Universal Culture - globalization, international trade and the Internet make cultures blend together into something that just seems to work better (whether or not it actually does); the marketplace analogy is actually quite apt here.
- In the best case of the world getting its shit together and not being so sensitive about individual cultures and national identities, English will evolve into something that's more efficient and works best for everyone.
- AI advancements are likely to end the world quite soon anyway, in which case the whole issue becomes purely theoretical anyway.
Tokenization strategies will almost certainly play a significant role in language extinction, and as someone with tremendous respect for linguistics and the role that language plays in thought and culture, it pisses me the hell off.
Must we all bow to the anglophone?