Also, this really should be more precise that it’s talking about neo-caveman. Legit caveman no speak English.
Also, this really should be more precise that it’s talking about neo-caveman. Legit caveman no speak English.
你不去,我也不去
You no go, I also no go
今天这里人很多
Today here person very many
下雨就不去
Fall rain then no go
Vulgar latin does correspond to modern romance languages (and english) much more closely.
Putting aside the question of whether English's larger training corpus would be a dimension of quality: content is "tokenized" before the model sees it. The "~4 English chars/token" rule isn't a strategy to optimize around because in practice, many English words become one token, the same way Chinese words do.
"I love you" and "我爱你" both tokenize down to 3 tokens. So no efficiency gains from a Chinese translation. "Telephone" is just one token, and "电话" (Chinese word for telephone that spans 2 characters) also gets condensed into one token. Sometimes there are minor gains, like "经济" being one token but "economy" being two. But it doesn't end up materializing in the gain you'd hope for by increasing meaning-per-character.
You can play around with OpenAI's tokenizer here. Great for getting particular about how to save tokens in prompt and tool definitions https://platform.openai.com/tokenizer
What do you mean? A lot of us already use Hanzi (汉字) or Kanji.
Or are you asking why non-Chinese speakers do not prompt LLMs in Mandarin?
Since the information density is about the same in all (spoken) languages, there shouldn't be much of a difference. But maybe one writing system is better suited for LLMs.
i think korean is the best, just learn the alphabet and you can do pretty crazy compaction by dropping honorifics and abbreviation, you can also sound out foreign words ex. 짱깨 if given the right context LLMs should have no problem understanding