Does Speaking to Agents Like Cavemen Save 65% of Tokens? We Test
blog.jetbrains.com
blog.jetbrains.com
alt.nerd.obsessive FAQ v1.4
In the episode where they were filming the Radioactive Man movie [1995], the comic book store guy tells Bart that he can find out the star of the RM movie. He promptly posts a message to alt.nerd.obsessive which states "Need know star RM pic". This information is relayed through the nerd world until it reaches a nerd hiding under the table at a meeting of movie moguls casting the film. He relays back the answer immediately (Rainier Wolfcastle).
https://en.wikipedia.org/wiki/Radioactive_Man_(The_Simpsons_...Also, this really should be more precise that it’s talking about neo-caveman. Legit caveman no speak English.
i think korean is the best, just learn the alphabet and you can do pretty crazy compaction by dropping honorifics and abbreviation, you can also sound out foreign words ex. 짱깨 if given the right context LLMs should have no problem understanding
What do you mean? A lot of us already use Hanzi (汉字) or Kanji.
Or are you asking why non-Chinese speakers do not prompt LLMs in Mandarin?
你不去,我也不去
You no go, I also no go
今天这里人很多
Today here person very many
下雨就不去
Fall rain then no go
Vulgar latin does correspond to modern romance languages (and english) much more closely.
Putting aside the question of whether English's larger training corpus would be a dimension of quality: content is "tokenized" before the model sees it. The "~4 English chars/token" rule isn't a strategy to optimize around because in practice, many English words become one token, the same way Chinese words do.
"I love you" and "我爱你" both tokenize down to 3 tokens. So no efficiency gains from a Chinese translation. "Telephone" is just one token, and "电话" (Chinese word for telephone that spans 2 characters) also gets condensed into one token. Sometimes there are minor gains, like "经济" being one token but "economy" being two. But it doesn't end up materializing in the gain you'd hope for by increasing meaning-per-character.
You can play around with OpenAI's tokenizer here. Great for getting particular about how to save tokens in prompt and tool definitions https://platform.openai.com/tokenizer
Since the information density is about the same in all (spoken) languages, there shouldn't be much of a difference. But maybe one writing system is better suited for LLMs.
Also it was tested on reasoning low, i'd have liked to have them tested on higher reasoning levels.
Grammar sounds off? Was the message correctly understood; is there enough left to apply error correction and regenerate the intended message?
Specifically abjads since as far as I am aware their difference from alphabetical and abugida scripts is that they infer the vowels instead of explicitly marking them.
What are your thoughts on it?
When Sabrewulf[1] come, some fight, some run ... some run in cave. No good. Bad.
Token cost. Seem smart less token. But already discount. No profit.
Use less token? Price go up.
What money?
You should be bothered by the persistent & popular use of the "caveman" terminology and the myth surrounding it. However, from a dedicated amateur linguistic point of view, the "caveman language" that so many people seem to know by default is kind of remarkable. Perhaps it points to (it doesn't) some deeper, ancient understanding.
[0] this is a joke because the caveman myth comes with a lot of misogynistic baggage and you should have a problem with it. Also because I made you read this. (this is not an explanation of the joke itself, that's up to you)
[1] the sabre-tooth cat is a trope that over-represents its relationship with people. I mean it's extinct, but it was around "back then", right?