(I did then go and check the book myself; ChatGPT in English was right, the name is there)
(I did then go and check the book myself; ChatGPT in English was right, the name is there)
For example (not actual output):
Input: "こにちは"(konichwa) Qwen Thinking: "Ah, the user has said "こにちは", I should respond in a kind and friendly manner.
Qwen Output: こにちは!
It quiiiickly gets confused in this, much quicker than in English.
That said, in my tiny experience, LLMs all think in their dataset majority language. They don't adhere to prompt languages, one way or another. Chinese models usually think in either English or Chinese, rarely in cursed mix thereof, and never in Japanese or any of their non-native languages.
So all information gleaned reading a glyph in the context of japanese articles would be totally different vectors to the information gleaned from the same glyph in Chinese?
It was pretty good at one shot translations with thinking turned off however, I imagine thinking distracts it from going down the Japanese only vector paths.
Why?
Because the creators want the reasoning trace to be human readable. And without a pressure forcing them to think in English, they tend to get weird with the reasoning trace. Wild language-mixing, devolved grammar, strange language-mixed nonsense words that the LLM itself seemingly understands just fine.