Free tier Gemini CLI literally writes Android app for me by just endlessly wondering in English. AGI's here. And it struggles with Japanese. How!?
Free tier Gemini CLI literally writes Android app for me by just endlessly wondering in English. AGI's here. And it struggles with Japanese. How!?
In some fields and languages, this is easier, but as soon as you need the LLM to follow rigorous instructions, it'll fail.
But for e.g. Spanish, if a bare subjunctive verb (making the verb "<x>" something like "could <x>," "should <x>," or "would <x>" without specifying) is used in a sentence, there's no way to know from that sentence who or what that subjunctive is being applied to. The unwritten rule is: 1) if it's obvious who or what you've been talking about, then it applies to that person or thing, 2) if there are two or more things it could be about, then you should figure out a way to add more information to specify which, and finally 3) if there are no good candidates, the person is referring to themselves.
I've heard that there's a lot more of that in Japanese, a language that I don't think really has personal pronouns at all. For example, iirc when you refer to an emotional state and don't specify who you're talking about, it's automatically assumed that you're talking about yourself. I'm not too familiar with Japanese, but languages are simply different from each other, they're not substitution ciphers.
LLMs have trouble staying aware of the context as things go on even for a fairly short time, and in some languages, the meaning of the conversation infects every utterance. Seems a little similar to how LLMs are bad at lifetimes in Rust. I'm sure they'll gradually get better.
It doesn't mention mistranslating, so it's difficult to know the root of the problem is AI "struggling".
> It doesn't follow our translation guidelines. > It doesn't respect current localization for Japanese users, so they were lost.
I believe this is the root of the problem. There are define processes and guidelines, and LLM isn't following it. Whether these guidelines were prompted or not is unclear but regardless it should've been verified by the community leaders before it's GA'ed
That's very much in the past now. But it'll linger in the training data for a while.