In one instance, I was having it correct akkusativ/dativ/nominativ sentences and it would say the sentence is in one case when I knew it was in another case. I'd ask ChatGPT if it was sure, and then it would change its answer. If pressed further, it would again change its answer.
I was originally quite excited about using an LLM for my language practice, but now I'm pretty cautious with it.
It is also why I'm very skeptical of AI-based language learning apps, especially if the creator is not a native speaker.
I was asking only for the meanings of the words and phrases, though. I didn’t ask for things like pronunciations, grammatical categories, etc. In the past, when I’ve tried to get that kind of granular information from LLMs, there were indeed errors, presumably because of tokenization issues.
A few days ago, I ran some similar tests with Japanese, asking for readings of kanji and jukugo in an extended text. All of the models I had tried before for such tasks had screwed up. This time, however, ChatGPT o1 scored 100%. It also was able to analyze sentence grammar accurately, unlike the other models I tried. I was impressed.
At current API prices, though, o1 might be a bit too expensive for such a task.
1. Role based "agents" with a router and logs (for auditing reasoning and decision making).
2. Cross validation and redundancy with the translation "agent" using a 2nd language (that is not English) that you are also native in to check if the translation carries the same "meaning" (sentiment) and cultural significance (Turkish is especially rich in symbolism and cultural memes).
YMMV: I am a car salesman irl and have no formal training.