We're using LLMs and good ol'ML. These systems are never going to be 100% accurate. Then again humans are also not 100% correct and we're working hard on ironing out the kinks like the one you just discovered.
I certainly don’t point it out to be discouraging - on the contrary, I feel like I am the exact target audience for a product like this and would happily pay for it if it can reach or be near the trustworthiness of a human teacher.
As you no doubt have considered, but if you define you mission by the tool (AI) and not the outcome (accurate, faster language acquisition) you'll have a great tool and lesser outcomes. Whatever you prioritize, that's what you'll get! :)