I was thinking about a piece of this problem earlier and I'm curious if you figured out element that I didn't:
How do you provide feedback on how someone mispronounces a phrase?
How do you provide feedback on how someone mispronounces a phrase?
Again, I'm curious if someone solved it. Theoretically this is within modern video/audio AI, but that doesn't mean we have the models or even organized data.