What do you use for validating pronunciation?
Or do you just use a Speech To Text model and hope the text that comes of it is the expected phrase?
Or do you just use a Speech To Text model and hope the text that comes of it is the expected phrase?
No comments yet.