How would this actually work in practice? Do I ask the user to utter specific words then train on that? How is it different from the traditional speech recognition that I need to 'train' to work better on my voice?
The Holy Grail would be to train the model while using it, without any friction. I don't think these methods support that though.