Vocalizing internally is your muscles doing what they would do when you speak, but moving ever so slightly or not at all.
I guess the device was first build to recognize actually spoken speech with the signal that is transmitted to the muscles. They gradually reduced the speech levels, and now the system is efficient even with tiny signals.
Producing speech is an incredibly complex task; taking most humans more than 2 years to acquire. Loads of different muscles are involved. I would not be surprised if the different signals they collect in their device are more than redundant for the task.