The study didn't use special characters and also used the same laptop keyboard, which I would imagine significantly inflates the percentage compared to a real world situation for this to be effective.
The particular method suggested is about using correlatable known text with audio situations to build the accuracy for other recordings of the same typist. I.e. a meeting with a chat feature or an office app may have correlated exanples.