I'm a little surprised to have not seen any mention of the Speech Synthesis API (https://caniuse.com/#feat=speech-synthesis).
It's supported almost everywhere that the Web Audio API is.
It's supported almost everywhere that the Web Audio API is.
[1] https://w3c.github.io/speech-api/#dom-speechsynthesis-getvoi...
(However, one of the scenarios for a version 2.0 was to implement the same API, additionally to the 'native' one, to be used as a fallback solution. While I actually had implemented this already, I don't think it may be that useful, while it increases file size quite a bit.)
Edit: Viable points for this may be still a) reliable performance and interaction, and b) known voices (even, if they are a bit robotic), c) use in offline applications. Using an analyser node for animations may be yet another.