In my city I can recognize that google has two different 'voices' or voice libraries. They sound slightly different. I'm curious how that works and why it's not all done with one.
I've noticed this as well. My working hypothesis is that one is for high(er) bandwidth and the other for low bandwidth situations.
My understanding is that the "low-fi", more robotic one uses an offline TTS engine for when there is no connectivity. When connectivity is good, it will switch to better, cloud-based one.