A lecturer who is fluently multilingual might indeed smoothly switch accents when pronouncing foreign words. But it's still the same voice, and (if they're well practiced at it) they don't have to pause in mid-sentence to switch languages, as text-to-speech systems usually do. And eSpeak can switch languages while still being the same voice, since it's a rule-based, parametric synthesizer. But, at least with NVDA, a mid-sentence HTML span with a different lang attribute still causes a (short) break in the intonation on either side. That's too bad, because a multilingual parametric synthesizer like eSpeak could be like the ultimate polyglot speaker, impressing us all with how smoothly it switches languages.