Eleven Labs is still very far above open source models in quality. But StyleTTS2 (MIT license) is impressively good and quite fast. It'll be interesting to see where this new one ends up. The code-switching ability is quite interesting. Most open source TTS models are strictly one language per sentence, often one language per voice.
In general though, TTS as an isolated system is mostly a dead end IMO. The future is in multimodal end-to-end audio-to-audio (or anything-to-audio) models, as demonstrated by OpenAI with GPT-4o's voice mode (though I've been saying this since long before their demo: https://news.ycombinator.com/item?id=38339222). Text is very useful as training data but as a way to represent other modalities like audio or image data it is far too lossy.