There’s a difference between listening to an audio book or proofreading text, for example. In the latter case, the user may want to hear “capital X”, not “ten” in “Act X, scene III”.
For proofreading, blind people often prefer their TTS system to be predictable rather than smart, but often wrong.
For example, Apple’s TTS tended to pronounce the “Read” in “Read me.txt” in the past tense, but that didn’t bother users much. Its imperfect smartness around abbreviations (“Dr. Mulholland” vs “Mulholland Dr.”, “St. Albans” vs “Albans St.”, to mention a few) bothered them more.
(Weirdly, the iOS TTS system seems to handle these worse than MacinTalk did around 1990)