We've got a few decent speech synths, but information about how things should be read out isn't passed through to them. That's handled by a screen reader program… except screen readers can't represent half the semantics they should, so people regularly bypass them, which leads to (a) UI inconsistency; and (b) the systems being useless if you need something other than a screen reader. AI scraper bots are the straw that broke the camel's back, so virtually no (current) website is accessible via a basic web browser any longer. UI customisability was low in the Windows 95 days, but we've managed to go backwards from there.
We might as well go the whole way, and design something that's actually usable, then put together case-by-case compatibility layers. Here's how we translate Home Office Design System HTML, here's how we translate Stacks Design System HTML, here's how we translate MediaWiki HTML, here's how we translate Wordpress Gutenberg HTML, here's how we translate Moodle HTML… here's how we represent the OpenDocument content model for reading and writing, here's how we represent the SVG content model for reading and writing, here's how we represent a login flow…