My guess is screen readers can have trouble telling what it really says.
It's actually the humans who don't know what it really says. We see 'some weird looking glyphs' and mentally translate them to letters we know. This quirk of the human brain is how we can read a random jumble of letters from different alphabets as English words.
Screenreaders on the other hand know exactly what it says and they read it out precisely how it's written.
I've seen it many times here at least. Just add an addendum!
E.g. 饾枡饾枍饾枎饾枠 get read as "Mathematical bold fraktur small t, mathematical bold fraktur small h, mathematical bold fraktur small I, mathematical bold fraktur small s"
(* I usually hate when 'just' is used this way, I'm sure it's a minefield in some respects. But still more robust than trying to optically interperet what amounts to captchas in some situations :)
iconv('UTF-8', 'ASCII//TRANSLIT', $text)
in my work to make sure I could accept arbitrary Unicode from users, and also send that same info as ASCII to legacy systems that mangled UTF. It worked really well.When I wrote the parent comment, I tried it real quick with a REPL on my own machine, but apparently I don't have the right combination of environment variables or locales or God knows what. Didn't feel like chasing that rabbit into his hole.
It should aboslutely be "smart" enough to simply say "fraktur this" -- to recognize a string of letters of the same style and specify that once, and to interpret the text as a word.
It is the job of screen readers to convert the visual meaning to an auditory meaning as clearly and concisely as possible. If it isn't doing it concisely, then the screenreader is failing and needs some better engineering.
<span aria-label="like this example"><span aria-hidden="true">饾暆饾暁饾暅饾晼 饾暐饾暀饾暁饾暏 饾晼饾暕饾晵饾暈饾暋饾暆饾晼</span></span>
The alternative would be restricting the use of these characters, but this would kill many valid use cases and you can't just forbid non-latin characters on an international service.I've written a post[1] about this and thought I'd implement it in one of my services as an example. I haven't seen any example of it live yet.
Getting those services to include such markup in the copy/paste version would probably solve this problem for most occurences. Note that (a) the original text is known to the service and (b) copy/paste content doesn't have to match what is displayed (e.g. "pastejacking")
https://twitter.com/jypnati0n_/status/1303012163766743041
Seems like the screen reader could have a "say what they mean" mode to pronounce "饾挄饾拤饾拪饾挃" as "this". I assume the legitimate use of Unicode mathematical symbols like "饾挄饾拤饾拪饾挃" is much less than the pseudo-font usage.
I was doing it to work around the limited formatting available here, but I noticed a couple hours later that my comment was transliterated into ASCII. Given the time delay, I assume it was done manually. (Dang, was that you?)
Understandably though, screenreader software probably isn't as fast moving as other industries. If there was an open source screen reader I think I could easily make a pull request to solve this particular problem, if it is a problem.
ASCII English "Latin" letters. Not "the" alphabet or "normal" letters.
If people are going to do strange things like this en masse, perhaps the screen readers could implement some normalisation rules.
I assume he's talking about using unicode characters that visually similar to english letters, but are not actually english letters.
They're chosen because of some visual elements, and are hard, if not impossible to pronounce for text-to-speech.