I see two types of content. Published content and User generated content.
Published Content is any content produced by employees who run the website or customers who are hired by the employees to work for them. Here content is controlled and there is some sort of oversight as content gets published and there is responsibility on publishers. Generally there are some monetary incentives to the publishers to do that.
User generated content is any thing posted by any user. There are no monetary benefits to the users themselves. Not much oversight. Though there might be community guidelines, they are often not enforced. Consequences of not adhering to the policies might not matter much to the user.
From accessibility standpoint, there is not much we can do for the user generated content other than educate them. In the end, it's user's decision whether to adhere or not.
In case of published content, website owners have total control on what goes out to begin with. They can proactively enforce accessibility guidelines.
For the comment in discussion, my observation is that glyphs are popular for user generated content rather than published content. And from accessibility standpoint, it's an OK trade off.
But as you said, for an industry that systematically ignored the problem for 10yrs, I would rather push them for generation of better accessible published content than user content and hence i said itβs a trade off.
I also wonder how ML can potentially help here. For example, for images without alts, automatically populate them.
I assume he's talking about using unicode characters that visually similar to english letters, but are not actually english letters.
If people are going to do strange things like this en masse, perhaps the screen readers could implement some normalisation rules.
My guess is screen readers can have trouble telling what it really says.
It's actually the humans who don't know what it really says. We see 'some weird looking glyphs' and mentally translate them to letters we know. This quirk of the human brain is how we can read a random jumble of letters from different alphabets as English words.
Screenreaders on the other hand know exactly what it says and they read it out precisely how it's written.
I've seen it many times here at least. Just add an addendum!
E.g. ππππ get read as "Mathematical bold fraktur small t, mathematical bold fraktur small h, mathematical bold fraktur small I, mathematical bold fraktur small s"
(* I usually hate when 'just' is used this way, I'm sure it's a minefield in some respects. But still more robust than trying to optically interperet what amounts to captchas in some situations :)
iconv('UTF-8', 'ASCII//TRANSLIT', $text)
in my work to make sure I could accept arbitrary Unicode from users, and also send that same info as ASCII to legacy systems that mangled UTF. It worked really well.When I wrote the parent comment, I tried it real quick with a REPL on my own machine, but apparently I don't have the right combination of environment variables or locales or God knows what. Didn't feel like chasing that rabbit into his hole.
It should aboslutely be "smart" enough to simply say "fraktur this" -- to recognize a string of letters of the same style and specify that once, and to interpret the text as a word.
It is the job of screen readers to convert the visual meaning to an auditory meaning as clearly and concisely as possible. If it isn't doing it concisely, then the screenreader is failing and needs some better engineering.
<span aria-label="like this example"><span aria-hidden="true">ππππ π₯πππ€ ππ©πππ‘ππ</span></span>
The alternative would be restricting the use of these characters, but this would kill many valid use cases and you can't just forbid non-latin characters on an international service.I've written a post[1] about this and thought I'd implement it in one of my services as an example. I haven't seen any example of it live yet.
Getting those services to include such markup in the copy/paste version would probably solve this problem for most occurences. Note that (a) the original text is known to the service and (b) copy/paste content doesn't have to match what is displayed (e.g. "pastejacking")
https://twitter.com/jypnati0n_/status/1303012163766743041
Seems like the screen reader could have a "say what they mean" mode to pronounce "ππππ" as "this". I assume the legitimate use of Unicode mathematical symbols like "ππππ" is much less than the pseudo-font usage.
I was doing it to work around the limited formatting available here, but I noticed a couple hours later that my comment was transliterated into ASCII. Given the time delay, I assume it was done manually. (Dang, was that you?)
Understandably though, screenreader software probably isn't as fast moving as other industries. If there was an open source screen reader I think I could easily make a pull request to solve this particular problem, if it is a problem.
ASCII English "Latin" letters. Not "the" alphabet or "normal" letters.
They're chosen because of some visual elements, and are hard, if not impossible to pronounce for text-to-speech.
Instead of "clickable clapping hand sign quick graphic emoji" or whatever ridiculouslness it's saying...
...shouldn't the screen reader just say "clap", perhaps in a different pitch or intonation that one signifies "emoji"?
The tweet can be perfectly read by a human being as "If clap y'all clap don't clap" etc... with the "clap" in a higher pitch or more emphasized intonation.
The emoji are fine. Screen readers just need to improve how they read emoji. Which doesn't seem particularly hard to engineer.
There's also the problem of Emoji misuse, for example, using , pronounced as "globe showing americas", as the emoji for Earth.