If I write "ceci n'est pas du français" here for instance it'll remain untagged.
I think the only websites where you can reliably expect the text to be tagged correctly are dictionaries and the like, where it can be done unambiguously and at scale (wiktionary seems to do it for instance: https://en.wiktionary.org/wiki/coin ). In these case it seems like it could genuinely be useful.
I've also looked at a few language learning apps (such as Duolingo) and none of them appear to use the lang attribute at all, even though it seems like it would be a perfect use case for it.
I guess my point is that I'm not sure it's fair to blame a format that strives to be minimalist for having such a shortcoming for lacking such a very specific feature.
Of course if you're a blind linguist you might have a very different opinion on the subject...
Imagine saying accessability hardware - like corrective lenses - is niche.
Then "accessiblity" in general is not niche, it effects everyone or almost everyone.
But "accessibilty" is a broad concept, applying to all sorts of different unrelated needs. Vision impairment, cognitive differences, mobility needs, etc.
Some of them effect more people than others. Some of them will be "niche", effecting few people. it is just a fact, right? It is faulty logic to say that because "accessibility" effects everyone, every single accessibility accomodation is also therefore of use to everyone. It's not so.
How much need is there for multi-lingual screen-reading? I have no idea. It's obviously not universal -- I don't need it for instance, although I may need other kinds of accessibility attention -- but there may be lots of need for it! But we can't prove it one way or another just from language games around the general concept of "accessiblity".
Whether developers ought to accomodate even niche accessibility needs is a separate ethical or practical argument.
it would be nice to have more accesible interfaces for everyone. helping one person doesn't mean you fail to help everyone, in fact, helping one person in terms of accessibility generally has the effect of helping everyone.
we're asking for some access. not all the access. scoped properly, it doesn't seem to be that insurmountable of a development goal, honestly.
How does facilitating multilingual screen-reading help someone who is neither multilingual nor uses a screen-reader? How can it possibly help everyone? (like literally 100% of people, you are suggesting?) What am I missing?
- better support for users who are not multi-lingual (a client could try to auto-translate sections that are not in the user's native language)
- better search capabilities (find me all documents that contain a French passage)
- better voice-assistant support (if a document is being read out loud for any reason, it would be nice to be able to handle multiple languages well. This also combines with the auto-translate capabilities above.)
Basically, all of the reasons why Gemini includes a "lang" attribute in the first place, except recognizing that parsers and clients want to be able to make use of that on a section-by-section basis rather than purely at the document level.
Sometimes documents are big and expand across multiple languages. It's just a bad abstraction in general to assume that each document would only have one language type[0]. Documents/books in the real world don't work that way, and a document format that can't accurately represent a giant portion of classic literature isn't a very good general-purpose format.
[0]: yes, technically Gemini allows you to specify multiple top-level languages, but that's not useful for anything. If I'm parsing a document, I don't want it to tell me "hey, there's some French in here somewhere, good luck!" Tell me where it is.
A text only web is pretty niche. Can this project really afford to start slicing off additional portions of its userbase? I agree that multilingual blind users are in the minority, but those users are probably the most likely of anyone to be interested in a text-only web.
And even ignoring blind users, it's just... what's the point of any of this if Gemini is not going to be semantic? It's like religiously adhering to the GNU principles of always printing human-readable text on the command line, and then not shipping a shell with a pipe command or any text parsing capabilities. From just about any perspective, whether you're worried about accessibility, or parsing, or search -- it's useful to know what language a document/section is in.
One of the few advantages of having a limited, well-defined, non-extensible spec is that it's easy to work with and write parsers around. And immediately Gemini is just saying, "filtering sections by language? Good text-to-speech support? No need for that."
If nothing else, Gemini is theoretically built specifically for generality[0], and it's already guaranteeing out of the box that it can't be used well with voice assistants. Those things aren't a fad, a general-purpose document format kind of needs to support them. And trying to have a voice assistant that intuits the current language is just a really bad system for everyone.
I agree with you that those behind Gemini should take this issue seriously, but maybe something like what I've mentioned is what they have in mind as their answer. Not sure.
That's one reason inline links are not allowed (they have to be a separate line): it would complicate the parser.
It shouldn't be too hard to add a "language switch" line, so that sections (but not words within sections) can be different languages.
Making Wiktionary or Etymonline in Gemini would require you to add a line break for each foreign word, which seems annoying to read.
I don't see though, why the spec can't allow links that are written on a newline for markup simplicity but displayed inline? And then the same could be done for foreign words.
A better example would be things like wheelchair accessibility for instance, which is indeed very much not implemented everywhere.
To be clear I'm not saying that it's a bad idea to implement accessibility features, It's actually quite commendable. I just feel like in practice it's a lot of additional work to get it right and even then it may not do much good at all. Again, while the "lang" attribute is technically supported by HTML, I struggle to find places where it's used in the wild, even in places where it doesn't seem like it would be technically difficult to do it (like dictionaries and language learning websites).
If I was a gemini developer I'd definitely listen when people complain about bad accessibility, but I would also take the time to understand how people deal with things like multilingual content in practice and how to make it easier, because clearly at the moment "lang=" is not it.
That's my point :)
> Also it's a pretty bad example because very rarely do you need to worry about "corrective lense accessibility", except maybe for things like VR helmets.
Ahh, but what corrective lesnses points to is degregation in eyesight. So ways to make viewing/experiencing content _without_ glasses is definitely a concern to between 50%-70% of adults.
I've always loved that the Kindle, even with such a simple use case as reading a book, supports so many ways to see it and interact with it, so even people without vision or motor problems can use it the way they like best.
But at Apple we were well aware that accessibility features like closed captions were often used by people with no disabilities. We would add stuff knowing that all sorts of people would find uses for it.
I can easily imagine a language tag would also be useful for filtering out a German-language gemini for instance.
This was a double edged sword however. Like with closed captions/subtitles, we knew that many people using them that had hearing issues were older and also likely had out of date glasses prescriptions.
But this didn't stop the AppleTV designers from lowering the contrast of the font by displaying it over a transparent background, lowering the font size, using Helvetica instead of an accessible font, and removing the positional text feature that indicated who was speaking. And in typical Apple style, there was no option to make the text more legible.
All of these choices were made to appease people without disability who were using the feature for whatever reason. I didn't like it.
So I guess what I'm saying is, accessibility features are good for all sorts of reasons but don't lose sight of the much, much smaller group of users who require the feature in order to use the product at all.
I think it is important to be honest and upfront when we talk about accessability in software because I've seen so many times when developers and product owners dismiss it entirely. I really want to stress the point that thinking about accessebility as a niche is exclusionary and unethical.
I think about my time working at Net a Porter - UK upmarket online fashion retailer - where I brought up some accessability concerns especially around their terrible in-house captcha and it was dismissed as "they arent our target market" which is so blanetly untrue, but all stems from this myth that accessability is only for a small group of users.
> don't lose sight of the much, much smaller group of users who require the feature in order to use the product at all.
The people who need to use captions/subtitles are not a "much smaller group". I think about foreign language media, where i need the subtitles (like when I watched Dark on Netflix), or even primarily English media but has some forign language (I watched Lost recently - there's a fair bit of Korean and Arabic in that). Or even when you're watching TV around someone who's sleeping.
Everyone needs "accessability", and if you don't now, you will eventually.
This is uncharitable. They were specifically talking about multilingual accessibility. You responded to something they didn't write.
Yes, and a minority of those users are visually impaired, hence it is a niche. Maybe you considered "niche" to mean "preference" or "novelty"?
" niche, adjective
denoting products, services, or interests that appeal to a small, specialized section of the population. "
The problem with this you'll get a million other people saying "Gemini is a really good simple format that the internet needs, but you need to just add feature X". But if Gemini added all those features, it would become as bloated as HTML/CSS/JS!
Yes, and the other million people are going to say their proposed features are "simply doing the right thing".
It may be that marking up what language a short except of text is in is a good feature to have (and would be fairly easy to achieve if Gemini used the Markdown format as you could simply us a <span> tag). A counterargument would be that traditional text doesn't do this, so a text file format ought not to unless it is aiming at being something more than what text does.
But add too many features, the project becomes something else, is maybe less useful & doesn't gain traction. A Gemini project that is too bloated and thus gains no traction helps no-one, disabled or otherwise.
Gemini does aim at being more than text[0]:
> The "first class" application of Gemini is human consumption of predominantly written material - to facilitate something like gopherspace, or like "reasonable webspace" (e.g. something which is comfortably usable in Lynx or Dillo). But, just like HTTP can be, and is, used for much, much more than serving HTML, Gemini should be able to be used for as many other purposes as possible without compromising the simplicity and privacy criteria above. This means taking into account possible applications built around non-text files and non-human clients.
Lots of documents and books contain passages in other languages, it's reasonably common -- I don't get why people are saying this is a niche concern. Books get around that problem because they have controls around presentation, and they're not designed to be consumed visually, not as a semantic, parseable format. But Gemini, at least in theory, is designed to be parseable.
I.e. it's not enough for something to be possibly usable by a large portion of people - the userbase have to actually be interested in utilizing it and consider it more important than other things instead. I strongly think if you follow this logic you'll find it's a very niche group that cares this deeply about the semantic language tags. Maybe you truly strongly don't and consider it something everyone would love if only it was there.
First, I do think this kind of thing falls into a category of... maybe 'polish', for lack of a better word?
This is something that you're not going to notice you're missing until you think "I want to write one of these quick and dirty display clients that I've been encouraged to build over the weekend, and wow, embedded Japanese is just literally impossible to render well." If you have a document format that's designed to be adaptable and flexible, predicting what you're going to need is difficult.
So maybe the idea is that you're not going to throw transcriptions for comics/manga into this, and nobody would need to search documents for embedded passages in other languages, and stuff like auto-translations in a client wouldn't be useful. But to me it kind of just smacks of, "this is designed for the use cases we could think of right now." I don't know what features will be critically important until I try to do something creative, that's why I value a tightly constrained format that's simple to understand but still offers at least some flexibility.
The other objection that keeps coming up in my mind is: is Gemini actually swamped for feature requests right now? Are they being forced to prioritize features? I'll definitely concede that if you surveyed 10,000 people using Gemini, stuff like a hard limit of 2 subheaders is very likely annoying more people right now than anything to do with language types. But Gemini also doesn't look like it's trying to fix its subheader limit. The impression I get from the FAQ is that this is not a spec that's actively evolving much right now.
And that kind of loops me back around to thinking about polish again. On one hand, I'm inclined to think you're right, this is probably not the first thing on anyone's mind. But on the other hand, "a document with multiple languages in it" is a reasonably common thing that will show up in multiple settings. So I'm still looking at a text format that even just on the surface level doesn't seem to be very good at describing text, even in cases where the solutions seem pretty simple. Having a language switch tag in the spec or just allowing multiple top-level media types in the document doesn't seem like it would complicate the parser spec or make it any harder for me to build a client.
Part of this is, if I'm going to adopt a format, I want it to be well thought out -- I want the authors to have spent more time thinking about it than I have. So if I can immediately see problems before I even start using the spec, it makes me wonder what else I'm going to run into where I'll just be scratching my head over why it was designed that way.
On some level I get where you're coming from, and rendering multiple languages is not the end of the world. You can even potentially work around it, maybe. But it would also be very, very easy to get this right, and I don't understand why the first thought from a spec designer putting together a content type header wouldn't be "what if the content type changes mid-document". This is maybe unfair, but the immediate thing that it makes me think is, "these are people who have not spent much time considering how text documents work, even for simple cases."
But if Gemini was in heavy active development, my opinion would probably be different. I don't know, maybe it is? Are they actively evolving the spec right now? Maybe I'm just being over-critical. But I feel like I shouldn't be able to immediately think of use cases involving basic text that Gemini just can't handle, it should take me longer to see the limitations of a general-purpose text format.
Yes they do, and since Gemini allows Unicode, it can cope with this. Note that books typically don't mark what language foreign words/phrases are in.
> But Gemini, at least in theory, is designed to be parseable.
If I was designing Gemini I'd design the source format as something very like Markdown and have it compile to a format which would be a subset of HTML. This would allow <span lang=...> constructs.
Other text-centric media often emphasize foreign words in some fashion. For instance, in an English language novel foreign language text is often italicized. Same in many articles.
Ooof. This kind of wording makes me flinch away.
There are differences in the display (font) of the same Unicode character.
https://en.wikipedia.org/wiki/Han_unification#Examples_of_la...
I find Gemini pretty accessible overall. There are no images, so including alt descriptions isn't even a concern. There's no CSS, so you can't make controls that visually look like check boxes, but are much less accessible. In fact, it's way harder to make an inaccessible Gemini site than to make an inaccessible website. The only gripe I have is the possibility to introduce ASCII diagrams, which usually aren't accessible.
The provides an alt text tag to preformatted text blocks for this purpose:
> Any text following the leading "```" of a preformat toggle line which toggles preformatted mode on MAY be interpreted by the client as "alt text" pertaining to the preformatted text lines which follow the toggle line. Use of alt text is at the client's discretion, and simple clients may ignore it. Alt text is recommended for ASCII art or similar non-textual content which, for example, cannot be meaningfully understood when rendered through a screen reader or usefully indexed by a search engine.
CAD software exports .stl files which can be processed into instructions for machines. That is, there is already software that can take lines on a screen and convert it to instructions for a machine. There's probably a way to annotate ASCII images of drawings (schematics, architectural) and output text "text* drawing of an object 14 centimeters by 5 centimeters, ..." Stuff like vinyl cutters (and even the proprietary cricut) can take vector graphics and convert them to instructions for a machine. so it, in theory, might be possible to just paste the ascii drawing text into some software: render ascii as .svg/vector, convert to 2d stl ("a layer"), convert stl to descriptive text. Since the software renders, you don't need OCR for the annotations, which are mostly instructions.
if there's no annotation, assume the '-', '_', '|', etc are to scale, and just use "units" as the dimensions instead of centimeters or whatever. |----| = 'two verticals separated with a 4 unit line'
Am i crazy?
Obviously humans are a bit better at contextual clues, but.. the information has to be there for a human user to make the distinction.
If anybody ever opens an S&M themed bakery, the problem will be accentuated.
Now it won't be a surprise.
It might be better then to invest effort in improving language inference heuristics, which would help in all applications (pdf documents, text files) than to try to build in support in each underlying protocol.
It may be that the language which a quoted foreign-language word is meant to read in was mentioned far back in the text. Human sighted readers will know how to read it, but computers won't yet (and perhaps won't for many years if not decades). However, there are visually impaired people now who need to be guaranteed the same Gemini experience as sighted users.
This means that if the screenreader picks the wrong language, then the user has to guess the spelling from the way it's pronounced to understand the text.
Though I'll concede that right now on the web it isn't much better; but that's a failing of publishing tools, not the format.
Indeed, so why not let the human writer clarify it?
Why do you prefer a resource intensive guessing algorithm that will never have a 100% accuracy over just having an annotation that takes less than 10 bytes and is trivial to parse?
In my opinion there's no reason to even consider the first option, especially with gemini's focus on simplicity in mind. This "what can another little JS library hurt" attitude is what lead to gemini in the first place.
Add language inference heuristics to cover all kinds of formats and now you've got a dependency. Then some people need different settings so you add config files and parsers/dependencies for those. At some point you update the model and now it can't properly differentiate Mandarin, Cantonese and short Japanese phrases anymore. You start adding cross-platform support and other features, and now some people say it's too slow on their Raspberry Pi Zero setup. Bloggers complain as they have to rearrange some quotes via trial-and-error so the heuristics pick up the correct language. Unfortunately that makes it worse for people running the older version 0.7.
After this rant it should be obvious, but: I prefer just adding a ["en"] or similar and stop worrying about it.
Not always, but often enough that it's an informal standard.
On the other hand, there will also be a bunch of cases where people don't do the work to mark this up (maybe not even realising they need to)...
I do agree though that this is a far bigger problem space than HTML's `lang` attribute is capable of solving. It's not like knowing the language of the text always tells you the unique pronunciation of a string of characters.
Reading, as a discipline, is hard at the edges, and lots of humans don’t get it right either!
<sentence lang="en-US-x-Pittsburgh" intonation="falling"><word pos="subject-pronoun" ipa="hi">He</word> <word pos="verb-simple-past" ipa="ɹɛd">read</word> ... </sentence>Okay, now try to read this comment with a screen reader accurately. Someone who knows chinese, japanese, and english will be able to pronounced both of those words correctly. There's no heuristic that can do the same.
I should be able to annotate each of those characters with whether it's meant to be chinese or japanase.
I often see it on elements/sections of sites that link to content in another language.
Here, German is the "main" language and lots of documents are available in German but not English. If you need to link to a German page/doc from an English one, it's not uncommon for the link text to be in German too.