<i>: The Idiomatic Text element
developer.mozilla.org
developer.mozilla.org
- <i> makes text italic
- no, presentational elements are bad, use <em> instead
- okay, fine, you’re all still using “i” so we’ll put it back in the standard
- here’s a convoluted meaning to pretend that it’s in the standard for some reason other than “we tried to remove it and failed”
<i> as italic was great like Tailwind is great: sometimes I just want to design something basic without having to context switch to CSS. I want italic, I use the <i> tag.
What has semantic HTML ever done for us? Does the browser care that I used <i> instead of <em> for italic? Do screen reader even care?
You can contact me directly using my handle at my domain, which is 21337.tech
We need to treat the screen like a PNG which an AI can parse to determine what the text is, what text is relevant to read to the user, where and what the buttons are, how to emphasize reading the text based on the visuals, etc. It's the only thing that will scale and while difficult it seems that creating something like this is within reach.
This actually gives me a little hope that AI-interpreted UI could trigger a trend back to visually obvious and unambiguous controls.
All operating systems and platforms have pretty fluent and in-depth APIs for exposing an “accessibility tree” for assistive technologies to use. It’s your responsibility as a developer to build interfaces for actual users, and there’s very little reason not to.
That's actually a really good idea.
It's 2059. We're using browsers that render pages created with Dreamweaver and Frontpage using advanced machine learning techniques to determine that <blockquote><font size="+1"> is an entry in the table of contents, while <p><font size=+2> is a level 1 header.
I also used to believe that rigorous semantic tagging was The Way To Do Documents, and it was certainly a useful crutch, but we absolutely need to move beyond that.
Prefab[1] and especially Bubble Cursor [2] look like be super useful additions for the flat-UI era.
Here's an explanation of the accessibility features I've added to my canvas library[1]. These features may not be the best approach, or even the most appropriate approach ... but at least they're there, and they can evolve into something better as/when people offer feedback.
For us small things, like enable reader views. For people with disabilities, it lets them effectively use screen readers.
Also, from a purely selfish aspect, I vastly prefer documents with at least a minimum of semantic organization, even if it's just header, footer, main, and meaningful heading levels.
Without those, well, it's just a document full of undifferentiated text with various attributes to distinguish it visually or typographically, and usually inconsistently used. Is this piece of text styled "bold with fontsize 16" supposed to be a second or 3rd level header?
<em> for things that are emphasized, as in emphatic expressions within text. For example "You'll get nothing and you'll like it."
> It's worth noting that the W3C specification says that a reference to a creative work, as included within a <cite> element, may include the name of the work's author. However, the WHATWG specification for <cite> says the opposite: that a person's name must never be included, under any circumstances.
(The heuristics only apply now, as there are two kinds of emphasis, which introduces some uncertainty of how things should be expressed. Moreover, as type styles have moved to CSS, there's a chance that a screen reader presentation may miss an intended separation of text.)
Edit: There's a clear meaning to the use of distinctive type styles. I don't see an intrinsic value in pretending and that there is no such intended meaning and that there should be an artificial separation, rather than generalizing on presentation styles.
Edit #2: In Antiqua, italics isn't just oblique text, it's an entirely different script. So what is the use case and the intended meaning of using a different script? Isn't this more conceptual than just using a fancy visual presentation style, which may be happily ignored in any other representation as there's probably no intended meaning to this?
Fun fact: the original HTML definition [1] has "typewriter" as the only styling element and uses surrounding underscores ("_are_") for emphasis.
[1] http://info.cern.ch/hypertext/WWW//MarkUp/Connolly/complete....
But then HTML became the rendering layer for a universal sandboxed VM, and all that went away.
Why, it has liberated us from European colonialism. <i> being no longer italic, you can wrap it around any script from the world and enjoy the semantics of it.
The phrase "What has semantic HTML ever done for us?" is a riff on the quote "What have the Romans ever done for us?" from the movie Life of Brian, about a Jewish-Roman man mistaken for another Messiah.
The Romans are, of course, Italic peoples.
Italic type is derived from a form of semi-cursive writing, and is entrenched in the European printing tradition. No punning with Italy intended.
So what does it mean if you have <i> around Chinese, or Devangari or what have you? Do you just apply shear to make the text slanted?
I suspect that it bothered some "woke" types that HTML contains Euro-centric typographical directives, so they have been repurposed to have some sort of, culturally neutral semantics that is more inclusive of the planet's diversity (pardon me if I'm not nailing the terminology here).
In short, <i> no longer belongs to whitey and his writing system.
There's actually a tag for this. I can't remember what it is, but it puts dots above the ideograms, which serves the same function.
https://en.wikipedia.org/wiki/Ruby_character#HTML_markup
With CSS `text-emphasis` we can do the dots you are speaking of:
https://css-tricks.com/almanac/properties/t/text-emphasis/
Probably there are other useful bits of HTML/CSS for Asian languages. (right to left and/or vertical writing, anyone? I always wanted to fart around with pretty-printed Chinese poems and, like, calligraphy script fonts with procedurally generated jitter.)
Edit: there are also Unicode entities for these:
Actually, yes! <i> is still defined in the standard to use the italic font style: https://html.spec.whatwg.org/#phrasing-content-3
And if there isn't an italic version of the font (often the case for Chinese etc), https://w3c.github.io/csswg-drafts/css-fonts/#font-style-pro... says the browser should programatically shear it:
> If no italic or oblique face is available, oblique faces may be synthesized by rendering non-obliqued faces with an artificial obliquing operation.
Just use the <em> tag.
The OP addresses "idiom in another language" as one case, but if it is within a scholarly quotation, can one change the typography to a different convention?
I suppose it’s easier to reason about, but it feels like it was invented by that grug brain dude who likes solving lots of problems but hates thinking.
If look at Tailwind's website and examples, you'll see that it's one of the very few CSS frameworks that properly uses semantic HTML everywhere.
And there's no such thing as "semantic CSS".
Whereas with Tailwind or HTML < 5 you name you elements based on what they look like. ".font-bold" or "<b>"
> Whereas with Tailwind or HTML < 5 you name you elements based on what they look like. ".font-bold" or "<b>"
No, you don't. With Tailwind you use the semantic elements in HTML. And you style them using generic primitives
It works quite nicely.
I think the parent meant “[semantic HTML] and CSS”, not “semantic [HTML and CSS]”.
H2 or H3? A purely semantic distinction with Tailwind.
Use an OL for a custom list of items? You start with a clean slate, no overrides required.
And you are discouraged from writing custom CSS that causes unexpected styling as well.
Every HTML tag with an empty class attribute is a clean slate. You only have to care about the default display mode and about allowed descendants.
It’s not an MDN interpretation; this has been part of the HTML5 spec since day 1.
I linked to a 12 year-old article earlier in this thread [1] describing this very thing when HTML5 was new.
The i element represents a span of text in an alternate voice or mood, or
otherwise offset from the normal prose in a manner indicating a different
quality of text, such as a taxonomic designation, a technical term, an
idiomatic phrase from another language, transliteration, a thought, or a ship
name in Western texts.
[1]: https://html.spec.whatwg.org/multipage/text-level-semantics....I think the word you are looking for is “contrived”
Going from HTML/XHTML to HTML5/XHTML5, several elements were redefined, including <i> and <b>.
It’s nothing new.
Here’s an article from HTML5 Doctor explaining this more than 12 years ago [1]. I hope web developers serious about their craft aren’t just now discovering this.
If you read the spec, it mentions several different uses for <i>, depending on the context. Otherwise, we’d need a different element for each context, bloating the spec by having special purpose elements instead of general ones.
For example, <i> is used for ship names:
<p>They came over on the <i>Mayflower</i>.</p>
It wouldn’t make a lot of sense to have an element just for names of ships in a general purpose markup language like HTML.Turns out, if you look at the Wikipedia article about the Mayflower, the name is marked-up with the <i> element [1].
In general I think many UXs are too shy about using italics.
“Are you sure you want to delete the user you piece of shit?”
Options:
> Yes, you worthless dildosaurus
> No! Stop! I hate you!
The practical thing to do was to just re-define <i> and <b> to mean what <em> and <strong> were defined as then add additional elements for when more fine-grained meaning if needed. Instead, countless of hours had to be lost on s/<b>/<strong>/g instead.
This is exactly the sort of "argument from purity" that lacks any practicality that I meant.
1. Add an "i" element to SVG for isometric path drawings
2. Update the HTML parser algo to special case this "i" for inline SVGs
Instead, they should have introduced a new tag, e.g. <idiomatic>, while deprecating <i>.
Most internet standards are descriptive, not coercive. The semantic web's advocates have been coercive from the beginning, but despite the repeated failures of coercion, they keep on at it.
What does "idiomatic" mean anyway? If I say that idiomatic markup is a waste of time, do I have to put "waste of time" in an <idiomatic> element? How is text in "another language" idiomatic text? [Checks for another example] Oh, that's the only example they give of something that's "idiomatic".
It seems clear to me that "idiomatic" isn't what they mean at all; what they mean is some text fragment that should appear differently, because in some way the "mode of speech" is to be treated differently. It looks as if they've hijacked the <i> tag because they want it to stop meaning what it means; they've chosen a meaning that starts with the letter "i" for that reason.
The way I parse it, "idiomatic" text is any kind of text that would normally be set apart by being presented in italic. That is, it's not even semantic markup at all; it's a politically-correct gesture to the semantic markup die-hards.
But now I'm unsure if I should consider these distinctions <bullshit> or <horsecrap>?
Trying to force writers to instead overlay a layer of meta-meaning on their prose, so that the meaning is expressed both in the text and the markup, is a fool's errand. It's a bit like a programming system that requires the developer to express his meaning in both Forth and Python; it's just asking for an author to write text that directly contradicts the markup.
It's a stretch.
This is really hard to take seriously, even as someone who advocates for semantic HTML.
Surely all these nuances mean something tangible, but I don't see it.
Compared to <b> vs <strong> vs <em>, <mark> and <i> seem comparatively sensible.
> The i element represents a span of text in an alternate voice or mood, or otherwise offset from the normal prose in a manner indicating a different quality of text, such as a taxonomic designation, a technical term, an idiomatic phrase from another language, transliteration, a thought, or a ship name in Western texts.
Looking at stuff like this, I think the results are pretty obvious. Despite the “diversity”, the value of what Mozilla offers hasn’t simply stop increasing, it has now started decreasing and in cases like this actually brings negative value to the internet overall.
This is what firing/alienating the actual people who did the actual job does. And no amount of “diversity” can make up for it.
What a shame. Mozilla used to do so much good.
> The u element represents a span of text with an unarticulated, though explicitly rendered, non-textual annotation, such as labeling the text as being a proper name in Chinese text (a Chinese proper name mark), or labeling the text as being misspelt.
I'm not really sure what unfair is, though, to answer your question.
Or are these actually becoming official standard names?
I'm truly enjoying the sheer absurdity of it all.
https://html.spec.whatwg.org/multipage/text-level-semantics....
> B: Renders as bold text style.
The element (and other style-related elements) were deprecated due to CSS. For b and i the strong and em elements were created as semantic alternatives. Now the b and i elements have been given specific semantic usage and are now not deprecated.
From https://html.spec.whatwg.org/multipage/text-level-semantics.... (with _ added for the relevant section used by MDN):
> The b element represents a span of text _to which attention is being drawn_ for utilitarian purposes without conveying any extra importance and with no implication of an alternate voice or mood, such as key words in a document abstract, product names in a review, actionable words in interactive text-driven software, or an article lede.
From https://html.spec.whatwg.org/multipage/text-level-semantics.... (with _ added for the relevant section used by MDN):
> The i element represents a span of text in an alternate voice or mood, or otherwise offset from the normal prose in a manner indicating a different quality of text, such as a taxonomic designation, a technical term, _an idiomatic phrase_ from another language, transliteration, a thought, or a ship name in Western texts.
Deprecated, but it will be back.
<strong>: The Strong Importance element
https://developer.mozilla.org/en-US/docs/Web/HTML/Element/hr
<hr>: The Thematic Break (Horizontal Rule) element
> The <strong> element represents text of certain importance, <em> puts some emphasis on the text and the <mark> element represents text of certain relevance.
This really confused me. Isn't this all basically the same?
"Semantic" HTML seems to be pushed after the fact as a justification for CSS to have syntax completely separate from HTML, reserving attributes for "behavior". The idea that markup attributes shouldn't contain style info is laughable when they're there for the exact reason to associate typed properties with elements.
"Bring levity in kind"
The real problem, however, is that many writing systems have never had any kind of oblique style for emphasis. Or cursive (which is a different thing). Or they have cursive, and use it all the time in print, but technically that's a different font without any connections to the main one. So any person making some kind of template should be aware that the effect of <i>-as-italic can range from “obvious thing everyone expects and understands” to “does nothing at all” depending on language and font. I think they should have got straight to that point in the article.
It is pitiful how those unreflected small choices can put fences in one's mind. Everyone knows the font selection dialog in the browser, and probably thinks that the world is made of serif, sans-serif, monospace, and some Comic Sans. But that's nonsense, even for Latin.
Even more pathetic is inability to use basic punctuation marks. Something as barebone as Fixedsys font had all the quotation marks and dashes since Windows 3.1, long before programs became Unicode-aware. All that time absolutely nothing has been stopping you from pressing the key and getting the proper symbol just like you type any other… except the typewriter key labels hardware makers still use, and the default system layouts being just as stale.
I honestly find it embarrassing rather than funny.
"The foundations of the modern web were built by people who couldn't possibly have imagined the importance their inventions would play in the future, and the sheer breadth of applications they would be used for. One of the consequences of the web's convoluted history is that tag names like <i> are historical artifacts that don't make sense in the context of modern semantic elements."
Is it really so hard to admit that? Why make up bogus explanations?
I think I'll hang onto the original Season 1-4 interpretation for my headcanon.
Italicized paragraphs or long spans are more common than you might guess. Sometimes they're used to indicate speakers in a dialog or to signify editorialized interjections. No idea if browsers are smart enough to render nested <em>'s this way now without specialized CSS.
I've seen <i> used for icons(!) but that's a bridge too far for me.
By the time the web is finally "semantic" and everything is marked up appropriately, we'll have AIs that just infer the semantics and don't need all the meticulous tagging.
meanwhile, yes, google does have to use AI to figure out webpages because the semantic web is a failure.
I will spend time considering if something other than a <div> is more appropriate for an element. I'm not going to spend very much time pondering the difference between <b> and <mark>.
Add something to the document that tells it to ignore screen readers as long as a <screenreader> tag is present that puts everything in one place.
Because nested layers of DIV are apparently "semantic"
BTW: The <u> element has been rebranded as "Unarticulated Annotation". :-)
https://developer.mozilla.org/en-US/docs/Web/HTML/Element/u
(So traditional emphasis is now both idiomatic and unarticulated.)
In this specific case, it’s probably just a bad mechanical translation, possibly from long ago as part of a batch i → em change, or possibly recent in their conversion from HTML to Markdown (might have done both em and i → *). Ill-conceived in either case, but not the end of the world.
If MDN docs is autogenerated, that's regrettable. I swear I'm not going to fall back on W3Schools.
Incidentally, whatever happened to W3Schools? There was a time when every google search for anything remotely similar to web development came up with the top 5 links being to W3Schools. Nowadays not so much. Did google finally decide that W3Scools content amounted to disinformation?
these should include the lang attribute to identify the language
but then don't do that in the Example (with *vini, vidi, vici*)...Although this is not quite correct. "Semantic" is what makes sufficient distinctions for a given case. For example, assume I'm going to publish Ashby's "Introduction to cybernetics". It has an interesting structure: numbered chapters ("2"), smaller divisions within a chapter with a title but without a number, and yet more smaller units that are numbered like "2/17". Some of these smaller units have a title, some just a number. The point is that all this is rather unique.
So I'm about to mark up the text to indicate all this. A semantic way to do that would be to invent a notation that is as unique as this book. There will be "<ashby-chp>", "<ashby-div>", "<ashby-unit>" and such. We do not specify how this is going to be rendered, but at least we faithfully describe what we have without losing anything and without adding anything irrelevant.
Now we want to render it on a visual media. Here we have a different notation that describes fonts, styles, spacing and so on. These are very different distinctions from the structure of the book, but are very approriate for typesetting. We decide how we represent our "ashby" distinctions with these tools and write a transform from our notation into the visual one.
Same for a screen reader. Here we have a notation that describes pitch, speed, etc. No spacing or fonts, of course. We decide how we are going to represent our distinctions with these and come up with another transform.
At each step each element in our notations has a clear purpose. The "ashby" notation captures all the distinctions the author needs to make his point. The "visual" and "aural" notations are tools to express distinctions that can be made on a specific media. This is semantic. And this, by the way, is the original idea of XML (a multitude of notations) and XSLT (the notation transformer).
[And the description of the transforms also uses yet another notation with yet another clear purpose :)]
These are the symbols:
—
Italic:
∠ ANGLE Unicode: U+2220, UTF-8: E2 88 A0
⦢ TURNED ANGLE Unicode: U+29A2, UTF-8: E2 A6 A2
Bold:
⁎ LOW ASTERISK Unicode: U+204E, UTF-8: E2 81 8E
* ASTERISK Unicode: U+002A, UTF-8: 2A
—
I just made this encoding up. Ah, it would be more correct to say I just made this markup up. I made it up and nobody understands it. …Or yeah, actually most people will understand it?
If I have a label to some form element in italic, it would not be correct to write it as <label><i>Name:</i></label> , right?
There's no surrounding text it needs to distinguish from and the fact that it is a label should already point out to screen readers that it's a label, while adding an italic font to the label without the i should distinguish it visually, is it correct?
Also Markdown doesn’t support the distinction.
This is not to say that browsers should prevent styling <i> elements differently, but the intended meaning is still “whatever italics means”.
- Ship or vessel names in Western writing systems
Am I the only one finding this a bit overly specific?
They’ve probably gone to a school whose curriculum includes the terms “critical” and “theory” and “colonialism” more often than not.
This seems backwards to me. Doesn't "idiomatic" mean natural-sounding? Setting it off from the normal text seems kind of the opposite.
Other than the fact that they aren’t semantic, it does make some sense to use small tags if you are concerned about saving bytes.
<i class="fa fa-facebook"></i> <!-- renders the Facebook logo or something like that -->
I believe they presently just use <span> now.
Put another way, where's the assistive technology equivalent of OXO Good Grips? I'd imagine a tool that is part ad-blocker, part auto-summarizer, part navigation helper, part personal curation assistant; that does the Right Thing (tm) 90-95% of the time, even on highly dynamic web applications (since that's already _way_ better than the proportion of sites / applications built with a11y in mind); that people in general find to be a superior user experience, regardless of whether they have disabilities or not.
I read the spec, but there's not much of a hint there.
> idiomatic text, technical terms, taxonomical designations
These were usually put in quotation marks.
For book titles, wrap them inside 《 》
In Unicode, Chinese, Japanese, and even Korean and Vietnamese characters are combined into something called "CJK Unified Ideographs",or just "Han".[1][2]. So they all get the same treatment for something such as italics. Not sure what HTML currently specifies for that.
However, a problem with sarcasm tag is that it wouldn't really help accessibility compared to say saying "Sarcasm:" or something like "(The preceding remark was sarcastic.)".