<article> vs. <section>: How To Choose The Right One
smashingmagazine.com
smashingmagazine.com
Which means, if you care about accessibility and want to pick the best element for the circumstances, you must know that you are deciding not on a tag but on element’s role.
Which means, the moment you find yourself asking “should I use <section> or <article>?”, or “should I use <nav> or <toolbar>?”, etc., you… might as well throw a <div> there for right now, and as soon as you’re done with the task head to ARIA docs. It’s a warning sign that you are lacking clarity as to the content hierarchy you’re dealing with.
Consult the source of truth (ARIA docs) instead of picking a role by proxy by reading HTML tag docs; give your div a well-informed choice of a role, and leave it to your linter to complain if it could be made briefer by changing it to a different tag that has the same role implicitly.
You could rewrite every edge case for every element in JS if you wanted, but why? Just understand what the elements mean.
I realize that not everyone is a native English speaker, but in the case of these two words it seems to me that the non-technical, casual definitions or article and section are enough to differentiate them.
Articles can go in a section, but an article "of" a section implies that being contained by a section is an ~identifying feature of the article, which is at odds with our intuitions of how articles and sections work. "I read the first article in the Editorials section" is, at least to my brain, a perfectly cromulent usage that puts the "article" subordinate to the "section", though.
You wouldn't call the article 'a section in the sports article'.
Similarly if you said 'there are ten articles in the section' we'd know it was a newspaper section or similar, that the articles were distinct pieces.
*Use custom elements.* =) We now have carte blanch in modern HTML specs and web browsers to invent our own elements when the built-in semantic variety won't cut the mustard. In one example, it used a <div> for a newsletter subscription area instead of <section> or <article>. Why div? You could actually add <newsletter-subscription> and now you have a tag you can style and utilize PLUS it conveys meaning within the DOM.
Note that I haven't yet mentioned web components…that's because you don't need to author a web component to use a custom element. A custom element in HTML starts life out as an HTMLUnknownElement object in the DOM. If you want to keep it there, that's perfectly fine! As an optional step, maybe you do want to attach some additional behavior via JavaScript. Now you can upgrade it using customElements.define and boom! It's a web component.
I've actually set up linters in some projects now where HTML is checked for <div>/<span> tags and you get an error unless you consciously opt-in for those particular cases. Again, in almost every respect (perhaps a little more relaxed if you're deep in the bowels of a component shadow DOM, but still…), I find reaching for a true semantic element or a custom element is really the way to go in 2022+.
1: Visual noise. I would prefer to edit this:
<div class=article-section>
<div>Hammer</div>
<div>
Buy this and everything will start to look like a nail
</div>
</div>
</div>
Rather than this: <article-section>
<article-name>Hammer</article-name>
<article-description>
Buy this and everything will start to look like a nail.
</article-description>
</article>
2: Web scraping. Div soup makes scraping harder than a doc of nicely named elements.But I can usually grasp very quickly "Ah yes, inside an article-section, the name comes first, and the description comes second".
Writers of scraping software don't like ":nth-of-type(3)" though. They prefer ".price". As it is easier to write and will survive more changes to the site.
discussed here https://news.ycombinator.com/item?id=24055733
<div class=article-section>
<div>Hammer</div>
<div>
Buy this and everything will start to look like a nail
</div>
</div>
<div class=price>123.45</div>
</div>
How would you use XPath, and how would it be more rubust than '.price'?But if there isn't a .price because someone has changed it to .amount for some reason. Or hey started using Material-UI or some other framework and they no longer have any meaningful classes anywhere which is what I most often see people whining about.
then you would probably get an absolute xpath to the item you want, then you would use the getRobustXPath(element, document) function for that element, using the linked library under discussion https://github.com/cyluxx/robula-plus to get a robust xpath. This xpath will probably be pretty unreadable but you can then test to see it gets your stuff.
How will it be more robust - well that depends but if I have something like this
<div class=article-section>
<div>Hammer</div>
<div>
Buy this and everything will start to look like a nail
</div>
</div>
<div class=price>123.45</div>
</div>
I might first off want to get products and their descriptions and their prices together. And I would probably say hey it might be that they are under this kind of class based structure but that can change but what is less likely to change is that a 'price' will beA number with either . or , as the potential decimal delimiter, with potential currency symbols either directly in the element with the number, not found, or in a directly preceding or following element.
It would be useful to find prices on the page before everything else, and from there to figure out where product names and where descriptions where in relation to the price.
Dependent on language of page product layout will probably change, but in a western language the price will probably come 'last' in our product container, which is exactly what you have done here.
A non-robust xpath for matching the above might look something like
product = //[local-name()='article' || local-name()='main']//[child::*[last()][xs:decimal(text())]
The above might have an error in the xpath because hey, writing without testing in environment, and also obviously really expensive as have not optimized. But it should also be clear enough that by removing the need to have matching elements be based on the incidentals of how we style those elements we can add up with data extraction that is less likely to break with site redesigns.
on edit: obv. article or main check would be optimized beforehand, just some sites use both, some use one, and some neither so would have to do all sorts of things first, etc. etc.
<div class=article-section>
<div></div>
<div class=price>{.}</div>
</div>
It is the most robust. It picks the price in the article and not some other price that might on the page. And if there is no price, you get an exception that there is no price, rather than an empty result <article>
<name>Hammer</name>
<description>
Buy this and everything will start to look like a nail.
</description>
</article>
And shouldn’t the doc version have classes for, if not all then at least the hammer div?1: https://developer.mozilla.org/en-US/docs/Web/API/HTMLUnknown...
<dl>
<dt>Hammer
<dd>Buy this and everything will start to look like a nail.
</dl>I guess div/span don't offer much but I assume the screen readers etc. at least have experience dealing with that vs. randomness.
Of course divs and spans do get the benefit that they are read as separate elements by screen readers, however this can also be a curse https://stackoverflow.com/questions/60869293/accessibility-r...
Don't know how it would behave with a bunch of readable-segment elements or what have you, as I haven't tested.
That don't work without Javascript and that have multiple unsolved issues ranging from accessibility to not playing nice with SVGs.
No. Use divs and spans.
It's not because custom elements "work without Javascript". It's because browsers are trying to work even in the presence of errors and invalid/unknown markup.
Javascript is required for custom elements.
<div> to me is like the generic building "block", not meant to convey meaning most of the time, and you can use aria-role when needed.
I don't think we needed <section> and <article>, many things in HTML-CSS naming/conventions don't make sense to me, after 15 years. Too many ways to do the same thing.
Under that understanding I think you can probably still find interesting semantic versions of these tags in the XHTML 2.0 mailing lists and schism discussions. They aren't relevant to HTML's present, but might be interesting for someone truly curious about the path not taken in semantic HTML (the path unlikely at this point to ever be taken).
The hgroup elements also seems to be related to this.
While any such interpretation is somewhat funny in the context of the parent comment, it may still turn out useful. E.g., if we were to scrape any content from an existing site in order to reintegrate it for a relaunch or a similar purpose.
And, as we're at it, a div is really just a technical means for applying something to a group of elements (e.g., in it's a original use, an attribute for centered text presentation), think of it as blocks in programming. Nothing semantic to see here, keep calm and carry on…
BTW, thanks for mentioning the hgroup, which is often overlooked, but really makes sense, when combining headings and subheadings, which are to be understood as a single item (like the head of an article, yes, an article in the common sense).
My issue with them is not with their roles, but with their names. And, from the article and from OP, it seems I'm not the only one. I think "region" as it is used by WAI-ARIA would be a better name. Also something like "contentinfo" instead of "footer". And "complementary" instead of "aside"...
> E.g., if we were to scrape any content from an existing site in order to reintegrate it for a relaunch or a similar purpose.
The spec call this "outline".
Related to divs. I find ironic that making pages with tables were frowned upon 20 years ago, yet it is hot again now, but we are calling them "grids".
They say, the lack of usage of hgroups is due the lack of support by screen readers. Another common use case is <h1>Chapter 1</h1><h2>Foobar</h2>.
(Tables should even be more accessible, since there is <th>, both in <thead> and with `scope="row"` for table rows.)
Something, I've been guilty of (sometimes) for emulating hgroup: <h1>Heading<br /><small>Subhead</small></h1>.
XHTML 2.0 removed 42 tags: acronym, applet, area, b, base, basefont, bdo, big, button, center, dir, fieldset, font, form, frame, frameset, h1, h2, h3, h4, h5, h6, hr, i, iframe, input, isindex, label, legend, map, menu, noframes, noscript, optgroup, option, s, select, small, strike, textarea, tt, u.
XHTML 2.0 added 18 tags: access, action, addEventListener, blockcode, di, dispatchEvent, heading, l, legacyheadings, listener, preventDefault, removeEventListener, ruby, section, separator, standby, stopPropagation, summary,
AFAICT, XHTML 2.0 reorganized tags into modules, yes, but didn't actually try to expand the set of semantic tags, except for XForms--the XForms module looks really complex. And those module groupings were more concerned with functionality, not content semantics, per se.
FWIW, here are the XForms 2.0 tags: action, delete, dispatch, group, input, insert, load, message, model, output, range, rebuild, recalculate, refresh, repeat, reset, revalidate, secret, select1, select, send, setfocus, setindex, setvalue, submit, switch, textarea, trigger, upload
By contrast, HTML5 looks to have added more semantic tags, and more incoherently. HTML5 has 111 elements (excluding math and svg).
HTML5 removed 14 tags: acronym, applet, basefont, big, center, dir, font, frame, frameset, isindex, noframes, param, strike, tt
HTML5 added 34 tags: article, aside, audio, bdi, canvas, data, datalist, details, dialog, embed, figcaption, figure, footer, header, hgroup, main, mark, meter, nav, output, picture, progress, rp, rt, ruby, section, slot, source, summary, template, time, track, video, wbr
Source: https://www.w3.org/TR/html4/index/elements.html, https://www.w3.org/TR/xhtml2/elements.html, https://html.spec.whatwg.org/multipage/indices.html#elements...
I think the problem of bad architectural decisions has been understated, possibly since the dawn of computing.
An example of greatness is HTTP and SMTP. I'd argue most cases don't even need HTTP/2 or HTTP/3 -- getting rid of all the tracking/bloatware on the modern internet would deliver more value for the end-customer but that's not the direction we're going in. The industry has boomerang'd back to nearly static sites (I think there was progress made, IMO SSR is a local maximum), so honestly HTTP/1 is often good enough.
An example of sadness might be Bluetooth. Every time I sit down to look at some docs that involve it (mobile, otherwise) I am horrified anew.
Bad standards basically set the entire industry back person-decades.
So normally I'd probably go look at something like:
https://tdg.docbook.org/tdg/sdocbook/5.2/article.html
and
https://tdg.docbook.org/tdg/sdocbook/5.2/section.html
for some inspiration.
However - tfa was a lot more interesting and pragmatic - giving good advice on accessibility ; something that is actually worthwhile and not just silly bikeshedding...
Which would have been fine, except even the most obvious ones (header, footer) were given very idiosyncratic definitions, and others (article, section, aside) were seemingly thrown in at random.
This led to absurd examples where, as is still in the spec, a blog comment is an <article> and the comment's header is a <footer>. This of course undermined the original premise -- that these elements were just 'paving the cowpaths' of how people were authoring HTML, but the spec author would have none of it and shipped them all the same. And here we are. :)
I met the guy on the IE team who created the <span> element. He just had the notion, coded it up over a weekend, merged, and then shipped.
Yes, <span> was useful and needed. Bravo.
And yet, I was so grumpy that he and the whole IE team treated this stuff so casually, so ad hoc.
This was back when I still cared about standards, correctness, interop, compatibility, following the rules.
It was a phase. I'm over it. I've accepted that programming is now archeology, forensics, and scrapping.
I would start with <block>, which would take the place of <div>. But, big twist, you could put anything inside of <HERE> and get default block behavior. That way you get <article> and <aside> for free and people would know that it’s arbitrary, thus feeling more free to be descriptive with their tags.
</rant>
Custom elements are part of HTML5: https://html.spec.whatwg.org/multipage/custom-elements.html#...
And you can override the "display" property with CSS.
Is there any SEO advantage to using header/article/section/footer type tags ?
<section> once drove the model where we were going to kind of soft-kill the h1-h6 elements and have only h1s where their heading level depended on the section nesting level... but my understanding is that this didn't really work out and is functionally dead, and the current recommendations are all to just use the range of numbered headings as before.
Do screen readers use a different user agent? Seems like a11y would be simpler in many cases by just delivering a text based result that's tailored to screen readers. Most of the issues seem to arise from the structure chosen in order to style things rather than the content itself.
And just expecting that vaguely guessing what the other wants will work.
And probably some optimization on the engines to sort the fetching of the resources and painting the screen.
> Should I Bother? Yes, you absolutely should. [...] Firstly, this article by Mandy Michael reveals that browsers pay attention to your HTML structure to generate a Reader mode which strips the page of unnecessary information, images and background.
Reader mode is not a universally good reason to be using article and section elements. It only matters if you care that much about Reader mode, and Reader mode is not a part of any web standard. Make your webpage so readable that Reader mode isn't necessary.
> Secondly, It helps you actively think about how you present your content.
Depends on if it does. I've used both article and section since time immemorial and I honestly couldn't tell you if they helped me create better markup. I just use them because they seem logical and they're there, but I wouldn't really miss them if they were deprecated. Article is only really useful if you have multiple article-like places in your markup, since it actually has an "article" role, but sections are just like any other region in terms of accessibility and aren't that useful.
There is no 'outside of accessibility' - or shouldn't be.
Sure, there should be nothing outside of accessibility. /s
If you read the post (and I can't really blame you for not reading it all the way through), you'd see that the author is suggesting that there are reasons besides supporting accessibility features for using these semantic HTML tags. I just don't agree that they are objectively good reasons for why one "should" do so.
Oh yeah, and semantic HTML is by no means the only way to write pages that are accessible. Everything that a semantic tag can do, the role attribute can do, too. I never said to not care about accessibility.
If you use main, article, etc is easier to write apps to process the page. Firefox reader mode use the main element to know how get the text from the page.
<button> does far more than <div role=button> can do.
Even if you meant to only refer to landmark roles, why write <div role=main> or <div role=region> when you can write <main> and <section>?
As someone else pointed out, Reader modes don’t recognize roles and user stylesheets don’t either.
<html/>
That's the most accessible website in the world. You can print it without paper, character encoding is no issue, RTL is no issue, you can see it when you are blind, you can listen to it spoken without connecting your headphones, even if you are deaf you can hear it spoken, it's printed in braille on the screen (or any flat surface), it works in all languages spoken and written and sign language. It really is the pinnacle of accessibility. It's even available in dead languages such as Linear A and Etruscan. Every website with any sort of content is strictly less accessible than this website (which also is as utterly useless as it is accessible).The minimal “valid” HTML page is actually this:
<!DOCTYPE html>
… though it’s only valid in contexts that don’t require an in-document title, such as email text/html parts (the Subject header from the MIME message is used) or frames (no title is needed); for a top-level page in a web browser, to be “valid” it needs a non-empty title, such as <!DOCTYPE html><title>x</title>
In both of these, the <html> open and close tag are omitted, as they’re both optional in HTML syntax.If you’re not caring about validity or quirks mode, then you might as well just serve a zero-byte text/html body.
As for XML syntax which you seem to have tried to use, <html/> wouldn’t be a valid HTML document because you haven’t used the http://www.w3.org/1999/xhtml namespace.
I never said that it doesn't.
> If nothing else because it is a legal requirement
Whether something is a law or not has nothing to do with whether something is a moral obligation or a correct decision.