Fun Facts on Producing Minimal HTML
blog.notryan.com
blog.notryan.com
Web standards specify how to render valid HTML in a standard way, and browsers have become better and better at that over the years.
Invalid HTML destroys all that. Each browser will have to guess how to repair and fill in the blanks on malformed incomplete tag soup, and there is no one right way to do that. Browser X will make different guesses from Browser Y, and the next version of each will be different too.
Please just write actual HTML that is valid and your website will render much more consistently and reliably across shifting platforms, browsers and versions. It is not the size of HTML that slows the web down anyway.
https://html.spec.whatwg.org/multipage/parsing.html#parse-er...
In addition, most of the suggestions in the link (leaving off <html> and <head>, leaving off quotes for attributes, not closing <p>) are valid HTML in the first place.
I write scrapers a lot (not the irresponsible kind and never for monetary gains) and invalid HTML, while technically parseable and to spec, are often a pain in the ass. You have to bring in a full blown HTML5 parser, and they could be way slower (e.g. lxml.html vs html5lib). Depending on your language of choice there might not even be a to-spec HTML5 parser available.
So, just close your damn tags (except self-closing ones), and close them in order, it’s not hard, the size increase is minimal, it will help with your own sanity and people will thank you for it.
But in any case, you are never getting all the web sites in the world to validate, so what does it matter if you can get some of them to validate? You still have to handle the rest.
Writing unclosed/out-of-order tag soup is extra work, not less work.
What a weird hill to die on.
It's not a question about being pro or con closing tags or which is "best". The question is about the billions of pages which already exist and is not going to change.
If you write a web scraper, that is just the reality you work with.
Did I ask people to not follow bad advice (intentionally not closing non-self-closing tags just for the hell of it)? Absolutely.
Parsing an HTML document without a correct HTML parser is just wrong. (I decline to call it an HTML 5 parser, because it’s not strictly that any more, and this stuff is from more than ten years ago: get with the program!) I’d lump it in with using regular expressions to parse HTML: useful in certain situations, but not wise for the general case.
So you have a fast but wrong HTML parser? Find a correct HTML parser. If lxml.html doesn’t parse correctly (I don’t know whether it does or not, I haven’t checked—but suppose it doesn’t), then I maintain that you should never under any circumstances use it in new code. It’s an artefact of twelve years ago, and it’s bad.
So you have a correct but slow HTML parser, and you’re not happy with that? I could say “if you care about performance why are you using Python anyway”, but that would be disingenuous—though some of those sentiments can still reasonably apply. Anyway, your remedy is to find a fast and correct one. I confess myself a tad surprised to find no stable or featureful bindings to html5ever, which is one of the best things along these lines. https://github.com/SimonSapin/html5ever-python and https://pypi.org/project/htmlpyever/ are two things that at least start on this, but it looks like nothing interesting’s happened in this space for about three years. Huh.
But anyway: I decline to adjust my authoring practices because you refuse to use the right tools. Enough other people use the right tools that I don’t need to care. It’s rather like the XHTML/HTML situation if you squint: the only reason to use XHTML, which would reject an invalid document (while HTML would let your nominally invalid document do something useful), would be if something you were interacting with required it.
Ah, but this post is (seemingly accidentally) recommending you use quirks mode instead of HTML5:
> * <!DOCTYPE html> is required by the HTML spec (but not the XHTML spec). Most of browsers will not care if you omit it. It was created for backwards compatibility with prior HTML versions. You can omit it in minimal documents.
HTML5 specifically permits tag omission in many cases [1], including but not limited to both the opening and closing <html>, <head>, and <body> tags. The parsing algorithm provided in the standard generates a DOM tree that includes the omitted tags in the correct place(s).
For example, the following document is valid HTML5:
<!DOCTYPE html>
<title>Hello</title>
<p>Welcome to this example.
[1] https://www.w3.org/TR/html52/syntax.html#optional-tagshttps://developer.mozilla.org/en-US/docs/Mozilla/Mozilla_qui...
(I’m not trying to be hostile, just seeking that misinformation be corrected, and you’ve made a claim that doesn’t make sense to me unless it’s from quite some years ago.)
Now it this was about XHTML you would have a point.
People often mix-up optional syntax and error recovery. You can leave out the <head>-tags, the parser can infer it - this is well-defined by the DTD. It is not an example of error recovery.
Please don't. If it's plaintext serve it as such and if it's html serve it and format it as such.
> <!DOCTYPE html> is required by the HTML spec
Yes, it is and do it. Disregard everything the author wrote after that on this point.
> <html>, <head>, <body> are not required by modern browsers.
Agreed, they should be left behind. This is a valid HTML doc:
<!DOCTYPE html>
<title>test</title>
<p>test doc here</p>
> You don't need to close your tags.Some tags need closing, some don't. This is documented in the standard. Follow it, don't just freestyle it.
> If you do not define <meta charset="utf-8">, then most browsers will default to ASCII or Windows-1252. So if you are confidently in the ASCII range, skip the declaration.
Please don't. Set the charset in your headers (and it should be utf-8 unless you have a very good reason).
> Using a preformatted block (<pre>) of links
This is just bad. You have a list of links but you don't use the exact element created for creating lists?
Legitimate question: why? If I'm not planning on using non-ASCII characters, why bother?
Because not all character sets are ASCII compatible, and you don't know that your user's default is, even though most browsers' defaults if not customized are.
IF you have any user input at all (forms...), you will end up with major headaches.
<!doctype html>
<meta charset="utf-8">
<link lang="en" title="All blog posts feed" rel="alternate" type="application/atom+xml" href="feed.xml">
<title lang="en">Blog posts</title>
<body lang="en">
…
I’d rather just do it once: <!doctype html>
<html lang=en>
<meta charset=utf-8>
<link title="All blog posts feed" rel=alternate type=application/atom+xml href=feed.xml>
<title>Blog posts</title>
…I prob wouldn’t on a major project but only because there’s negligible cost savings.
I find the browser arguments (no pun intended) kinda weak. You can test on a couple platforms and have 100% of an internal corporate environment and 99% of the public internet user agents covered.
All programming is choices and opportunity cost. We make decisions based on maxing the value/cost equation. It may make sense for me to cover 99% of user agents and say screw the 1% if my cost benefit warrants it.
It does not mean I won't follow best practices for critical systems.
If anything, this lets me understand better why there is a best practice for something compared to blindly following them no matter what. For instance I might run into issues when I don't follow them for low critical systems.
- opening <a> tags close the previous <a> tag... but without an href they do nothing. Use them as if they were closing <a>s to save on slashes
- <select>...<select> is completely equivalent in html to <select>...</select>. Save more slashes this way
- formatting elements won't be closed automatically, they rear their head again like so many hydras until your burn the wounds with end tags. But! there's a way around this: consider '<div><b><b><b><b></div></b></b></b>X'; that's funny... why isn't the X bold? it turns out that only three identical (down to attributes) formatting tags are remembered. Use this to your advantage when nesting identical formatting tags inside themselves.
- since we're saving on slashes: <table> (while in a row, ie <tr><table>) closes the last table, and opens a new one. You might think: "I don't want to start another table so soon!" fear not! the browser will move anything you put in a table, above the start of the table, right up until you start putting cells in it. This also avoids having to open the table later, you can start <tr>ing right away.
- you saved characters by dropping the doctype... but was it worth it? only with a valid doctype declaration will you <p> tags be closed automatically when you open tables. Just 4 closing </p> tags will make you wish you included that doctype. still think it's worth it?
<p>one<p>two<p>three<p>four
create four <p> elements nested together? does this only work with <a> tags?your third point about the triple-identical (reminds me about TCP ACK and Re-transmit haha) is pretty nuts though
No, because “A p element's end tag may be omitted if the p element is immediately followed by an address, article, aside, blockquote, details, div, dl, fieldset, figcaption, figure, footer, form, h1, h2, h3, h4, h5, h6, header, hgroup, hr, main, menu, nav, ol, p, pre, section, table, or ul element, or if there is no more content in the parent element and the parent element is an HTML element that is not an a, audio, del, ins, map, noscript, or video element, or an autonomous custom element.” https://html.spec.whatwg.org/multipage/syntax.html#syntax-ta...
1. thanks for the link. that documentation is realllly put together well.
2. for anyone else who goes to the docs, scroll down to the table bit. Official HTML never looked so close to markdown before. But I guess it's legal, with the <td> omissions, etc.
cool stuff!
EDIT: posting code here
<table>
<caption>37547 TEE Electric Powered Rail Car Train Functions (Abbreviated)
<colgroup><col><col><col>
<thead>
<tr> <th>Function <th>Control Unit <th>Central Station
<tbody>
<tr> <td>Headlights <td> <td>y
<tr> <td>Interior Lights <td>x <td>y
<tr> <td>Electric locomotive operating sounds <td> <td>
<tr> <td>Engineer's cab lighting <td>x <td>z
<tr> <td>Station Announcements - Swiss <td>x <td>y
</table>- Use relative URLs when possible, i.e., /page.html, no need to specify the protocol.
- Use shorthand CSS properties like font, background, margin, border, padding, and list.
- Use lowercase tags as they compress better. See: https://encode.su/threads/1889-gzthermal-pseudo-thermal-view...
- Use shorthand hex colors.
These optimizations are harmless compared against some of the ones recommended in the linked article...
But you're right; relative by itself would refer to relative to the current location. Error on my part in not being specific enough. In retrospect "site relative" is not good terminology.
Example, for "http://mysite.com/dir/page1.html":
* Relative: "../other_dir/page2.html"
* URL relative: "/other_dir/page2.html"
* Absolute: "http://mysite.com/other_dir/page2.html"Protocol relative URLs can also be used for different domains as long as protocol remains the same, which is often the case with https: so widespread, e.g.:
<script
src=//code.jquery.com/jquery-3.5.1.slim.min.js
integrity="sha256-4+XzXVhsDmqanXGHaHvgh1gMQKX40OUvDEBTu8JcmNs="
crossorigin=anonymous>
</script>Also worth mentioning, the ultimate minimalism "hack" is to simply serve a txt or md file w/ Content-Type: text/plain
Closing tags are optional for the following tags:
html
head
body
p
dt
dd
li
option
thead
th
tbody
tr
td
tfoot
colgroup
The following tags are self-closing and should not have closing tags: meta
img
input
hr
br
Attributes can be left unquoted if the following characters don’t appear in the attribute value:- Single quote (')
- Double quote (")
- Space ( )
- Equal sign (=)
- Greater-than sign (>)
> This snippet is copy and pasted a whole lot around the web. Most people don't explain how it actually functions though. "width" sets the initial width to the mobile's physical display width in 100% pixels.
I still don't get this - if the browser is the width of the device, why doesn't the site flow to 100% of the browser?
The tag sets the width to the devices actual width (although pixel scaling can change the display resolution) and the initial scale sets the scale to not imitate a desktop.
The viewport meta tag let's you adjust this virtual window size. You can set it to whatever value you want but "device-width" will just use the screen's current width according to its current orientation. You can add it to a page's header and in many cases it'll look way better on mobile even without specific CSS targeting mobile.
The default viewport size of the original iPhone was in place to let it use the "normal" web. At the iPhone's introduction this was a big deal because other smartphone browsers worked best with simple "mobile" layouts. Safari on the iPhone rendered web pages as they looked on the desktop. Making the viewport sized to the display usually makes it so a user doesn't need to zoom in to read text or horizontally scroll to read wonder content. This behavior makes modern mobile browsers render pages more like old mobile browsers where the viewport was usually the portrait screen width (320px typically).
Besides setting the viewport width you can set a couple CSS properties on block elements to keep stuff more mobile friendly. I suggest `max-width: 100%;` and `overflow-x: scroll;`. This would keep for instance those `<pre>` blocks on your page from causing a horizontal scroll on mobile.
was that so hard?
Not always true. I’ve run into numerous issues caused by the lack of closing tags, and just did earlier this week.
<div>
<p>
<h1>
<h2>
without closing tabs create the DOM tree div->p->h1->h2if you were actually developing production code and misplaced, let's say, <p>'s closing tag, then that would mess up the rest of your tree (from your perspective -- the computer doesnt care)
Many of the rules for unclosed tags are more there so that browsers can agree on what to do with garbage first, and for you to rely on only incidentally! They defer to historical practice before common sense!
In order to predict this reliably, you essentially need to have the list of content categories[1] memorized (or look them up). Not all of them are ... necessarily intuitive.
[1]: https://developer.mozilla.org/en-US/docs/Web/Guide/HTML/Cont...
For example, <p>A<p>B</p>C</p> looks like two nested <p> but it is parsed as 3<p> next to each other: <p>A</p><p>B</p>C<p></p>.
"The end tag may be omitted if the <p> element is immediately followed by an <address>, <article>, <aside>, <blockquote>, <div>, <dl>, <fieldset>, <footer>, <form>, <h1>, <h2>, <h3>, <h4>, <h5>, <h6>, <header>, <hr>, <menu>, <nav>, <ol>, <pre>, <section>, <table>, <ul> or another <p> element"
(i.e. the things that aren't valid children for a <p>, iirc)
so i believe your example would end up parsing as
<div>
<p></p>
<h1></h1>
<h2></h2>
</div>
and not as <div>
<p>
<h1>
<h2></h2>
</h1>
</p>
</div>
(i'm not sure how exactly the <h1>/<h2> would behave - an <h1> can't have an <h2> child, but they don't have self-closing rules. it's probably in the spec somewhere) <meta charset=utf-8>
Without this, your document may be interpreted using an implementation-defined ASCII-like encoding (e.g., windows-1252 for English-speaking locales) if served without a Content-Type.My vision is poor, and reader view is an amazing help.
I also like the readability of a document where all closing tags are here. Some tricks can arguably make the code more readable (ommitting head and body) but some tricks require effort to understand if you are not used to them. We write code for human beings first.
Maybe an HTML minimizer could be used if one wants to save bytes?
Also, why?
If your own site is https only (which it should be) then <a href="//example.org/">example</a> is the same as <a href="https://example.org/">example</a>
Nit-picking, but rather than saying "99% of browsers" (and I am not sure there are 100 browsers out there to get that particular stat) it would be best to just mention the ones it doesn't work on.
It would be simple to have a browser or plugin detect and render these, assuming they don't already.
HTML Code Golf - How to make really small HTML that doesn't break Firefox or Chrome, currently at least
FTFY
<meta name=viewport content="width=device-width, initial-scale=1">
The `, initial-scale=1` has been unnecessary for a few years now (sorry, not searching for the citation now, hopefully you can find it if you’re interested), so it’s slimmer to use this instead: <meta name=viewport content="width=device-width">
If you drop the quotes, it’ll parse the same way (and parsing is well-defined in HTML now, so you can be confident all browsers will handle all parsing the same) but be nominally non-conformant. Up to you how much you care about non-conformance, but I avoid writing non-conformant documents, though I do regularly hand-write minimal HTML (omitting html/head/body, skipping unnecessary closing tags, unquoting attribute values, &c.)----
<!DOCTYPE html>
I strongly recommend against omitting this, because removing it throws you into quirks mode.Also I recommend spelling it `<!doctype html>`, because that will regularly save a byte or two in gzipping due to the much greater frequency of lowercase letters.
----
<meta charset=utf-8>
I strongly recommend keeping this; it’s a safe bet that your text editor is working in UTF-8, so specifying the charset thus ensures that if later on you insert some non-ASCII (e.g. pasting a quote that includes curly quotes) it will work properly.You can also save one more byte by spelling this `<meta charset=utf8>`. There’s fun history around that in the encoding spec, where that used to not be a valid value, but based on observing people spelling it that way sometimes they added it. So it’s now valid, but particularly old browsers might not like it.
----
Using <pre> for line breaks? Please no. Just don’t do this. The side-effects are awful. Use <br> if line breaks is all you want.
----
> Both single-quotes and double-quotes are valid for tag parameters. This is useful for producing valid HTML output in programs without resorting to escaped double-quote.
I like to do minimal encoding. Within a quoted attribute value, the only characters that need to be escaped are & and the particular quote used, so for an attribute that includes a large blob of JSON I like to use single quotes, so that I only need to escape single quotes and ampersands:
<a data-json='{"department":"R&D"}'>…</a>
This yields a smaller and more human-readable result, which is also nice.----
> You don't need to close your tags.
Well, there are three cases to consider here:
1. Self-closing tags like <meta>, which don’t have a closing tag (so it’s actively wrong to include one).
2. Tags for which the end-tag is optional, depending on what follows it, e.g. <p> doesn’t need </p> if it’s followed by various elements such as another <p>.
3. Non-conformant documents where the well-defined parsing behaviour just happens to produce what you want, despite what you’ve written being probably nonsensical and something where a human wouldn’t be sure what you meant.