What is the HTML compressor using? It’s missing quite a few opportunities.
Take this input HTML:
<!doctype html>
<html lang="en-au">
<head>
<meta charset="utf-8">
<title>This is a title</title>
</head>
<body class="foo">
<h1>Well.</h1>
<p>I wonder…</p>
</body>
</html>
It produces this output: <!DOCTYPE html><html lang="en-au"><head><meta charset="utf-8"><title>This is a title</title></head><body class="foo"><h1>Well.</h1><p>I wonder…</p></body></html>
It converted “doctype” to uppercase (equivalent, but bad for compression). It didn’t strip <head> and </head> as superfluous. It didn’t remove unnecessary quotes around attribute values. It didn’t remove the unnecessary </p></body></html> closing tags.Here’s what I say it should have emitted:
<!doctype html><html lang=en-au><meta charset=utf-8><title>This is a title</title><body class=foo><h1>Well.</h1><p>I wonder…
30 characters shorter, and structurally equivalent.