Making a Website Under 1kB
tdarb.org
tdarb.org
I just recently stumbled upon a site that had it all: webp/avif everywhere, minified CSS, even ditching unused classes from used frameworks, CSS data:-embedded and subsetted fonts (I think it even used a recent version of FontAwesome 6, but it still managed to make the WOFF2 be only 2kB in size because they just utilized like two dozen logos), only one request each for CSS and JavaScript (everything concatenated and with nice cache policies) and the site was still usable/viewable even without either one of these if you wanted to. Everything was automated in their deployment pipeline even. It only came to my notice because they wrote an article about it. I can't find it in my history but those things will stick to your head for a while.
500kB honestly doesn't require much effort at all. My WordPress blog with a popular theme and a few plugins has each almost every article below 500kB, despite looking like any modern blog and having at least 1 image per post.
Actually, if I remove the Facebook share button, they drop to almost 250kB. Time to remove it I guess.
To me it looks like it has always been a fairytale made up by Facebook to spread their analytics scripts all over internet.
I removed it today after seeing the impact on page size.
I guess it works for some niches like news, online quizz, etc.
And with those historic words, the vi vs emacs debate was finally won.
I don't know about that, because my RISC-V SBC with a 1-core in-order 1GHz CPU and no GPU can load Emacs GUI just fine.
(I had 32MB of RAM on the computer I learned Emacs on and it was totally fine.)
But it's microEmacs. Still, it is of course GUI (like nearly everything on the Amiga), and it works fine.
It shipped in the AmigaOS 1.3 extras floppy. Probably also bundled in other versions.
I made an entire organ synthesizer in under 1024 bytes of html/js (use your keyboard):
https://js1024.fun/demos/2022/23
source / instructions / background:
https://github.com/ThomasBrierley/js1024-mini-b3-organ-synth
The javascript was significantly reduced in size by using regpack, which is a regex based dictionary compression algorithm targetting very small self decompressing javascript. Writing for regpack takes some thinking because you have to try and make the code the most naturally self similar, which often means consciously avoiding more immediate space saving hacks in order to make larger sequences of characters identical - e.g it would usually make sense to not duplicate an expression, unless it's very short and only used a couple times, it usually makes sense to store it in a variable, but the character cost in defining the variable is actually larger than simply duplicating the expression through regpack (even if it's used a lot).
This might be worse than the lzw type compression used in HTTP, I haven't tested, but this was written for a code golfing competition where source size counts not transmitted size... Then again, conceptually this should lend itself to any any dictionary based compression, so zipping the original source (minus the whitespace), may also work out very well.
Here's a web interface that includes terser and regpack: https://xem.github.io/terser-online/
Note that it's usually best to hand minify for regpack, terser is only convenient for removing whitespace, but advanced automated js compression libraries usually cause the size to inflate when regpack is the final stage.
[edit]
zipping the source comes in at 952 bytes - so it appears this technique applies when targetting dictionary compression in general, including the commonly used compression in HTTP.
The latency is almost zero. I've always thought of "browser applications" as these heavyweight mammoths where each keystroke needs at least a couple hundred milliseconds to process.
Good to know these kinds of things are still possible. Browser synthesizers have become my new point of interest :)
http://www.captiveportal.co.uk
And yes, that’s supposed to be a non-https link.
I think the entire site, including favicon, might be under 5KB. You can check here:
Anyway, that didn't happen here, because I subconsciously knew that every link would load before I could even think of it, and that none would make coming back one step a pain in the ass, and that was refreshing and maybe even made me nostalgic. But more than anything else, it allowed me to read with more focus than I remember having the last few years. So yeah, I love that "bare ones" design.
(PS: I also realise my comment is so long it could have been it's own blog post. Maybe I should start one...)
neverssl.com solves this by redirecting to a random subdomain (for some reason that isnt clear to me near midnight)
a .co.uk equivalent is a great idea though, if it can be made accessible to users with hostile browsers
SSL/TLS/https (the padlock symbol) prevents this.
This site will never use those technologies.
yet it has https? strange hahahttps://www.fastmail.help/hc/en-us/articles/1500000280141-Ho...
Not a thing wrong with haiku, but for a whole art movement, we'll get more mileage out of a one-trip website.
Ignoring a bunch of caveats I won't get into, a normal TCP packet is no larger than 15kB, for easy transit across Ethernet: the header is 40 bytes, leaving 1460 for data. Allowing for a reasonable HTTP response header, we're in the 12-13kB range for the file itself.
That's enough to get a real point across, do complex/fun stuff with CSS, SVG, and Javascript, and it isn't arbitrary: in principle, at least, the whole website fits in a single TCP packet.
Typically the TCP initial congestion window size is set to 10 packets (RFC 6928), hence ten packets can be sent by the server before waiting for an ACK from the client.
So under 15kB or so (minus TLS certs and the like) website loading has the minimum latency possible given any other network factors.
If there's a session/ticket resumption, the server won't send a certificate, but it does still need to send a negotiation finished message, and it may likely want to send new tickets (I'm not sure if you can delay that though). In TLS 1.3, the client may send the request as early data, if not the request will come as the beginning of the second round trip, so the congestion window will have opened more for the response.
If it's a full handshake, the certificate is part of the first round trip, and the content is in the second round trip; the cert won't count against the congestion window, because it must have been received before the client sent the http request.
They can easily be in the 500 to 1000 byte range, taking up most of the first TCP packet. e.g this HN page has a 741 byte header. I suppose if you control the web server you could feasibly skim this down to the bare minimum for a simple static page - not sure what that would be.
Optimizing for 1 kB can be a fun creative exercise, but I think it's practically a bit meaningless. It's better to target something like a 28.8 Kbps connection and try to get the page to load under a second (including connection handshake < 20.1 KB), which is more than enough to have a rich web experience.
Out of curiosity, I recently wrote a little Elisp function to compute the "markup overhead" on a typical NY Times article, i.e. the number of characters in the main HTML page versus the number in a text rendering of it. It turns out that the page is 98.5% overhead. That doesn't even count the pointless images, ads, and tracking scripts that would also get pulled in by a normal browser. Including those, loading a simple 1000-word article probably incurs well over 99% overhead. Wow!
five minutes later
Nope.
By:
- removing quotes around attribute values
- replacing href values with the most frequent characters
- sorting content alphabetically
- foregoing paragraph and line-break tags for a preformatted tag
I was able to bring it from 730 bytes (330 had you compression enabled) down to 650 bytes (313 bytes after compressing with Brotli). Rewording the text might get you even more savings. Of course I wouldn't use this in production.Here it is: https://jsbin.com/cefuliqadi/edit?html,output
If you’re using the HTML parser (e.g. served with content-type text/html), activities like including the html/head/body start and end tags and quoting attribute values will have a negligible effect. It takes you down slightly different branches in the state machines, but there’s very little to distinguish between them, one way or the other. For example, consider quoting or not quoting attribute values: start at https://html.spec.whatwg.org/multipage/parsing.html#before-a..., and see that the difference is very slight; depending on how it’s implemented, double-quoted may have simpler branching than unquoted, or may be identical; and if it happens to be identical, then omitting the quotes will probably be faster because there are two fewer characters being lugged around. But I would be mildly surprised if even a synthetic benchmark could distinguish a difference on browsers’ parser implementations. Doing things the XHTML way will not speed your document parse up.
As for the difference achieved by using the XML parser (serve with content-type application/xhtml+xml), I haven’t seen any benchmarks and don’t care to speculate about which would be faster.
Unless presented with concrete steps to reproduce what you’re talking about, I refuse to believe you.
(Mind you, I’m not denying in this that there are differences, just that they’re even measurable this way on even vaguely plausible documents.)
Sorting content alphabetically and that sort of thing to improve compression may be silly code golfing and impractical for page content, but on the other hand I don't see that it costs you anything (aside from time experimenting with it) when applied to the <head> / metadata. https://www.ctrl.blog/entry/html-meta-order-compression.html
I think that both of these methods could be used in production, and I intend to do so when possible.
Edit: another version, 59 bytes: https://twitter.com/jonsneyers/status/1375828696846721031
Here's an entry I hacked up together: https://coffeespace.org.uk/colour.htm
It comes in at 1015 bytes and converts HTML colours into their shortened form (i.e. #00F for blue) and displays the colour visually.
Also added https://coffeespace.org.uk/tools/avatar.htm for generating avatars in the same spirit.
Tried again and it works on Firefox.
Instead of:
<link rel="icon" href="data:,">
Try
<link rel=icon href="data:,">
Another fun trick is using <!doctypehtml> since the spec says to pretend a space is there if not present for whatever reason (https://html.spec.whatwg.org/multipage/parsing.html#parse-er...)
The first error read "Bad value data: for attribute href on element link: Premature end of URI."
The 2nd error read "Missing space before doctype name."
Depending on the context these hacks may still be useful, but I personally think that both production sites and code golfing should require valid HTML.
I have implemented some fixed error pages for my company including its logo in svg below 1KiB.
The whole point of the 1MB club and other such efforts are to show what can be done without the equivalent of multiple copies of Doom[0] worth of JavaScript to display what is just a static landing page.
There are completely legitimate uses for "web apps", things that are actual useful applications that happen to be built in a browser. No one is saying that web apps aren't a totally valid means of delivering software.
The problem is every website written using the same frameworks Facebook and Google use for their web apps to build sites that could easily just be some static HTML[1]. We have devices in our pockets that would have been considered super computers a few decades ago. I have more RAM on my phone than I had hard drive storage on my first PC. My cellular connection is orders or magnitude faster with orders of magnitude better latency than my first dial-up Internet connection.
Despite those things the average web pages loads slow as shit on my phone or laptop. They're making dozens to hundreds of requests, loading unnecessarily huge images, tons of JavaScript, autoplay videos, and troves of advertising scripts that only waste power and bandwidth. I find the web almost unusable without an ad blocker and even then I still find it ridiculous how poorly most sites perform.
I love waiting at a blank page while it tries to load a pointless custom font or refuses to draw if there's heavy load on some third party API server. I also absolutely adore trying to use some bloated page when I'm on shitty 4G in between some buildings or on the outskirts of town.
It would be nice if I didn't need the latest and greatest phone or laptop to browse the web. It would be nice if web pages rediscovered progressive enhancement. Add JavaScript to improve some default useful version of a page. You don't need to load some HTML skeleton that only loads a bunch of JavaScript just to load a JSON document that contains the actual content of a page.
[0] https://www.wired.com/2016/04/average-webpage-now-size-origi...
<html> <head> <title>My Wonderful Website</title> </head> <body> <p>What a wonderful page</p> </body> </html>
Ha! There! 110 bytes. A whopping 914 bytes left for user content.
PS: (I think HN hates code blocks)
The linked article is bigger than 1kb, and the <1kb site is just a list of links...
Before: <link rel="icon" href="data:,">
After: <link rel=icon href="data:,">
Checking my own site with `find . -name '*.md' -exec wc -c {} + | sort -h` I find that 30 Markdown files are under 1KiB and 40 are over 1kiB. The largest file is a 18266 bytes post which is still quite small compared to most other blogs AFAICT. It also excludes the possibility of including images with more than a dozen pixels.
But the initial round trip can be 14kb, 10 packets. So under that the user will not need anymore trips.
With brotli level 11 I’m down to 330 BYTES.
They actually mean 1kb per page, which is pretty slick and decent even on dialup.
The two rules for a web page to qualify as a member:
Total website size (not just transferred data) must not exceed 1 megabyte
The website must contain a reasonable amount of content / usefulness in order to be added - no sites with a simple line of text, etc.
The github repo just says: An exclusive members-only club for web pages weighing less than 1 megabyte
So which is it? Are sites with multiple pages under 1MB (but then the total for all pages exceeds 1MB) allowed, or must the entire site weigh in less than 1MB?Then again in age of JS frameworks maybe it is an achievement for the new developer that was gaslighted into thinking 500MB of deps to make a simple site is normal
My site is 0.306 kB.
Narrator: It was extremely easy feat.
Just make spec-invalid webpage and skip all the heads, bodies, htmls and rest of it.
You also don’t need to quote attribute values. You don’t need closing tags for many html elements.