No one as far as we know, inside or outside Google, ever figured out how to use server push to consistently speed up loading pages.
No one as far as we know, inside or outside Google, ever figured out how to use server push to consistently speed up loading pages.
Vroom [1] is a research system to accelerate page loads. It identifies dependent resources for your page and uses a combination of client hints and and server push to improve network utilization. I believe that only important, non-cacheable content use server push. See the paper for results.
To be clear, web performance is not my main field and I haven't kept up with the latest happenings. I only bring it up as evidence that push can be useful, but I can't actually say whether it's worth the complexity.
EDIT: fixed some typos.
[1] https://web.eecs.umich.edu/~harshavm/papers/vroom_sigcomm17....
The webserver could just parse the HTML and send http push for each dependency! What was the problem with doing this?
To me this seem very similar to "open document format" are storing html with all its dependency in a ZIP package. So with HTTP/2 push you just push the whole package and the client tell the server if it already have some of the files.
Well… the client knows which resources are already in its cache. The server does not. You just suggested that the server should always send resources that client almost certainly already had cached, which is wasteful for everyone involved.
> So with HTTP/2 push you just push the whole package and the client tell the server if it already have some of the files.
That doesn’t make sense. Once the files have been pushed, the bandwidth and time has already been wasted. The client doesn’t get to tell the server anything in that scenario.
If it does already have that file, client simply close the stream it doesn't need to download the file or send request to server to say it doesn't need the file.
I guess I hadn’t realized that clients could RST_STREAM on these pushes, but it doesn’t change the outcome here.
What you describe isn’t a win for anyone except a client with a cold cache, and then they start losing immediately after that. That’s why it isn’t done. That’s why HTTP/2 Push is going away.
The spec say
"The server SHOULD send PUSH_PROMISE (Section 6.6) frames prior to sending any frames that reference the promised responses. This avoids a race where clients issue requests prior to receiving any PUSH_PROMISE frames."
HTTP/2 Push is such a cool concept, but the idea of it going away also makes perfect sense to me after years of not seeing anyone find benefit.
Server push, as opposed to preload/hints, only makes sense when you know a) a client will absolutely need the data and b) you're reasonably sure the client does not have the data yet (e.g. data that is uncacheable).
Stuff like resource media queries for instance. How the server can know if user wants dark mode or light mode CSS file?
On your particular question, there's the client hint Sec-CH-Prefers-Color-Scheme: https://web.dev/user-preference-media-features-headers/
I don't see this header being sent in Chrome 104!
Note that the server has to request that header using accept-ch.
curl --head https://sec-ch-prefers-color-scheme.glitch.me/
HTTP/2 200
date: Fri, 19 Aug 2022 18:55:43 GMT
content-type: text/html; charset=utf-8
content-length: 936
x-powered-by: Express
accept-ch: Sec-CH-Prefers-Color-Scheme
vary: Sec-CH-Prefers-Color-Scheme
critical-ch: Sec-CH-Prefers-Color-Scheme
etag: W/"3a8-iEB3drxZIB7EZYHQL534qwomDuI"Cookies? Custom header?
How exactly modern web ended up in a situation where, first time browser is displaying a single page (with maybe 3 pictures and 2 paragraphs of text), it has to download 300 resources from 12 different servers, including 2 megabytes of minified javascript?
Not sure about typical, but here is rough estimate for the rock bottom of it: "feh" image viewer on Ubuntu/X11:
$ strace -e openat -o >(wc -l) feh -F \
one.jpg two.jpg three.jpg
105That said, I would not be surprised if XFree86 was, at some point, significantly larger than contemporary web browsers, given the enormous amount of functionality that was loaded into it. Of course, modern web browsers definitely trump any old X server, so is this really relevant? I guess it depends. I get a huge sense that what people are longing for is a return to the "good ol days" when software was simple and fast, and X is a damning counter example, though it certainly doesn't mean that there's no point to be made at all, just that it is perhaps less black and white than people suggest.
I am not trying to say that feh is as complex or as slow as a web browser, only that in general, people treat the division of "native code" and "browser code" such that code running in a browser is always more bloated and slow. Somehow though, this is demonstrably not true. Looking at only very simple software with intentionally very limited scope is not going to give the whole picture. When you look at fairly complex software, the additional burden of having a browser underneath can fade away compared to algorithmic complexity issues and poorly optimized code. Your text layout and font shaping routines are probably going to struggle to compete with a browser if you have a very heavy layout to handle, and if you're doing accidentally quadratic things it's going to matter more than JIT or GC overhead.
And this time it's orders of magnitude worse.
Strawman argument.
I just want to read a damn article. Why does it need to download hundreds of resources from dozens of origins? Reading a PDF article on my desktop touches exactly one file, the PDF. Ok, maybe some dylibs from the OS related to PDF rendering. But surprise! Browsers already have everything built-in for HTML rendering! No need to download an HTML rendering engine, because it’s already there: it’s the browser!
The real waste is in the time it takes to create and destroy the connection. (hence why http/3 switched to UDP, and also why multiplexing is a thing.)
Thing is, combining js and css on the server side into 1 file pretty much negates this speed benefit. I've been building sites this way forever, and performance is great. First page takes longer, but then those files are in the cache anyway.
The actual problem is Ads. They require tracking, and of course displaying. If you're looking for waste, or indeed why a server would bother with http/3 then look no further. No ads, no problem.
Unfortunately for those sites that rely on advertising for revenue, I don't have an alternative suggestion. I get it, ads pay for lots of the Internet I "browse". I'm not a fan of pay walls and subscriptions.
So http/3 is a solution to an economic reality - and really only benefits the ad-supported Web - but for them it is important.
The wife's family living away from the main city have intermittent connection and slower speeds. We have only one submarine cable and poor infrastructure further out. So heavy, bloated sites eat up people's pre-paid internet packages quickly or have trouble loading altogether. Prepaid is around $1.25 for 500MB or $10 for 5.5GB.
I think lower bundle sizes should be a goal. Or lower bandwidth websites should be available such as a 'm.' Domain. IMO
Statements like this reminds me we live in different planets. I'd argue we should still care, just like we care about accessibility in general, but obviously it's indeed a lost cause. Anyone not on 1Gbps and 32 GB of RAM is to be relegated
The problem is that who make websites, instead of making documents, maybe built dynamically by the server and with some interactive elements in it, doesn't sent you a document but instead sends you an application that builds the document into the browser, basically turning an application that simply displays you document into a virtual machine that runs an entire graphical environment.
The fact is that most of the websites to function doesn't need a line of JavaScript, and still are downloading Mb of libraries to do what?
Most of the problems wouldn't be even present if we reason in terms of documents, and in terms of operations that you can do on these documents (GET, PUT, DELETE, POST). You find out that for most website you don't need JS!
One could argue that JS is an optimization because you don't load up again the whole page each time you navigate to a different URL. That is false: if you cache the assets that doesn't change between pages such as stylesheets and images correctly, downloading just the HTML of the page is something trivial. Look at the number of requests that a webpage that uses JS does: is that more efficient?
Mostly to polyfill around the fact that basic DOM manipulation is garbage and the default UI widgets are ugly.
^Who cares about 2 MB with 5G? That's exactly how we get into current shitty situation. Network is fast, memory is cheap so let's do whatever we want. In 2004, I only had 512MB memory on my computer, it run pretty smoothly. But now I got 64GB of memory and I run of memory now and then. I'm wondering how soon I'm going to put 1TB of RAM into my computer, not going to be too far away I reckon.
What you describe was a huge pain, most page forced us to download 15MB of data and load it in memory just to display 2 small images and some text.
A format like PDF would have been much better because we could have read enough byte until we are able to render the visible part of the document.
But instead we had to download and execute 30 javascript files.
> it has to download 300 resources from 12 different servers,
> including 2 megabytes of minified javascript?
This is "best practice" that is used widely to work around HTTP 1.1 Head of Line Blocking [1]. HTTP/2 and to a greater degree HTTP/3 (as it also alleviates TCP head of line blocking) are fixing the underlying problem making the former best practice an anti pattern (but only for those users able to use modern HTTP implementations).
[1] https://en.wikipedia.org/wiki/Head-of-line_blocking#In_HTTP
Fortunately, Reader Mode is available in a few user-friendly browsers.
If it's just a normal non-interactive website then it shouldn't have any javascript in the first place.
For example:
* If you have file A requiring file B requiring file C requiring file D you still get a chain of requests.
* Compression works a lot better the more data you throw at it, especially similar data which is usually the case for JS/CSS. SDHC was supposed to solve this, but it was never adopted by any major players besides linkedin (and google itself).
There is an argument for better congestion control. Which is funny coming from Google where they control the kernel on the server and the clients. Yes, it takes more time to deploy better congestion control if it needs kernel updates, but if it's really important to update congestion control without kernel updates, you could make interfaces for usermode congestion control.
Ads.
1) Load page through javascript-enabled browser (headless chrome, etc), and record resources accessed from the same server;
2) Save this list somewhere, where server can read it, keyed by URL;
3) When user requests same URL, push resources from the list.
The alternative (Service worker/<meta> preferch/preload tags) on other hand are much easier to handle although have one extra round trip. Because they are just text files that you need to upload to the server.
And the idea is that client would already have those files by the time browsers parses `<head>` it already has them all. Except...browser cache exists, so who cares.