Improving Performance with HTTP Streaming
medium.com
medium.com
The main question I had was how that’s changed with HTTP/2 in terms of being able to send something like RST_STREAM to trigger client side error handlers on e.g. a fetch promise so the normal failure paths would be triggered.
The other reason things like this became moot for many sites was that fronting caches became pervasive. For things like product pages, the content would be served by an intermediary as quickly as the network could deliver it so the benefits mentioned about seeing things like script or style tags faster are less significant outside of dynamic or uncommon pages, which people tend to see only after they’ve cached core static assets. At that point, the big performance wins would come from shipping less code and making it more efficient but they’re using React so that ship has sailed.
Those qualifiers were what I was thinking about - in most cases where this is relevant you don’t have the content length in advance, and transfer encoding chunked might not trigger this if the fault aligns with chunks or it’s being done by some middle layer.
They should get rid of the popups before starting other optimizations.
The real alternative is to not track at all. It's compliant, easy to implement, performant (the topic of the post) and does the right thing
No idea what that would do to any business value Airbnb gets from tracking, but the idea is that the lost analytics data is outweighed by the benefit of removing new user friction.
Cookie popups and third party JS are two of the biggest culprits for slow sites next to large images/videos and general bloat. Some of these tracking scripts and cookie popups require you to put them into the head. It’s terrible.
I always suggest my clients to simply remove them. Users are retained when the site loads and navigates fast. Spying on them doesn’t help…
Very large sites actually get leverage out of some analysis and tracking, however they should be doing that inhouse.
For everyone else it’s simply not worth it except that is your business model (not our clients).
[0] it’s called optimization, but it’s mostly just measuring, cleaning up, moving to standard practices and deleting alot of code. The times I get to optimize queries and algorithms are fun, impactful but rare. Most sites just suck because they are bloated and unpolished.
Perhaps the advice is good, but it would be like taking budgeting advice from a guy who's always borrowing money from you.
The truth is that most people do not care (and they're probably right, as they're not heavily dependant on performance).
For the past year or so I've started thinking most OSS maintainers should just pick bolder defaults so we can push the web forward more easily.
This. For all simple cases it doesn't matter. For many not-so-simple cases the buffering is an advantage – it might impose an imperceptible delay on receiving the first data back from the initial request but make that request overall more efficient without significantly affecting anything else. This is why it has become the default in most cases. Only if your page/app is fairly complex (by necessity or by bad design) do you even need to think about buffering being a problem.
> so we can push the web forward
Not buffering at all is a step backwards in many instances. By all means make this the default in your setup instructions (or docker images etc) for your app/site, but it shouldn't be the default for, say, nginx packages in distros that affect a great many other projects.
What this is doing is selective buffering anyway: the buffer is still there, but it is being explicitly flushed at key points. Sometimes this is circumvented by parts further down the chain which you can't control in your code, hence they change the buffering behaviour in nginx. This could be even further beyond your control, perhaps imposed by some proxy that your users sit behind, so keep in mind that this trick does not universally work.
One trick I've used for internal utilities (that is a bad idea for more public systems) is to throw out a string of difficult-to-compress content in a comment before flushing. This breaks through small buffers elsewhere in the chain without needing to reconfigure the relevant parts, trading a bit of bandwidth for an apparent latency gain. It is useful for utility scripts that return a lot of data that is then presented in a fancy JS table or graph, send the initial page layout with a holding area saying “loading…”, flush, then send the data. This saves the extra request to get the data while letting the initial UI render while you wait.
My guess is debugging this would be more complicated than other approaches. It's very nice they got around slow backend queries, but I wonder if this same level of effort could be applied to fix that problem too.
As they mention in the article, the main goal of streaming your response is to deliver the <head> tag to the browser sooner, so that it can start downloading assets sooner.
However it has many downsides like explained in the article (can't change response code or headers once they are sent etc), which requires many hacks.
103 early hints allow to asynchronously send "provisional" headers, including `Link rel=preload` headers, which the browser can use to start downloading assets, but without restricting in any way the final response rendering.
That's basically the same goal than HTTP2 Server Push (which was deprecated), but much simpler to implement (works with HTTP/1), and allow the browser not to download resources it already have in cache.
ref: https://developer.mozilla.org/en-US/docs/Web/HTTP/Status/103
Implementing this is a pain in the ass, any new intermediary (haproxy, nginx, ...) introduces a potential new source of buffering or IO loop that attempts to helpfully batch up your writes. But when the technique is implemented, it is legitimately an amazing way to reduce latency at the client.
I use this technique in an app where an opening chunk is used to ensure the browser has begun fetching/executing the app JS bundle before a list response is rendered. After the list is rendered (incrementally as each chunk becomes available) it is again used to execute a slow summary query without having to break out a separate request. All that fits in a single HTTP response body, no additional latency is introduced by making separate API requests for the list or for the counts. Because of the initial chunk causing the JS to execute, it is also possible for the app to display certain dialogs (depending on query string parameters) that are immediately ready for user input even before the subsequent responses have finished rendering or downloading. In the case of this specific app, the technique means some dialogs are usable 300ms earlier than otherwise
Another issue not touched on in the post is how all this newly unbuffered incrementally-arriving data interacts with visual rendering on the client, there is potential for a lot of flicker and burned CPU continually re-rendering UI elements.
It sounds like Airbnb are disabling compression to make their implementation work. So long as CPU isn't a problem, you can also implement compression in the backend, and flush the compressor each time some useful unit of work is produced (some task completes, or some amount of data e.g. aligned to the size of an Ethernet MTU is available for writing). Assuming all the stars align, this should produce compressed output in a form the browser can act on every time it manages to read any amount of data from the backend.
All this effort seems like nonsense until finding you have a perfectly functional app on a crappy airport wifi network where it was otherwise impossible to even get Google to load
I think it saves the same main problem which is to trigger assets download early.
However you are right streaming the response has a few extra advantages.
I still think getting most of the results with a tiny fraction of the effort is a better solution, but your mileage may vary.
I think, nowadays the majority of web apps react apps. The html generation logic is at the front end and the front end only does rest API calls. So this kind of optimization is not very useful.
We had this patent https://patents.google.com/patent/US20150012614A1
It's absolutely useful. If you wait for the HTML page to load, to fetch the script, to start the react app, to have the react app make the rest/graphql calls your perf will be awful.
You want to fetch and stream in the response to your query part of your document response. Me and my coworkers gave a talk on this and it's used on facebook.com: https://www.youtube.com/watch?v=WxPtYJRjLL0
In some situations you even want to stream the response within your query response within your document.
[1]: https://techlab.bol.com/en/blog/our-ride-to-peak-season-fron...
It's evolved since the full migrate to a react app but the high level concept hasn't changed.
When I see posts like these, I become even more afraid of what the web is becoming. I don't want to have any website fetching itself piecemeal, freezing my browser window for ~5 seconds while the browser logic scrambles to reconstitute content from a million pieces.
The reason people moved away from it is error handling: it was basically a cliche that you’d see an error message halfway through a page which had started out fine because there’s no way to go back and retract the HTTP 200 & start of the page which you had just sent.
The negatives you mentioned are already happening but that’s due to the widespread JavaScript culture of trying to put as much in the client side code as possible while not measuring performance except by having developers ask whether it seems fast enough on their M2 MacBook Pros with fiber connections.
ASP (and I assume the others, my memory of them is more hazy) could buffer but it wasn't the default. You could control when the buffer was flushed if needed, to push initial content while something larger was being produced, heading towards the same compromise documented here (selective buffering) but from the other direction.
I expect AirBnB is already unpleasant on such a machine, this won't make it any worse!
Blazor has this inbuilt now in .NET 8 preview 4: https://devblogs.microsoft.com/dotnet/asp-net-core-updates-i...