300ms Faster: Reducing Wikipedia's total blocking time
nray.dev
nray.dev
But if I'm logged in its much slower - there's perhaps a second or so lag on every page view. Presumably this is because there's a cache or fast-path for pre-rendered pages, which can't be used when logged in?
If you’re logged, a number of things have to be recomputed before the page can be rendered.
I run mediawiki at home, the difference is even more stark (i have a small home server).
Logged in users still get cached pages, but the cache is only part of the page, its not as fast a cache, the servers are not geolocated (they are in usa only for main data center), and depending on user prefs you may be more likely to get a cache miss even on the partial cache.
Also, Firefox container tabs might be a nice solution: normally browse in a container where I'm logged out, then to edit, change the tab to my Wikipedia container where I'm logged in.
The main things that are different:
* the user links (shows what user you are logged in)
* whether you have a new message notification
* your skin preferences (this will totally change the entire html)
* your language preference (normally this only affects the UI not page content, but some pages it affects page content)
* other misc preferences can affect wikipage content (e.g thumbnail size), although there has been a general effort to avoid adding new ones and remove existing ones that arent used.
* some stuff makes assumptions about logged in users not being varnish (e.g. how csrf tokens work for js based actions), but most of those can be changed if push came to shove.
That said there might be a reasonable argument that much of this isnt needed and it would be a net benefit to kill most of those features.
I debated whether to go with Astro or Next.js for my new blog, but because I've not had experience with the new React server components, I decided to try Next.js. I've had a goal of making it have the least bloat possible, so it is kinda interesting to compare these side by side now.
I picked random two posts that have similar amount/type of content:
* https://www.nray.dev/blog/using-media-queries-in-javascript/ [total requests 13; JavaScript 2.3 kB; Lighthouse mobile performance 100]
* https://ray.run/blog/api-testing-using-playwright [total requests 19; JavaScript 172 kB; Lighthouse mobile performance 100]
It looks like despite predominantly using only RSCs, Next.js is loading the entire React bundle, which is a bummer. I suspect it is because I have one 'client side' component.
One other thing I noticed, you don't have /robots.txt or /rss.xml. Some of us still use RSS. Would greatly appreciate being able to follow your blog!
As projects scale this matters more. For example, hard page transitions will accrue the cost of re-initializing the entire application’s scripts for every click.
another low-hanging fruit for optimisation that nray.dev is not using, is the new inline-stylesheets option: https://docs.astro.build/en/reference/configuration-referenc...
using the 'auto' option usually reduces it to a single global sheet and page-specific inlined css. especially if you are writing a lot of short style blocks in your components, this is very very handy to reduce the amount of served css files.
When you are a small team, it is easy to create a performant site because you work in a tight circle with shared conventions, rules, etc. and it is always the same people that work on the project.
When you are a large organization (think Facebook blog), there will be hundreds of engineers/designers/copywriters/... that are exposed to the codebase over the course of many years. Each will need to make a 'quick change' that over time compounds into the monstrosities that you are referring to.
The best you can do is add automated processes into CI/CD that prevent shipping anything that does not meet certain criteria. This might hold for a while... but as soon something borks the shipping velocity and some higher-up starts asking for names, those checks will be removed "temporarily" to unblock what that's burning.
As an engineer who takes pride in developing accessible and performant software, this killed the drive/joy for me of working for large orgs.
After dealing with this in a previous job I've found great joy in working on legacy code - when a project has hit rock-bottom it can only get better.
If the JS files are set to load asynchronously, the initial load should be almost equally as fast, and React should load in the background. Afterwards, additional navigation should be near instant.
Look for a tag like this: <script id="__NEXT_DATA__" type="application/json">
https://github.com/prettydiff/wisdom/blob/master/performance...
I would love to hear a bit more details about query selector performance. Could you maybe provide some details or a pointer where this was tested?
Also about websockets vs HTTP calls: HTTP calls do not open separate socket connection even when you're using HTTP/1.1 with keep alive, with HTTP/2 it's a default way of working, so should not be a problem. I guess something else is a culprit here?
As for HTTP the issue is two-fold. The first thing is that HTTP imposes a round trip: request and response. That means the connection traffic starts with a request, waits for the application at the remote end to do what it needs, the application at the remote end sends a response whether you want it or not. Secondly, in the case with keep alive all requests on that socket must wait for the prior round trip to fulfill before sending the next request.
The primary advantage of HTTP 2 over HTTP 1 is more massively parallel requests, which greatly flatten the waterfall chart of asset requests and thus speeds page load time. In HTTP 1.1 there is typically a maximum of 8 parallel requests.
The advantage of Web Sockets is two fold. First, there is no round trip. Each message is fire and forget. This allows messages to be sent more rapidly without waiting (with a caveat). Secondly, Web Sockets is completely bidirectional. The local end and distant end can send messages independently of each other or other timing.
The big limitation of Web Sockets is message processing speed. In the Web Socket library I wrote for Node.js performance is memory bound and I can send messages as fast as 470,000 per second on my laptop with DDR4 memory per socket. The problem is not sending, but receiving. Firefox tends to separate frame headers from frame body so you have to account for that. If messages come in too rapidly Node.js will concatenate multiple messages together, so you have to account for that. In TLS the maximum message frame size is 16384 bytes irrespective of the actual message size, so you have to account for that too. There are a few other scenarios to be aware. This processing takes time to parse the binary frame headers and examine the message body length. Even still I can process up to 9000 messages per second on the same machine.
The greatest limitation with web sockets though is message queuing. I solved for this on the Node side because there is an event, the drain event, to account for a message leaving the buffer associated with the socket on the local end. As a result I know when to send a message immediately or queue for sending after the current messages completely leave the buffer. Without a queue any message that comes into a socket before a prior message drains from the buffer will overwrite the prior. On the browser side there is no queue and no event to indicate when to send from a queue, so you just have to manually slow things down with a timer.
Wikipedia, like many web sites, makes it really easy for mobile users to get redirected to a mobile-specific version running on a mobile-specific domain.
The problem is that this mobile-specific version is not good for browsing on a full computer. It's not even great on a tablet. But there's no easy way to switch back when I've been sent a mobile-specific link other than editing the link by hand. Mobile links end up everywhere thanks to phone users just copy/pasting whatever they have, and desktop users suffer as a result.
Please, anyone who develops web sites, stop doing the mobile-specific URL nonsense. Make your pages adaptable to just work properly on all target devices.
If you insist on doing things this way with mobile-specific URLs, at least make it equally easy for desktop users to get off the mobile URLs as you make it for phone users to get on them.
They asked for the mobile version - but they're on desktop.
Do you serve what the asked for? Or what you think they should want?
What if they actually want the mobile version - if you always send desktop to desktop, then they can't get mobile...
The simple answer is that you don't encode device specific intent in the URI, you put in in style sheets where it belongs.
m.wikipedia isn’t the default, so it’s different from regular wikipedia.com redirecting a mobile user to the mobile version.
It’s easier for them to go to m.wikipedia.com than to change their user agent, especially if they aren’t that technical.
Note also that WP isn't alone in their choice. Facebook, another website with a large non-Western user base, also maintains the same behavior. Go to https://m.facebook.com, you won't be redirected to https://facebook.com .
There are tradeoffs either way, and no matter what WP does, they'll be making some users unhappy. Not all of these users fit the same profile, either. Wikipedia has a global user base, so what's best for most American sites isn't necessarily what's best for Wikipedia. Ultimately, it's not the case that one option is right, and the other is wrong.
- Desktop: 388kB (https://en.wikipedia.org/wiki/Rickrolling)
- Mobile: 221kB (https://en.m.wikipedia.org/wiki/Rickrolling)
I love user customization but this is extremely niche.
You rely on saving the user display preference as a cookie that persists across sessions.
They declare their preference on a device a single time on the site and that's remembered. None of this has anything to do with which URL is being used.
(Not to mention that following a link on a page like "en.m.wikipedia.org" isn't even user intent, it's author intent at most. But usually not even that, because the author just copied the link without even thinking whether it had an "m." in it or not.)
"Remove .m. subdomain, serve mobile and desktop variants through the same URL" https://phabricator.wikimedia.org/T214998
It would be considered a non-tracking essential site function value too, so you wouldn't need to beg permission (contrary to what people who want us to be against privacy legislation will claim), and the site is probably already asking for permission anyway for other reasons so even that point is moot.
Or are you suggesting mainstream browsers are blocking same-origin session-level cookies by default now? I'm not aware of any. And if you have a browser that is blocking such things, the worst that will happen is the current behaviour (repeated mis-guesses because the preference isn't stored) continues.
Things you touch regularly should be fine.
And apparently it only affects mass local storage, not cookies which are most often used for season management (so you might stay logged in but the app need to reset data previously called in local storage).
That Safari does this is useful information that I may need to warn users of one of my projects about, as it means intentionally offline data has a much shorter expiry date than on other platforms.
I'm not defending Safari's policy, by the way, just describing. I think it sucks, and a conspiracy theorist might note how it favors native apps over web apps.
It is a work-around improving the UX on the second and subsequent requests, not a fix for the root cause.
Fixing this might be a good way to reduce bandwidth.
On Wikipedia's mobile view, collapsed sections are super useful (as the table of contents is not visible via the sidebar) and media viewer makes it possible to view details of an image/thumbnail without navigating away from the page.
Yes and no. On the hand some sort of table of contents is useful (but note that you could also just display it inline, the way it used to be done in previous desktop skins), on the other hand those collapsed sections break scroll position restoring when reloading the page after somebody (your browser or directly the OS) kicked it out of your RAM. This is because your (absolute) scroll position depends on which sections where expanded and which collapsed, and that information gets lost when the page reloads – all sections end up collapsed again and the scroll position your browser remembered no longer makes sense.
(There is some sort of feature that tries to restore the section state, but a) it only works within the current session, but not if the OS low memory killer kicked the whole browser out of your phone's RAM and b) even when it does work, it runs too late in relation to the browser attempting to restore the previous scroll position.)
So now that the mobile Wikipedia's full JavaScript no longer runs on e.g. older Firefoxes (e.g. one of the last pre-Webextension versions), the lack of a TOC is somewhat annoying, but other than that, somewhat ironically my browsing experience has become much, much better now that my browser can finally reliably restore my previous scroll position because now all sections are permanently expanded.
But I guess it's not there yet.
There's this pattern on HN: people value a feature as having 0 utility and then become annoyed that someone has paid time/performance/money for them. Well duh, if you discount the value of something to 0, it will always be a bad idea. But you're never going to understand why people are paying for it if you write off their values.
At my last job there were countless pieces of UX to make things smoother, more responsive, better controlled by keyboard or voice reader, etc.. that required JS. It was not possible to make our site as good as possible with CSS, and it certainly wasn't worth the tradeoffs of loading a big faster (not that it couldn't have had it's loading time improved--just, cutting JS was a nonstarter).
The js is unnecessary if you can achieve the same result with plain css.
But to play devil's advocate: just because you can, doesn't mean you should.
In many scenarios I'd argue CSS would require more bandwidth. It can get quite verbose.
Wikipedia is consulted every day by a lot of people, i guess that a large number of those people are running older browsers.
<dialog> has only been available in Safari for about a year.
Wikpedia is one of the most popular sites on the internet. It needs to be as compatible as possible so that means using JS.
considering the dominance of few browsers (chrome , safari on iOS) will most users notice any difference? The first one (with that UA) to the site with a new build will have they cache key warmed up ?
All to get rid of a tiny bit of JS.
Plenty of sites aspire to be JS free for a variety of reasons it is worthy goal
Also, if you want to further speed up your site, just like you said, the fastest way to speed up the site is to delete JavaScript, get rid of jQuery.
<details> is completely supported.
But if I can do it with all html, that's better than a function to add / remove a css class and document.onclick = a function to find which recipie you clicked on and adjust the css class.
I assume the author of the blog post just wanted to optimize the current situation, not completely change how these features work (which would most probably be a much more elaborate change).
Same for dropping jQuery - that will probably be a few weeks or months of work in a codebase the size of Wikipedia/Mediawiki.
Only if you want fancy animations. I sometimes do, but I think wikipedia can do without(and they don't) and use <details>
https://developer.mozilla.org/en-US/docs/Web/HTML/Element/de...
And media viewer I would naturally do with js as well, but I am certain you can also do it easily with CSS. (Many ways come to mind)
Good article about some issues (with a link at the top to his previous article about them). https://www.scottohara.me/blog/2022/09/12/details-summary.ht...
[0] https://caniuse.com/?search=details
[1] https://www.mediawiki.org/wiki/Compatibility#Browser_support...
(Anyway, if you are using CSS + the relevant semantic HTML elements, it can be more accessible, not less, because you are expressing priority and emphasis, so they can skip over collapsed stuff. Although I have my doubts whether screenreaders etc make any good use of it, given that they apparently still do not do basic things like stripping soft hyphens.)
The fact they’re still using JQuery, probably for similar compatibility reasons is good evidence of that.
Now there are ways to use polyfills that only load when necessary but just about everything is very difficult at wikipedias scale. We can’t solve their problems from an armchair.
I've been developing websites professionally since 1996. HTML/CSS/JS and SQL.
I am still amazed there is the crowd out there who is "anti-JavaScript". They run NoScript to only allow it in certain places, etc.
It's 2023, a small amount of JavaScript on a page isn't going to hurt anyone and will (hopefully) improve the UX.
For the record, the last site I deployed as a personal project had 0 JavaScript. It was all statically generated HTML on the server in C#/Sqlite that was pushed to Github Pages. So I get it, it's not necessary.
For my little personal site. I'm also the senior lead on an enterprise Angular project.
JavaScript is fine, it's not going anywhere.
And yes, there are way too many React (specifically) websites that don't even need a framework at all, but it's become the go-to. That annoys me too. But some JavaScript in 2023 is fine.
It's a signal to noise ratio thing. There's a small reason to have JS switched on and a large reason to block it. I'd love it if the reason to have JS switched on was even smaller, and that is in fact the recommendation (that JS should enhance an otherwise functional site). I enable JS on sites where I regard the reward as worth it, but most of the time it just isn't, and I don't trust all the rubbish that gets included in the average site.
The worst offender is substack combined with many comments. I tried opening a Scott Alexander blog post in mobile Chrome on my aging mid range phone:
https://astralcodexten.substack.com/p/mr-tries-the-safe-unce...
It's not that it's rendering slowly, it apparently doesn't finish rendering at all. At first a few paragraphs render, then everything disappears. I basically can't read articles on this blog.
This website says the above has a total blocking time of 15.7 seconds:
https://gtmetrix.com/reports/astralcodexten.substack.com/3y8...
That's on more modern hardware I assume.
I also tried https://pagespeed.web.dev/ but unfortunately it timed out before it finished. Great.
This doesn't load fully in Safari on an M1 Pro either. I scroll down and the content just ends. It hijacks the scroll bar too.
EDIT: _not_ fine on M2 Pro: Chrome works, Safari doesn't.
As long as I'm here...so much ink has been spilled on Safari and it's a bitter debate.
I really, really, really wanted to make Safari work for a web app and had to give up after months of effort.
So many workarounds and oddities that simply aren't present in other browsers. After 15 years orbiting Apple dev...I've finally acquiesced to the viewpoint that carefully filing & maintaining a backlog of bug reports, checking for fixes, and providing a broken experience isn't worth it to me or users
Kind of silly to debate over such things. On desktop, the built in browser's only job is to let you download a better browser.
I use Firefox mobile, if I scroll too fast on substack I sometimes have to let it sit 10+s before the blankness renders
(Annoyingly, they don't show any general contract opportunity on their website. Luckily in the Google Play Store they have to supply a support email address. Got it from there.)
What's happening here is the browser did a first paint, and then the javascript started eating up all the CPU. You can scroll because of asynchronous pan/zoom, but when you get to the edge of the rendered region, it asks to render more and it never gets rendered because the page is stuck executing javascript. Recording: https://share.firefox.dev/3C2DE9C
It certainly has the overall feel and appeal of being done by hand, but I'm not sure.
If it's software, does anyone know which software, what's the name of this style, etc?
OTOH if something else should happen, that event handler can call event.stopPropagation(). This stops the event from reaching the delegated event handler, so no inefficiency there.
One of the hot methodologies in the JavaScript world is event delegation, and for good reason.
A lot of that crunching is either whittling through DOM traversals.
I've heard that rule of thumb that there are 5-9 machine instructions before JMP or CALL which you can find if you search "instructions per branch cpu"
If I had a database of 4,000 links and instantiated a function event handler to each one, that would be slow. But if I could invert the problem and test if the clicked object is inside an active query resultset, that could be fast. It could be automatic.
So one guy stood up and asked why our main page is not compliant with W3C guidelines (doesn't pass HTML validator test). L&S answered that they need to shave off any unnecessary byte, so that it loads faster. That's why they didn't care about closing tags etc.
How the world and perception has changed since that time... It's just sad that one needs a huge JS framework just to build a simple website these days.
Oh and the size of the homepage was in the order of 20 KB then, IIRC.
I'd recently looked at battery usage and found that my preferred browser (Einkbro) uses ten times as much battery per hour of reading than Neoreader (the Onyx BOOX stock ebook reader).
Einkbro is a solo-developer effort (based on the FOSS browser), so I'm not sure how much energy-usage optimisation it has or is lacking. It's sufficiently pleasant to use (relative to other Android browser options) that I don't use other browsers significantly ("Neobrowser", a re-branded Chromium, and Firefox Fennec are both installed, and used occasionally). But it really sucks down juice.
That's on top of the fact that many web pages are larger (storage, memory use) than a full-length book. Typical e-pubs run ~500 kB -- 5 MB, with graphics-heavy or scanned-in books running up to 100+ MB.
I also find that reading pre-formatted, laid-out content with pagination and navigation is increasingly far more pleasant than scrolling through a Web document. HTML as presently instantiated has some pretty sharp corners and limitations.
But then I tested in Incognito, and it was 350ms for the main document.
Disabling cache doesn't make a difference. I tried logging in on Incognito and boom, fast again. Somehow the backend is 3x the speed when logged in. I'm guessing anon users have to go through some extra bot checking or something.
I'm not entirely sure this can be so axiomatic...
That can definitely be a thing. However JS is usually a much bigger issue.
- Dynamic tables which can be sorted by column. (This is a feature I'd love to see baked in to HTML browsers generally.)
- The recent site redesign offering either visible floating or collapsed page navigation.
- Previews of both articles (Wikipedia links) and citations, such that these can be viewed without having to click through to a new page, or scroll the current article.
There are a number of others, I'm sure, but these three alone are quite handy.