Page Weight Matters (2012)
blog.chriszacharias.com
blog.chriszacharias.com
We did some crazy stuff to squeeze everything in. We would literally count bytes on every change - one engineer wrote a tool that would run against your changelist demo server and output the difference in gzipped size of it. We used 'for(var i=0,e;e=arr[i++];) { ... }' as our default foreach loop because it was one character shorter than explicitly incrementing the loop counter. All HTML tags that could be left unterminated were, and all attributes that could be unquoted were. CSS classnames were manually named with 1-3 character abbreviations, with a dictionary elsewhere, to save on bytesize. I ran an experiment to see if we could use JQuery on the SRP (everything was done in raw vanilla JS), and the results were that it doubled the byte size and latency of the SRP, so that was a complete non-starter. At one point I had to do a CSS transition on an element that didn't exist in the HTML, because it was too heavy and so we had to pull it over via AJAX, so I had to do all sorts of crazy contortions to predict the height and position of revealed elements before the code for them actually existed on the client.
A lot of these convolutions should've been done by compiler, and indeed, a lot were moved to one when we got an HTML-aware templating language. But it gave me a real appreciation for how to write tight, efficient code under constraints - real engineering, not just slapping libraries together.
Alas, when I left the SRP was about 350K, which is atrocious. It looks like it's since been whittled down under 100K, but I still sometimes yearn for the era when Google loaded instantaneously.
It blows my mind this is not advertised and not the default.
- It’s non-personalized – meaning you get less filter-bubble effect, which is useful when researching topics where you want to see opinions that oppose your own
- It has less tracking applied.
Thanks very much for the tip off.
https://developers.google.com/custom-search/docs/overview?hl...
Way back when Google first got VC, before they invented AdWords and got profitable, they explored a lot of custom business partner opportunities. So there were a number of specialized Google search endpoints like /linux, /bsd, /unclesam (search over Federal government websites). These were maintained until about 2013 - they go through a different rendering path than the one that serves mainstream Google Search. /linux, /bsd, /unclesam, and a few others were decomissioned in 2013, soon before I left Google, but it looks like /custom got a reprieve for whatever reason. It's actually been deprecated by the Custom Search Engine functionality linked above, but nobody's gotten around to removing it.
Incidentally, you can also get a similar page (slightly different, its last update was around 2007 and it shows no ads) by setting your user agent to "Mozilla 3.0".
There was a day when I just wrote a small script to open all sites mentioned in Google’s robots.txt (which did not return 404) in a separate tab. That’s how I found it.
As far as I know, some partners actually still embed /custom (I’ve only seen it at a newspaper a few months ago), so removing it might be problematic.
> As far as I know, some partners actually still embed
> /custom (I’ve only seen it at a newspaper a few months
> ago), so removing it might be problematic.
An ISP local to here does. [0]http://centurylink.net/search/index.php?context=search&tab=W...
I see it every so often if a domain doesn't get resolved - they'll hijack the request and show you their own page. I was pretty surprised that that still happens.
Not to take away from your effort, but that sounds like something a minifier should be able to do, right?
Edit: Yep, you said it below:
> A lot of these convolutions should've been done by compiler, and indeed, a lot were moved to one when we got an HTML-aware templating language.
Have you ever seen something completely insane and everyone around doesn't seem to recognize how awful it really is. That is the web of today. 60-80 requests? 1MB+ single pages?
Your functionality, I don't care if its Facebook - does not need that much. It is not necessary. When broadband came on the scene, everyone started to ignore it, just like GBs of memory made people forget about conservation.
The fact that there isn't a daily drumbeat about how bloated, how needlessly complex, how ridicuous most of the world's web appliactions of today really are - baffles me.
The author knew the article would get good reception here, because of how much people complain about page latency, and I knew that page latency is a thing that's important to more people than just me, for the same reason. (My clients mostly don't seem to care, for some reason, but I bet their customers do).
So yeah, having these public conversations can make a difference. We just have to stay positive and constructive.
Thanks for reminding me of that; I was previously sitting here thinking, "oh geez, here we go again!".
Here, I'll beat a drum a little. Maybe it will inspire somebody.
I just wrote this tiny text-rendering engine, mostly yesterday at lunch. On one core of my laptop, it seems able to render 60 megabytes per second of text into pixels in a small proportional pixel font, with greedy-algorithm word wrap. That means it should be able to render all the comments on this Hacker News comment page in 500 microseconds. (I haven't yet written box-model layout for it yet, but I think that will take less time to run, just because there are so many fewer layout boxes than there are pixel slices of glyphs.) http://canonical.org/~kragen/sw/dev3/propfont.c
The executable, including the font, is a bit under 6 kilobytes, or 3 kilobytes gzipped.
IMO pretty clear why, winner-takes-most economics emphasizes dev speed over quality, and then we need to maintain compat with it
the way I read OP was the abandon rate was ~100% previously in areas with high latency/low bandwidth, and his experiment had a non-zero abandon rate that was shrinking as word spread. Perhaps I need to read it more carefully.
I'm in Bangalore and being forced to use mobile 3G for now. The data charges are atrocious and each page load saps an MB.
Well there is, or at least there used to be. The reason native apps in mobile became so popular is exactly because web sites are bloated. It might not matter in the PC but in a device running on batteries it matters a lot.
The real problem of web development is JavaScript. It’s a relic of the past that hasn’t caught up with times and we end up with half-assed hacks that don’t address the real problem. We need something faster and way more elegant.
I absolutely disagree. JS is plenty fast for >99% of the things being done with it. The DOM, and manipulations of it, is the slow part.
Here are a couple of real culprits:
* Advertising / analytics / social sharing companies. They deliver a boatload of code that does very little for the end-user.
* REST and HTTP 1 in combination. A page needs many different types of data. REST makes us send multiple requests for different kinds of data, often resulting with those 60-80 requests mentioned above. We could be sending a single request for some of them (e.g. GraphQL) or we could get HTTP 2 multiplexing and server push which completely fix that problem.
* JSON. Its simple but woefully inadequate. Has no built in mechanisms for encoding multiple references to the same object or for cyclic references. Want to download all the comments with the user info for every user? You have a choice between repeating the user info per comment, requesting the comments first then the users by id (two requests) or using a serialisation mechanism that supports object references (isn't pure JSON)
* The DOM. Its slow and full of old cruft that takes up memory and increases execution time.
That's not quite fair. They subsidize the content for the end-user. Perhaps that's a crappy status quo, but in many cases without the advertising and analytics the content wouldn't exist in the first place.
[1]: Just copied them from coldpie's news site screenshot comment. That was about 1/2 of them, there were many more.
Do you ever wonder why the landscape became the way it did?
A significant problem is that content costs a certain amount of money to be produced, and web content is unable to command those prices.
Ad fraud is a big part of it, and some of the companies in the best position to solve it (like Google) are benefitting so handsomely from ad fraud that I can't imagine them stopping it.
Ad blockers will hurt legitimate content producers, but because people defrauding advertisers don't use ad blockers, they'll continue to make money.
This is perhaps the most ridiculous comment I have read this afternoon.
Wow. Well, I'm glad you have your feet anchored down here in reality and aren't exaggerating at all.
1. Advertising is the problem.
It's creating technical problems. It's creating UI/UX problems. It's creating gobs of crap content. It's creating massive privacy intrusions and security risks. And for what? Buzzfeed?
2. Ultimately, the problem is the business model for compensating informational goods. Absent some alternative mechanism (broadband tax, federal income tax applied to creative works), I don't see this changing much.
https://www.reddit.com/r/dredmorbius/search?q=broadband+tax&...
3. Micropayments aren't the solution.
I'd like a system under which creators of quality content would be equitably and fairly compensated. The present system fails this.
1. Advertising is the problem.
It's creating technical problems. It's creating UI/UX problems. It's creating gobs of crap content. It's creating massive privacy intrusions and security risks. And for what? Buzzfeed?
2. Ultimately, the problem is the business model for compensating informational goods. Absent some alternative mechanism (broadband tax, federal income tax applied to creative works), I don't see this changing much.
https://www.reddit.com/r/dredmorbius/search?q=broadband+tax&...
3. Micropayments aren't the solution.
Or you send it like
{
"users": [
{
"id": "de305d54-75b4-431b-adb2-eb6b9e546014",
"name": "Max Mustermann",
"image": "https://news.ycombinator.com/y18.gif"
},
...
],
"comments": [
{
"user": "de305d54-75b4-431b-adb2-eb6b9e546014",
"time": "2015-10-15T18:51:04.782Z",
"text": "Or you can do this"
},
...
]
}
Which is a standard way to do it that works here, too. Because you can have references in JSON, you just have to do them yourself.But it still takes more power for your server to go through gzip every time. And it will take more RAM on the client to store those objects.
But not exactly pure JSON. Client-side, you need to attach methods (or getters) that fetch the user object to the comment. I suppose you could just attach get(reference) that takes this[reference + '_id'] and looks it up inside the `result[reference]`. m:n relations will be harder though.
Otherwise you can't e.g. simply pass each comment to e.g.a react "Comment" component that renders it. You would also have to pass the user list (dictionary) to it.
response.comments.forEach(comment => comment.author = response.user[comment.author]);
Preferably, even, you could do it during rendering so it can be garbage collected afterwards.I'm not going to say that the DOM is wonderful but … have you actually measured this to be a significant problem? Almost every time I see claims that the DOM is slow a benchmark shows that the actual problem is a framework which was marketed as having so much magic performance pixie dust that the developer never actually profiled it.
vs
Initialising about 300K invisible DOM nodes containing 3 elements, each with 3-character strings is ~15 times slower than initialising an array of 300K sub-arrays containing 3 elements, each being a 3-character string.
Additionally, paging through the created nodes 2K at a time by updating their style is just as slow as recreating them (the console.time results say something different, but the repaint times are pretty much the same and this is noticeable on slower computers or mobile.)
Thats a single benchmark that raises a few questions at best. But I think React put the final nail in the coffin: their VDOM is implemented entirely in JavaScript, and yet its still often faster to throw away, recreate, diff entire VDOM trees and apply a minimal set of changes, rather than do those modifications directly on the DOM nodes...
Similarly, React is significantly slower – usually integer multiples – than using the DOM. The only times where it's faster are cases where the non-React code is inefficiently updating a huge structure using something like innerHTML (which is much slower than using the DOM) and React's diff algorithm is instead updating elements directly. If you're using just the DOM (get*, appending/updating text nodes instead of using innerHTML, etc.) there's no way for React to be faster because it has to do that work as well and DOM + scripting overhead is always going to be slower than DOM (Amdahl's law).
The reason to use React is because in many cases the overhead is too low to matter and it helps you write that code much faster, avoid things like write/read layout thrashing, and do so in a way which is easier to maintain.
Also, total time to repaint seems to be equally fast whether we're recreating the entire 2K rows from scratch or changing the display property of 4K rows. That seems concerning to me.
But yes, a more fair comparsion would be to write a WebGL or Canvas based layout engine in JS. If its less than 4 times slower (its JS after all), then the DOM is bloated.
I certainly would expect that you could beat the DOM by specializing – e.g. a high-speed table renderer could make aggressive optimizations about how cells are rendered, like table-layout:fixed but more so – but the more general-purpose it becomes the more likely you'd hit similar challenges with flexibility under-cutting performance.
The most interesting direction to me are attempts to radically rethink the DOM implementation – the performance characteristics of something like https://github.com/glennw/webrender should be different in very interesting ways.
I think you'd find a much easier case made that its because:
1) discoverability and distribution through app stores,
2) access to native features that took years to be enabled in the mobile web (for example access to photos..)
3) ability to make games (access to graphics, etc -- this also kind of hampers the case being made about "bloated" sites when people seem very willing to download 100MB games)
4) access to ease-of-payment (CC info stored on people's phones that you can in app payment, vs having to type info into a web form).
So we could try to focus on clear issues the web has vs native, or I guess we could design yet another language that is theoretically faster thanks to >insert pet feature here< and hope that fixes the web.
There's a huge difference between downloading a 100 MB app once (plus occasional background updates) and waiting to load a 100 MB web game every time it falls out of cache.
Plus the web loading delay happens constantly while you're trying to use it. You wait on an app download exactly one time: when you buy/install it.
My 4G connection right now is speedtesting at 46.68 Mbps download.
An iPhone 6 benchmark I found on MacRumors can read at 6720 Mbps (840 MB/s) [1]
Given the choice, I'd take the option that's 140x faster.
[1] http://forums.macrumors.com/attachments/img_0005-png.514599/
BTW: /etc/hosts + dnsmasq, for Linux, is amazing. (dnsmasq reads /etc/hosts and will block entire domains if listed as same).
I think, not totally sure how NoScript works.
By comparison, noscript simply blocks javascript from third parties. It does include a number of anti-xss heuristics though.
It's also possible to directly address hosts by IP, though unlikely (Web protocols such as virtualhosts would fail).
I'm strongly favouring uMatrix for now. It takes some tuning, but you have fine-grained control over CSS, images, scripts, XHR, frames, and other bits, by domain or host.
Aggregators and CDNs confound things a bit (Akamai, Amazon's cloudy thing.)
I feel Jakob Nielsen used to majorly play this role amongst Web developers back in the day, but for some reason I haven't seen him come across my radar for many years now. The Web performance experts now are better than ever but I don't get the sense of the majority of the industry hanging off their every word as seemed to happen with Jakob 10+ years ago.
I'm not saying it's right. In fact, I would prefer if everyone took a look at their apps much like the author of the article. I'm just saying, this isn't a new phenomenon.
https://en.wikipedia.org/wiki/Simpson%27s_paradox
It also reminds me of the phenomenon, in customer service, whereby an increase in complaints can sometimes indicate success -- it means the product has gone from bad enough to be unnoticeable to good enough to be engaged with.
Feels bad for the engineer who spent all that time reducing the size and finding out it made YouTube much more usable across the globe. Amusingly, disabling buffering was probably some penny wise pound foolish way to save bandwidth.
After 6 months of banging my head against a wall, I realized the reason we weren't fixing page weight was because our product managers didn't care about the experience of users in poorer countries, because they didn't have any money to spend anyway. Even though we had lots of users in those countries, and even though we made a big show of how you could use this app to travel anywhere in the world.
If there's a lesson there, its that as long as cold economic calculations drive product decisions, this stuff isn't going to get any better.
I've used it to emulate what it's like on a high-latency or high-loss network. Relatively easy tool to use.
Today, I have 2 mbit and can use Netflix or Youtube just fine, but mere 4 years ago, I had 600k and, boy, that was hard. Hard as in loading youtube URL and go for a coffee.
UPDATE:
In case Duolingo developers are listening, please test your site on high latency and very low bandwidth scenarios. I just love your site, but lessons behave too strangely when internet is bad here.
You can't even to that now. Youtube videos buffer about 1m30 of videos and stops after that :(
http://tools.pingdom.com/fpt/#!/F4VDN/https://about.me/penta...
I wonder what would happen if for example iOS decided to visually indicate page weight, kind of like how you can see which apps use the most energy.
Perhaps the simplest solution is for Google to start penalizing heavy pages, but as far as I know, page weight isn't part of their mobile-friendly criteria.
I find video backgrounds ridiculous most of the time (though they can be done really well), but that's not the real problem here - the real problem is including 200-800kb javascript code that does nothing but track your user, and often enough doesn't do it for you! (Hi Facebook!)
The real problem is using massive js frameworks for the sake of adding dynamic functionality to your site that, often enough, isn't actually worth it.
The real problem is that very often, these "features" are only as necessary as the marketing team says they are... the people who have the ability to ask "why?" and the ability to understand "why not" don't have the voice (or guts...) to do so.
The marketing team was surprised to learn how big the file was since the outsourced team "ensured" them that the new website would be lightweight. Oh well... Now the video has been reduced to a 16MB at 720p, which is still ludicrous to me, but they just really like having the video on the home page more than they want a truly lightweight page.
And to be fair, the video does not load when viewed on a mobile device, so at least there's that.
Genuinely curious: Why is this better than a mobile-friendly site designed specifically with the constraints of a mobile device in mind?
Also, if you care about page and weight and optimization, your site will be light everywhere, and the criticism of shoehorning a big bloated desktop site into a phone won't apply. This is not that difficult to achieve.
But if you couldn't even load the pages to browse through the videos, that'd pretty significantly reduce your watching even further.
Liken it perhaps to a library where you could only get 1 book per week. Well that's not much, but at least you can get 1 book. But if you go to the library you have to wait in line for an hour, and after examining one book section of 10 books, you have to wait 5 minutes to examine the next one. Choosing the weekly book would become such a chore you'd probably not even bother to go anymore. But if you'd suddenly be able to examine the entire library with only 1 minute of total waiting, many more would be much more inclined to do so, even if they could still only borrow 1 book per week.
Beyond that in my experience, lots of requests in a slow connection often fails before completing fully. I don't know why exactly, I've never been a network engineer. While streaming in a slow connection eventually gets there. So it's entirely possible that even browsing normal pages was failing entirely before for some people, even though (if they ever did load the page and choose a video), the video download worked fine albeit at a very slow pace.
My favorite example of this is YouTube repair tutorials. If I need to get something done, these can be irreplaceable. If I had a slow connection, but my car was up on the jack stand, I'd just have to wait.
This was my first experience with the web in high school, with a modem, on AOL. I loved it.
My friends sent me e-mail with a youtube link and I literally waited one hour or so to view that. Just let it loading while reading something else.
So, the trouble here is not me having low bandwidth, but having connections on broadband places.
And even people with no connections in the first world want to see video clips online.
That's why I think AJAX, web manifest [1], indexedDB, localStorage, etc. need to be leveraged much more. Imagine most of your app loading without making a single request, except for the latest front page JSON, or the latest . You have a bunch already in indexedDB so you just ask the server "hey, what's new after ID X or timestamp T?"
So your two minutes just became a couple milliseconds (or whatever your disk latency happens to be), and the data loads shortly thereafter, assuming there's not much new data to send back. And if you don't need any new resources, you only had to make a single request.
In fact, if I knew of a way to decrease the size of the first load, at the expense of making the second load take longer, I'd probably do it.
http://emberjs.com/blog/2014/12/22/inside-fastboot-the-road-...
Also, isn't adding bloat "more work"? And just like in real life, I think losing bloat is often harder than gaining it. Why not worry about carousels or endless scrolling or video background when even just one person misses that stuff? Where are all the highly successful websites that started to reduce bloat after they got off the ground? It's not a rhetorical question or sarcasm, I am interested, but I honestly can't think of even one example, it doesn't really mesh with my own (admittedly rather pedestrian) experience. A site starts with a blank design doc, and empty file and a white screen, and adding something and later removing it isn't easier than simply not adding it in the first place.
lol, I'm in China and all what you're discussing just does not exists here. What you'll get is just a "Connection Reset", no matter how compact the page is.
I remember visiting microprose.com with my 14.4k modem in the mid-90s and being mad that they used so many images I had to wait for about 5-10 minutes or so. I couldn't effectively read it at home and usually ended up reading it at the library.