Why is the Internet so Slow?
blog.apnic.net
blog.apnic.net
The page this article is published on makes 87 separate network requests to 19 different domains. Even with a great deal of the content cached, it takes 7 seconds to reload over my WiFi today. This is because I'm using an ad blocker; without one, the page makes more requests to more domains and transfers more stuff, slower.
It seems likely that cutting bloat would have a much bigger impact on responsiveness than infrastructure upgrades. Cutting bloat would produce benefits even (especially) in times and places where signals or infrastructure are marginal. Cutting bloat could mean your old phone or tablet computer could browse the web tolerably instead of ending up in a landfill. Cutting bloat from a website today produces a benefit, for all visitors, worldwide, today. Improving infrastructure may be a good investment, but cutting bloat is a force multiplier.
The responsiveness of the web has been adjusted to optimize the number of eyeballs that see ads. If a site is too sluggish, people leave, but a site that's too fast is leaving money on the table. How many advertising, tracking, and affiliate domains would a web page like this connect to, if latency were cut by two thirds?
Their page has a high number of requests primarily because they do not bundle their javascript, CSS, or icons.
I see 95 requests that I see without adblocking, and about 20 of those requests are javascript coming out of their own WordPress instance. They're using 2 different full icon libraries and adding at least 5 more icons not contained in either.
Think about it. If you include a gigantic library in a compiled language, but only use 3 of its features, the rest of the library doesn't make it into your final program. Only those three features that you used do, because that's the linker's job. It takes only the features of a given library that are actually used for space efficiency reasons, and removed anything unneeded.
This is possible for JavaScript, but common language usage patterns work against this goal. If you're going to use a clientside library that's fine, but you should ship only the small fraction of it that your code actually calls into, not the entire honking library, 90% of which goes untouched by the client for every user.
Point here is, I agree with you, but this is a difficult problem to solve. You're asking a large number of frankly inexperienced programmers to deviate heavily from most JavaScript tutorials, and worse, most example code included in the documentation of most large frameworks.
The equivalent for dynamic languages is tree shaking, and is a harder or impossible problem depending on the language. When you can do something like the following in JavaScript, how do you determine what's dead or alive?
var fn = document.getElementById("function").value;
var input = document.getElementById("input").value;
var output = functions[fn](input);
document.getElementById("output").value = output;
[0] EDIT: But all the method and fields within a class are loaded, regardless of whether they are used. Which is why I called it a hybrid.Ignore if it can be detected or not, we'll pretend that the author has to certify that they followed your subset and didn't do X, Y, or Z. If they disobeyed your rules then all bets are off and the result is undefined.
The code in tracking scripts varies between "short and elegant" to "massive and slow code". There is also a piece of json where the tracker thinks it is necessary to call 10 other trackers: See for example http://ib.adnxs.com/async_usersync?cbfn=AN_async_load : AN_async_load([ {"url":"http://cm.g.doubleclick.net/pixel?google_nid=appnexus&google..., "tagtype":"img"}, {"url":"https://ib.adnxs.com/getuid?https%3A%2F%2Fcm.g.doubleclick.n..., "tagtype":"img"}, {"url":"http://c.bing.com/c.gif?anx_uid=xxxx&Red3=MSAN_pd", "tagtype":"iframe"}, {"url":"http://odr.mookie1.com/t/v2/sync?tagid=V2_4265&src.visitorId..., "tagtype":"img"}, {"url":"http://match.adsrvr.org/track/cmf/generic?ttd_pid=appnexus&t..., "tagtype":"img"}, {"url":"http://pixel.rubiconproject.com/tap.php?v=4894&nid=1986&put=..., "tagtype":"img"}, {"url":"http://sync.mathtag.com/sync/img?mt_exid=13&mt_exuid=xxxx&re..., "tagtype":"img"} , {"url":"http://p.rfihub.com/cm?in=1&pub=345&userid=xxxx", "tagtype":"img"}, {"url":"http://t.mookie1.com/rsp?dnv=[TIMESTAMP]&rurl=//ib.adnxs.com..., "tagtype":"img"}, {"url":"http://t.wayfair.com/a/vendor_sync/user?vendor_id=1&uid=xxxx..., "tagtype":"img"} ]);
The slowness that we experience is from delays caused by calling scripts that call scripts that call scripts, processing large pieces of javascript and network latencies. edit: and server latencies.
What percentage of my cellular data-consumption is specifically ads?
I need an app that tracks all data connections and tells me what percent went to which advertisers and a mechanism to block them or bill them to subsidize the cellular bill I pay for data, but that payment has a sizeable chuck allocated to my device displaying their ads.
Does anyone not recall the impetus behind the original paid cable tv model was that "you pay for the content of cable TV, thus we won't show you commercials"
The. They found out they could say "fuck you" to consumers, play commercials, hike up rates and also take in billions of government subsidies for infra upgrades that were never performed.
And companies wonder why some people want vigilante justice on such companies.
The page is 2.3kB on the wire, and the single other request it makes (for the image) adds another 70kB. Taken together, this is a two-order-of-magnitude decrease in page size, with exactly the same content.
Every byte is sacred! This webpage infuriates me.
Convert it to a 32 color 8-bit version and it looks identical. Then use ImageOptim (or like) to compress it down to less than 16 KiB.
convert -colors 16 Ilker_fast.png Ilker_tiny.png
pngcrush Ilker_tiny.png Ilker_out.png
Total page weight is now 15kB, 0.2% the size of the original page. Incidentally, that ratio is about the same as the ratio between the original page and the entirety of the game Overwatch.advpng -4 -z -i 15
http://motherfuckingwebsite.com http://bettermotherfuckingwebsite.com
In the spirit of these repeated posts, I'll repeat the typical critique. "The white background is harsh on eyes, and there is nothing wrong with images and color to help improve aesthetics."
"I would've even made this site's background a nice #EEEEEE if I wasn't so focused on keeping declarations to a lean 7 fucking lines."
http://www.webdirections.org/blog/the-website-obesity-crisis...
Using ublock origin, Ghostery on Safari with 2015 MacBook Pro 15"
For years all I've wanted is, say, 6Mbps & 5ms ping, but nearly everyone seems to focus on bandwidth.
It might lead to more ads on webpages. But it would also dramatically improve VPN and all "realtime" apps, i.e. VoIP, Skype, remote desktop, sometimes ssh, etc.
Read https://drive.google.com/file/d/0B6Xurc4m_PVsZ1lzWWoxS0pTNVE... for a general introduction about latency in Ethernet/IP networks and http://www.ieee802.org/1/files/public/docs2016/cm-chen-front... for an introduction to the "radio over ethernet business"
Even if you're not trying to send radio baseband data over ethernet networks -- there's lots of noticeable queuing going on, everywhere. It's not just the unmanaged and pathological "bufferbloat" queues (that occur when a device has a network link much faster than its other network link) -- there's little queues that are used for tasks like:
* frame aggregation
* holding onto data until a shared medium is idle
* for scheduled systems like DOCSIS or LTE, holding onto packets until an uplink grant has been given -- or holding onto packets until a station's downlink grant/timeslot becomes active
* inter-device communication (usually with ring buffers or the like)
* any processing/forwarding that isn't of the cut-through flavour usually involves queuing at least a single frame
* context switching / memory copying / interrupts
Retransmission mechanisms can introduce latency as well -- if there's no special channel for acknowledgements, they'll use up airtime, and if data needs to be retransmitted, it might induce delays for never-transmitted data that's waiting for its first transmit opportunity.
In general, unless you're actively being paid to stomp out every imaginable source of latency, there'll be little queues living everywhere in your system -- from TCP's queues right on down to the CPU's store buffers-- because queuing is the natural way to cope with subsystems that produce/consume at different rates.
BBR is a sender-only modification to TCP that basically solves the problem of large buffers.
Anything faster is just a cherry on top. Sure, files download faster, you can stream 4k, etc. But I'd say that's luxury which you could choose. Lower speeds are "enough" for a lot of people.
I can have an effective voice conversation with terrible quality audio if it's low latency. Extremely lifelike (high bitrate) audio is just uncanny (not helpful) if the latency and jitter are high.
Besides, high detail remote desktop are a small, specialised segment, relatively speaking. If you're a developer, most of the time you work with text. If you're an office worker, most of the time you work with fairly static apps like office. Sure, you can't really work with remote blender at the moment, but we're getting into improvements for the long tail now - which is great. But once we're past 10mbps, improving latency will help more people than improving bandwidth.
Right on the money. There is a deep synergy between AI and simulation. It's not just that they both run on GPUs, but simulation will become a major part of future AI. A robot needs to imagine scenarios before executing them. So they meet in the middle - gaming, VR, AI and robotics. Isn't it interesting that PacMan and other Atari games, that we used to play when we were kids, are now the playground of AI?
It's deployed in some places and working already. Citing from the webpage:
> SCION is the first clean-slate Internet architecture designed to provide route control, failure isolation, and explicit trust information for end-to-end communications.
Also it will reduce latency, increase throughput and more, as mentioned in the papers.
The project has really been getting traction in recent years, it may just become our new internet base layer.
Currently the people at ETHZ are working on verifying router hardware mathematically to guarantee SCION's properties.
They also don't implement any security headers: https://observatory.mozilla.org/analyze.html?host=scion-arch...
I don't understand what this has to do with a free internet, routing doesn't have any affect on that
We also don't take advantage of DNS enough. We should/could use it to pass along information about the client's ISP to locate a server one or two hops away from the client. DNS could easily solve the IPV4 with SRV records, but instead, we allow service location to happen at the TCP level.
The internet _could_ be fast, but it's not :/
If I had the choice, I'd take an additional image in that trade, thanks. The 'rich' 'experiences' I've been offered so far are inferior to bog-standard web pages. I'll grant you that it is friendlier to coders whose applications fit the model. Beyond that, they don't work without Javascript[1]. The "richness" is usually useless animation and similar, frequently employed because it is there, rather than actually adding any value (Anyone remember the Jquery animation explosion?). It breaks a lot of automation; see the JS comment. These are all things that are important to me.
I get that folks are fine with losing me as a user/customer, and that's their choice. But they should know why, thus my explanation.
[1] Which means I usually go elsewhere when I encounter it, because I default to leaving it off.
What you're really asking for is for the components being used, be they images, animations of frameworks, to be used with a deliberate purpose. Which I think is what everyone is asking for, and what good developers and designers strive for.
You missed my complaints about the uselessness of Angular apps without a JS interpreter. That's important, because it means they're basically worthless for consumption by nonhumans. (Sure, I can run Chrome headless, but that's an entirely different can of worms that play poorly with pipelines, not to mention an enormous, absurd runtime for what should be a trivial file transfer.)
Web automation on a personal scale is enormously useful. Angular breaks the underlying assumptions that make it work for sites that don't go out of their way to accommodate it. I get that may be an unintended benefit to some people who want to be control-freaky about how their offerings are used , but it enormously reduces the value of the web.
A connection just hit me - it is similar-ish to way back when, when some misguided designers wanted to publish PDFs on the web instead of HTML. They wanted full layout control, at the cost of basically everything else. Angular is similar, in that doing anything the creator didn't anticipate is very difficult, despite complying with the letter of web standards.
And also, as I said, I default to browsing with JS off. When I hit a blank page, I curse "Angular" and go back to the search engine. So that's annoying, but I've yet to encounter an Angular site I can't live without.
Add these lines to your nginx config file:
# Only support HTTPS and enable http2
# which will allow you to multiplex requests.
# In a separate config do a 301 from normal HTTP
# for all paths / subdomains.
listen 443 ssl http2 default_server;
listen [::]:443 ssl http2 default_server;
# Make TCP send multiple buffers as individual packets.
tcp_nodelay on;
# Send half empty (or half full) packets.
tcp_nopush on;
It reduced our page load time by around 75%, and more for people accessing the site from the other side of the world.It seems that you've followed [1] by the similarity of what that suggests to what you've suggested. tcp_nodelay seems to force it to not wait 0.2 seconds, which nginx does so that it doesn't send lots of really small packets. I can see this increasing speed by 200ms (a lot) but increases network traffic as a cost.
tcp_nopush seems to be about optimising each packet sent, reducing the total size of them.
Interesting stuff though!
[1] https://t37.net/nginx-optimization-understanding-sendfile-tc...
Of course you'll have to check your own site for your own performance issues, but while travelling from Toronto to Vilnius I found our site much slower and investigated why. We have a webapp running Ember with mix of largish assets, like images, and smallish assets, talking to a JSON API standard API written in rails and proxied through nginx (responses are usually around 1kb of JSON, but this gets compressed with HTTP2).
In reality waiting 200ms is an enormous cost and carries a hidden cost: Larger packets are easier to lose. If you're going mostly to browsers that is more and more often going through wifi or cell tower networks with at least 1% packet loss and when the resend happens you're stuck with that 200ms delay again.
That's my understanding, but the real truth is "it just works for me when testing around the world with real-world use cases, try testing it out yourself and please let me know if you learn something new".
The problem with webapps is a compounding effect of many requests.
a) dependency chains (asset A loads B loads C) mean the browser only knows to load C once B has been loaded, so the latencies (including nagle) add up. http/2 push is supposed to help with that.
b) statistics. your 95th percentile latency is irrelevant if you're loading hundreds of assets and want to know when the site has finished loading.
RTT hurts in my world.
Variation in page load times far outweighed any difference the config change might have brought.
So, YMMV.
From a theoretical standpoint I agree http/2 should help (although in practice I didn't measure any difference). I can't see why tcp_nodelay would help except in extremely niche applications (sending tiny responses with lots of pauses?)
Indeed, given that HTTP/2 forces TLS (for any mainstream browser) which forces an extra RTT, for light pages my evidence is that HTTP/2 makes things marginally worse...
Will wait for something else.
Why would you use nginx for security? I see over 100 CVEs for nginx.
This isn't exactly unexpected. Nginx is an incredibly complicated piece of software written for performance in C. That's basically the trifecta of security badness.
If you want security, why not use a (possibly slower) simpler product written in a safer language?
Also, lots of CVEs indicate only one thing: that lots of people are checking it.
Doesn't it also mean they're finding lots of problems with it?
Also that it has lots of vulnerabilities.
Instead, you could use a product that's secure by design.
Of course, it doesn't support HTTPS or HTTP/2, so... there's that.
That's astonishingly good. It's great that they think they can do better but bloody hell.
It's indistinguishable from magic.
Many car engines have an efficiency of around 20% for example, as compared to the maximum theoretical efficiency of 67%. If car engines performed only well enough to meet your standard of amazingness, your gas costs would be 7 or 8 times higher. Or if you used a modern Toyota engine for comparison, something like 14 times higher.
Obviously I took an example from a different sphere for comparison... if you really find this to be an astonishing result can you elaborate?
But in true there are production cars that pass 430 km/h, resulting in a factor of about 2.5 millions. A lot less than 10 millions.
sun -> mj-PV -> battery -> electric motors run circles around those.
Says the person who can't even photosynthesize!
As a trite comparison, it's a similar to arguing the theoretical maximum speed that a person can run. Interesting, but doesn't tell you anything about the maximum speed that a person can move.
The speed of light in a vacuum, though, is qualitatively different. It's not a local maximum, it's the global maximum.
Anyway, I reserve the right to be impressed. I spend enough of my life being jaded and cynical. Sometimes I see things like this and it makes me think about stuff that I take for granted.
I'm always on fast networks (100mbps or so), but I find myself waiting 4-5 seconds or more before Safari even shows signs of connecting, even to things like Google and YouTube. It seems to be limited to HTTPS sites, and seems to have been introduced when Apple started using a new MacOS TLS/SSL framework called something like "Apple secure transport" (I forget exactly) a few MacOS versions back, and it's almost certainly not DNS (I get the same issue with Google DNS as with my provider's).
I don't even know how to debug this, since I don't have a command line client that goes through the same TLS/SSL framework code as WebKit. Curl is fast, for example.
Speed of propagation in typical cable is around 66% of c. That is not too bad. More latency is probably caused by buffering.
Speed of light is just too low. We should get a lobby-group working on it and the limiting laws of physics.
The timing calculation approximation we used had no term for 'c' in it at all. The simple approximation is just an RC timing curve: total capacitance of the wire and gates that you need to change, through the on-resistance of the driving FET. The more complex version treats each wire as a transmission line and incorporates its inductance, but ultimately for electrical signals travelling in a medium near a ground the important factor is the dielectric, not the speed of light.
The other massive limiting factor in processor design is heat conduction.
With neutrino-based signalling you eliminate both the lower propagation speed and the suboptimal path.
Basically, lightning causes ripples in the Earth's electromagnetic sphere. Those ripples travel across the surface of the earth at an apparent speed that's faster than light, because the actual path is through the earth.
It's possible as far as the physics go but also impractical given the particle accelerator and detectors needed. And the post-processing the scientists had to do to tease out the timing information also would negate any latency advantage.
An accelerator- or reactor-based neutrino experiment can detect whether the source is on or off, so in principle you can send binary information. This has been demonstrated using NuMI/MINERvA [1]. It's more of a novelty than anything practical.
https://www.scientificamerican.com/article/china-shatters-ld... http://spectrum.ieee.org/telecom/security/two-steps-closer-t...
Because we don't know how to build an optical router with the same capabilities as the ones we use today and the alternative is to give up on net neutrality. I'm ignoring endpoint performance and protocol bloat.
There is work being done in quantum computing that might eventually lead to some CPU like device operating on photons alone and hollow core fiber optics also seem to be an active field. The work on upgrading the various protocols together with limiting network requests seem to be the way to go until our technology catches up.
You also get more hops thanks to some backbone networks abandoning MPLS in favour of native routing/packet labeling.
In Romania, as I understand, networks are largely free of L3 switching overhead thanks to large portion of networks running pure L2
L2 switching is faster than IP routing because even cheap "smart" L2 hardware come with hardware acceleration, but a typical non-core router for a mid-tier ISP is usually a consumer grade, PC hardware based router.
>100ms
Under load with anymuch sizeable routing table?
Linux is able to handle about 5 Gbps of traffic of small packets (64 bytes) with one core. This gives a latency of 100ns to keep the pace. You need to find a Linux router 10000 time slower than this to get a 1ms latency.
People who care about how fast the internet is are not the same set as people who know statistics acronyms, and they don't show up in a simple web search. You know how else we can make the internet better and faster? Publishing articles with software and licenses that make it easy for people to improve and share your article.
Some classes of issues will be better solved if the solver understands stats. M/M/1, distributions, percentiles, etc, yes. If only I understood them...
By my local major London paper somehow needs 4MB of crap and >60s to display a 2 para story. I don't need anything as subtle as "stats" to tell them that they're holding it wrong and why I avoid their site like the plague.
Good one.