For Static Sites, There’s No Excuse Not to Use a CDN
forestry.io
forestry.io
There are plenty of reasons not to use a CDN, not the least being that you might not want to give a third-party access to your traffic. Even static information might be sensitive, accessing some forbidden data can put people at risk. The central position CDN are increasingly taking on the Internet make then worryingly nice targets for snooping by surveillance bodies.
CDNs are beneficial for images and static asset files because pages often have dozens of those so the small speed up of going through a CDN is multiplied many times over.
The difference yielded by moving a single file (the primary page) to CDN is unlikely to make any measurable benefit overall unless you specifically run a site with a very geographically distributed user base.
Even for sites with a NA/European user base you can host in either US or Europe and it makes no difference. The latency is not significant compared to other performance issues.
I imagine data mining for CDNs is very lucrative given they interface customer AND business.
they block 3rd party scripts
"They" who? I am blocking 3rd party scripts with uBlock..Well, you can often pay much less for bandwidth than alternate solutions.
What we need is for gateway.ipfs.io to use geo-dns such that it can serve this purpose, and IPFS gateways can be run all over the world by regional ISPs, universities, and even individuals.
It does a pretty good job for USA visitors (which is where the majority of my traffic comes from). But it'd be ideal for IPFS to use geo-dns, along with FileCoin implemented so that I can pay for nodes in other countries to have my content pinned.
Also, my site uses Service Workers to cache all of the content + TurboLinks, which drastically speeds up every subsequent page transition. I'm not sure if that's doable if you're using an external CDN like Cloudflare for all of your resources.
As a domain owner, you have to give your HTTPS private keys, and all your users private data (authentication cookies, passwords, etc.) to your CDN, or you have to do a lot of careful dividing up the "static" resources from the dynamic stuff on different domains, serving them with different certs.
As a web user, some CDN's like cloudflare offer 'HTTPS', but with plain HTTP as the backhaul to the origin server. That tricks the user into thinking their connection is secure, which is, IMO, immoral.
That people can browse Wikipedia, and their ISP / government is prevented from knowing whether the article being read is about knitting or a political issue is a good thing.
It can't, which is why it's ridiculous to say "there's no excuse not to use a CDN".
Snark aside, it at very least dilutes the brand of SSL to give the pretense that they're giving it to you with the default option, and give you the pretense of giving you something extra on top of it with the "full (strict)" option.
For example, until recently many people used it to add a custom domain to sites hosted on github pages, but the connection from CF to GH was still encrypted.
Static / dynamic is orthogonal to private / non-private. It's important to consider that what public or static information a user requests may be private for them.
Consider for instance, if I go to WebMD's page on e.g. Lupus. If the static, public image on the page were pulled in from a CDN, the fact I visited that page is now leaked to a third party, for which they could infer whatever they wanted to. Or maybe sell that on, so drug companies can hit me up with targeted ads.
If you're using an ad blocker, wouldn't this be irrelevant? The drug company would be wasting money buying data that will be worthless.
No; using an ad blocker won't prevent this kind of data collection. Ad blockers prevent you seeing ads, and some can block certain page elements associated with data collection (e.g. social media buttons), but this type of data collection is indistinguishable from the content itself.
While there are still other users who don't block ads, the incentive to collect such data for targeted advertising exists, and I think it's unlikely you will be specifically excluded from the dataset on this basis. I think it's also worth bearing in mind that it's in the interest of a data vendor to inflate expectations of the relevance of the data they sell.
Mostly due to the fact that its 2 statically resources than can be loaded in around 0.3 seconds.
Foresty.io can load the HTML and the CSS in the same amount of time, but then another 8 megabytes of JavaScript and imagery follows.
Optimizing for ping time is premature optimization when website obesity is the elephant in the room.
> Communicating with the more distant server caused a 500% increase in round-trip latency. 250 milliseconds may sound like a short amount of time, but this time penalty applies to every request the user makes to that server: CSS, JavaScript, images, etc. Even small amounts of latency can negatively impact your site. If there were an easy way to eliminate this latency, wouldn’t you want to do it?
(150 assets | 8MB (lol)) may sound like a small amount of data, but bandwidth delay product and sequential asset requests can result in time penalties which are multiplied with latency to significantly increase page load time.
You can optimize your TLS handshake (you're using TLS, right?) to fit in fewer TCP fragments (don't include the whole chain), use OCSP stapling so the client doesn't need to make a separate request to your certificate authority before starting to load your site.
You can remove the penalty of sequential asset fetches using HTTP/2 to push / prefetch assets. This should almost completely eliminate "asset delay product" parts of your load time from having too many assets that aren't queued as part of the initial request.
You can cheat bandwidth delay product by using Google's BBR TCP congestion control on your servers, which provides much higher bandwidth in the face of packet loss (packet loss which will probably be more likely for clients with higher ping). You might not really notice this, but someone loading your site from Australia over janky undersea cables will thank you later.
At this point, if you look at the flame graph of your site loading, you may have realized that the steaming garbage heap of JavaScript and fonts is still taking several seconds to literally transfer over the network and parse. You should now meditate on what part of "static site" requires so much dynamic client code, and try to trim your site down to something that could at least fit on an effing floppy disk.
Maybe after you've looked at all of these, maybe, if you have globally distributed users and site load time is still a major issue for you, maybe then evaluate what putting a CDN in the middle would do for your load time. But loading your site from a CDN will never make up for the 5 second render time caused by your endless bloated chain of abstractions.
I mean I really agree with you and think more people should do exactly what you're doing, but comon, let's bikeshed this at least a little :)
No! Of course not!
I don't understand why every damn site add those annoying popup. I would love to know his many subscriptions sites receive from those popup.
You open one connection and then the rest flows. For static sites it might be enough.
Of course, in case of worldwide audience it might be a problem, but with DNS anycast can’t you resolve this with putting your website on local providers ?
Then you're effectively rolling your own CDN. Definitely an option for some cases, but overly complex and costly for most.
I think the most likely outage I might encounter would be due to operator error, eg. accidentally pushing a bad web server config to my 3 servers.
More details: http://blog.zorinaq.com/release-of-hablog-and-new-design/
> As a result my site has had 100% uptime since its deployment years ago
Assuming your individual hosts did go down occasionally, how do you know how many of your visitors waited 2 or 3 minutes for the browser to try the alternative IP?
I set up dozens of caching reverse proxies, distributed on a few VPS providers. Each VM then uses strongswan to route to my primary origin servers. In some cases, there is no extra bandwidth cost, if the caching VM's happen to be in the same datacenter as the origin servers, as I can use the private interfaces for my strongswan traffic.
If I want to poorly mimic the geographic DNS behavior of CDN's, I can use split views in DNS to very roughly send people to a closer caching proxy. It isn't perfect, but then neither are CDN's.
To take this a step further, I can use multiple domains with TLS SNI from different registrars to provide some take-down resistance.
The advantage to this model is that one CDN or VPS provider does not have control over my content.
The drawback is that I have to manage these nodes myself. Nowadays that isn't too bad, because each VPS allows for making API calls to spin up VM's with pre-built images. Ansible also allows for adding new nodes dynamically. There are community playbooks for most VPS providers.
If you read [The First Round Review](http://firstround.com/review), you've seen it in action.
GitHub pages is free and highly scalable, but unless you're a paying user GitHub has some limits. So for things like very popular projects you could simply move your stuff to a server+nginx or s3+cloudfront. However, these cost money so it's hard for many to justify doing that when GitHub Pages is free. In comes Netlify which provides the "better" GitHub Pages experience with the ability to use whatever you want that can be installed using go, python, ruby, or node and the resulting markup is distributed via a CDN all for free.
With such a nice deployment story the question becomes: why not just put every website like this? Mainly because it's all based on using Git to edit markdown files which usually non-technical people have issues with. Netlify offers NetlifyCMS as a solution to this and Forestry offers a more robust version of this with a more powerful and clean editor.
I find myself putting CloudFront in front of pretty much everything to unlock speed, security and now even functionality like auth thanks to Lambda@Edge.
I have a CloudFront add-on for Heroku that, in some cases, can double performance without any application changes.
https://www.mixable.net/blog/making-heroku-fast/ https://elements.heroku.com/addons/edge