Public CDNs Are Useless and Dangerous
httptoolkit.tech
httptoolkit.tech
Required reading, “More than you ever wanted to know about font loading on the web”: https://www.industrialempathy.com/posts/high-performance-web...
Edit:
> Developers need to use a system font stack which loads fonts from the system itself instead of remote web fonts
IMO the takeaway is "self-host fonts, make sure they're small files, load early, and don't block render," not "only use native fonts"
The default font display setting for Google fonts is 'swap'. Using font-display means the browser will block rendering for a short time, and then use the fallback font if the font hasn't loaded, and then swap in font when it does eventually load.
But, and here's the important bit, the block period for 'font-display: swap' is 0ms. In other words, Google Fonts works how you think it should already so long as the developer has included a fallback system font in the font-family.
I suspect you're attributing a blocking problem to Google Fonts when really it's caused by something else.
I go back and forth on allowing it to set background and text colors. Forcing black text on a white background does make some pages pretty ugly, but it does force everything to be readable.
EDIT: Unless there's a way to include metrics outside the font there really isn't a way to prevent reflows when it loads other than blanking the entire page until then.
(There are many ways to avoid reflows with font loading, eg font-display: optional)
- The page takes an extra 100ms to load
- ~ a minute later (when you're halfway through reading it) the content will jump by some large amount.
There isn't a nice way to do custom fonts on the web and you should just avoid them.
> The page takes an extra 100ms to load
Nope. Slow connections' floor for first paint is 100ms when using `font-display: optional` – not an additional 100ms. The page is still loading like normal in this period. (If you're already under a 100ms FCP for 3G users, you know all of this already)
> the content will jump by some large amount
`font-display: optional` is literally designed to prevent this from ever occurring. Either the font is ready for first paint in 100ms, or another font is used.
On Firefox, this is as simple as going to about:config and setting gfx.downloadable_fonts.enabled to false.
On Chrome, it requires adding the flag --disable-remote-fonts to however you're opening it.
It's not even clear if using Google Fonts is GDPR compliant (since you're leaking the fact that your visitor has visited your website): https://github.com/google/fonts/issues/1495
Even if you do want to use your own hand-picked font on your site, just self host it, or bundle it with something like https://fontsource.org/
They provide the tools to bundle the fonts and serve them yourself: https://fontsource.org/fonts/merriweather-sans
EDIT: The original comment mentioned "fortawesome" instead of "fontsource". I mixed up the names, my bad!
Using FA as a CDN is not GDPR compliant, either.
Generally speaking using any public CDN is not GDPR compliant. If you _could_ self host the files, then you can not meet the necessity test required for any of the relevant legal basis in GDPR art 6(1).
Google Fonts and FA are worse, because the personal data is shipped off wholesale to the USA. Neither FA or Google (for fonts) offers a data processing agreement that would make it legal to use.
They provide the tools to bundle the fonts and serve them yourself:
https://fontsource.org/fonts/merriweather-sans
Will update original comment.
> The disclosure of the user's IP address in the above-mentioned manner and the associated encroachment on the general right of personality is so significant with regard to the loss of control over a personal data to Google, a company that is known to collect data about its users, and the individual discomfort felt by the user as a result, that a claim for damages is justified.
[0] https://rewis.io/urteile/urteil/lhm-20-01-2022-3-o-1749320/
The user likely has no idea they contacted Google — it only happened because the code in your application executed it.
The reasoning was that as it is so very easy to self host Google fonts and with caching as an argument gone there is actually not enough reason anymore to use the hosted version by Google and transmit PD to Google.
Would the site have used a foundry that didn't offer self hosting the reasoning might have turned out differently if the site owner could have made a strong argument for using this specific font by this foundry.
At least that was my reading but IANAL.
Your website is malware that contains code which makes the visitor's browser connect to Google's server.
Though it arguably doesn't get enforced anyways, except for some big corps.
It doesn't often really work out that way.
Tax Law is especially interesting here, because you can basically depend on what the IRS has said but there's a chance that even if they said X, the tax court will actually rule Y (in your favor). Sometimes the risk can be worth it (if you made a good-faith effort to comply with a reasonable understanding of the law, and the court rules against you, you just have to pay the tax; even if you were found unreasonable you pay the tax + penalty which often isn't that much).
Browsers have always loaded a font from the system over a web font if it has a font of that name.
We could just normalize having a lot more free as in beer fonts installed on the average user's machine.
The unfortunate emphasis there is "average user's machine" as font loading times are already a a deanonymization vector/fingerprinting. Way back in the day IE made a big deal about bundling a group of fonts and calling them "Web Safe" and we need an initiative like that again. Take the Top X fonts from Google Fonts and just bundle them with browsers or operating systems and spread those as widely as possible so that it isn't a useful deanonymization vector.
This isn't actually accurate. It's possible for a CSS author to write @font-face rules that behave in this way (by including local(...) sources ahead of remote-font URLs), but that's by no means universal practice.
https://developer.mozilla.org/en-US/docs/Web/CSS/font-family
`font-family` rules apply ahead of `@font-face` rules. `@font-face` rules are used to source fonts not already installed on the system.
(ETA: local() is to provide alternate family names when a font may be installed under different family names.)
See https://www.w3.org/TR/css-fonts-3/#font-family-desc:
> If the font family name is the same as a font family available in a given user's environment, it effectively hides the underlying font for documents that use the stylesheet.
Given CSS such as
body { font-family: Palatino, serif; }
@font-face { font-family: Palatino; src: url(palatino.ttf); }
this will block the browser from using an installed Palatino font and tell it to load the remote resource instead. So it's possible to use @font-face { font-family: Palatino; src: local("Palatino"), url(palatino.ttf); }
to ask for the locally-installed font if present, falling back to the remote resource otherwise.Self-hosted web fonts are fine.
Going back and forth on allowing sites to set text/background colors.
Good content doesn't need graphic design tricks to remain interesting.
More like self-indulgent wankery.
https://addons.mozilla.org/en-US/firefox/addon/decentraleyes...
Given the web is about 65% browsed by a Google client, I'm surprised Google Font aren't directly installed into each system fonts on first load.
https://addons.mozilla.org/de/firefox/addon/localcdn-fork-of...
But sometimes if the webpage looks funny, you have to deactivate it.
An interesting take on which part of the dog wags.
In your model, developers seem positioned to serve designers.
Having written a book on design, and practised it, I have to say that web is very unusual in allowing designers such rein. In other disciplines, like sound design or product design, the constraints and requirements of the project come first, and designers need to fit their work to that.
A very great many web projects I have seen fail went south the moment "designers" got involved. Designers seems to wield an inappropriate level of power and influence, and operate without proper technical supervision, often breaking things or creating unreasonable demands.
I think this misalignment comes from the early days of the web, when anyone who could write HTML called themself a web designer and took on an entire site development. Even though post-CGI/database the server-side roles were better differentiated, the legacy of "aesthetic driven design" lingers.
I do not want to denigrate web designers, but times have changed and in my opinion, having a website that "looks great" just isn't that important these days. I'd much rather use simpler, accessible, secure, private, content-centred sites.
Why exclude aesthetics? There is no need for it. It positively affects usability, just like the other qualities you described.
It does indeed. "Aesthetics" in the most accurate sense of the word, as opposed to orderliness or just "looking right", is about how it makes us feel. Some people seem particularly sensitive to it and say they can't use a site or application that feels wrong. I respect that nuanced choosyness. Indeed I'm sure there's a biological thing going on, as in the way we select berries to eat or choose partners to mate with. Hard to introspect parts of our brain say "that's okay" or "that's off, don't touch". Aesthetics is very real.
But (there's always a but :) ... it should never be put ahead of functionality. As I see it that's what makes "Design" design, and not Art. It's blending aesthetics and elegance within a practical value set. And very often that "requirements set" is given to us designers by another... a film director, a product manager, an information architect etc. They say - "Here's what it has to do, (functional requirement) now go design it (non-functional)".
And what I am saying is that, in Web Design at least, the designers have gotten the upper hand, and that's partly the fault of developers, system architects, or project managers for letting themselves get pushed around like that.
That would be the role of a designer who is doing their job properly and paying attention to the specification for a simple, accessible, secure, private, and content-centred site.
- If the font has associations with a particular place or time that you want to imbue your system with.
- If you give the user themselves the ability to select a font that works best for them (for example a font focused on the dyslexia reading experience in an ereader).
Such as?
I checked the docs though and it seems like self-hosting is within the license.
Maybe material icons are different or maybe I'm misreading this page.
Isn’t the maximum concurrent multiplex streams on one connection, like, 1000?
(I suspect you didn’t mean to say multiplexing)
Edit: Firefox by default uses no more than 900 parallel HTTP connections (check your about:config page), and I assume this holds true for all windows/tabs combined - which is plenty.
But now that you bring it up, isn't the 6 resources per domain outdated? At least for HTTP2? When I've done performance testing on my site lately it definitely seemed to load more than 6 resources in parallel from one domain but I haven't looked into this topic in a long time (back when it definitely was limited)
If you've run the numbers, and that's a problem, hey, great, more power to you! There's web sites where that's a problem. But, you know, remember that 100 new users per sec is a scale of nearly 10,000,000 per day. Is that really your scale? There's an awful lot of sites, even busy, productive, and profitable sites, that are looking more at the "1 new user loading a fresh copy of these things per second", if not one per minute.
I remember working with "servers" in the low hundreds of megahertz, on software stacks a lot less optimized than today. Shifting around your static serving could do something then. Today? Choking even a single CPU server's capability to serve static content takes some doing. By the time that's your core problem, you'll know.
(There's still a lot of "best practices" banging around the community from the "low hundreds of megahertz" period of time. At the risk of goring some sacred oxes, I also consider the obsession with total statelessness to date from this era. While one must still be careful with it, judicious application of state in the modern era can be very helpful. It's less panic-inducingly terrifying when I can afford the equivalent of entire servers dedicated to each individual user, where "server" is defined as "a server class machine from the time this advice first percolated out". The software engineering considerations around state are relevant and must be considered, but the solution of "zero! none! never! not any!" is no longer the only viable or best choice.)
If your web page needs 10MB of static files to load, you're going to need a lot of round trips to get that to your users. Your server and clients may be on 10G ethernet, but that 10MB download is still going to be limited by 'slow start' congestion control unless the round trip time is low. That's why people want to use CDNs more than reducing resource use of serving static files (which is pretty low as you mention).
This does not negate your argument, in fact for many use cases doing your own caching could be a better way to meet increased demand (certainly a cheaper way). CDN-related issues can be incredibly difficult to troubleshoot, too, especially without a support contract.
But I can tell you from personal experience I've been asked to produce a "CDN plan" for content access on the order of once a minute during peak hours, and which had no reasonable business case for distributing enough to matter.
I am personally responsible for one system that does need CDN-like capability, because it does serve static files at a rate and a required reliability that a single server can't meet. (We don't use a "real" CDN because for our use case it would be absurdly expensive, but we're keeping an eye on the offerings out there.) So I do know they can exist. But even that system is, in 2022, on the lower end of what you would need such a thing for, and it is serving multi-gigabyte images out on a fairly routine basis to many tens of thousands of consumers (which are automated systems, not humans). That's what it takes to be on the "low end" of blowing out a normal static file server. 15 years ago these systems were a royal pain to maintain, at a much lower level of usage, and non-trivial effort was spent optimizing what they serve. Today, they mostly just hum away doing their job.
People were less obsessed with statelessness back then. There was some push for statelessness (peaking at the time of the 10K problem), but it was much less than now.
I believe the current wave is entirely caused by the publicity of the "we rent IaS, but not as commodity computers" cloud, and doesn't have any technical origin.
Cloud VM instances are for the most part very underpowered (and for the performance, overpriced). You'd have to go far back a lot of years to get to an era where a "beefy system" of the time was as weak as the smallest AWS instance (t2.nano?)
I seem to recall some people measuring the effect even before it was dead, and the hit rate was never all that great anyhow. There were so many versions of jQuery, so many other libraries, and so many CDNs, it still didn't hit as often as you might think.
The issue is this: many sites would download files from 10+ different domains and a single ‘long-tail’ DNS query, increasingly likely if you have more domains, will cause more delay than you could possibly save by using multiple servers.
(cdnjs, that I maintain and is referenced in the article, does both of these by default if you copy a script/link tag from our site.)
(ps, I dislike the name; you don’t encounter the term “subresource” anywhere else in webdev. It’s a “subresource” of the HTML document resource, eyeroll)
A PWA with a Service Worker could perhaps implement its own client-side "gateway", translating public gateway URLs into direct IPFS access. Without the Service Worker (or without JS) it would fall back to using the gateway.
It sounds like we'd need to have browsers implement IPFS protocol support directly for this to be a feasible alternative to centralized CDNs.
But that isn't how humans work, and there needs to be some grease easing the process at the start. So they still serve a function, even if it's not the one that the CDNs imagined at the start.
We’re all on a journey in our careers. Those of us with more knowledge and experience can and should help those with less, instead of gatekeeping and lecturing.
I'd much rather work with someone that just needed to learn a few things I take for granted than someone who decides it's their right to say who's in or out of IT haha. Teach or get out of the way.
There is a difference between capability and knowledge. We’re each on a journey.
> There is no single service that can go down [...]
> Each piece of content is loaded entirely independently, [...]
Both of these statements can be true, but in fact, what ends up happening is someone puts their site on a pinning service like Piñata or Infura, and then it's right back to being centralized and trackable. Filecoin helps decentralize who pins, but even then, one ends up using a pretty interface like web3.storage since using the real service is just too much overhead for just pushing a page.
I'm personally a big fan of IPFS, but I'd use it to make an entire site reliable for users who care to self-pin.
I haven't been following IPFS very closely recently, but this doesn't track for me at all.
So people pin some content on a big node.
Once you download it, now it's on YOUR node, and your friend can still get it off you. Your node is just as authoritative as the big one. They disappear and IPFS can carry on.
I remember people criticizing NFT's using IPFS links as "not solving anything" and I was truly truly baffled at why anyone would think that. You could easily hold on to your image on a thumb drive, or absolutely any method of backing content up. If every single copy on the internet got nuked, you could just publish from your usb drive again and suddenly it's back up, available for all to see. More to the point, anyone could do that. So anyone (including you) with even a slight motivation to preserve that data could authoritatively do so.
disclaimer: I do not like NFT's. I don't like them practically, I don't like them abstractly, and I don't like them culturally. It's just this is one criticism thrown at NFTs I don't agree with, IPFS makes this aspect work.
Eg skypack is just awesome, it lets you import any npm package onto a website as if it was a proper es6 module. Google fonts is amazing in similar ways. Dangerous, fine, but I empathically disagree with calling that "useless".
https://esm.sh - lots of options
https://jspm.dev - does import maps (with a polyfill) https://generator.jspm.io/
https://esm.run - jsdelivr hosted esm builds but only ones already uploaded to npm
If you copy the tags from cdnjs.com this attribute will be added automatically.
Very much performance for the time invested.
But why?
I'm confused. Where is the difference between using these companies as CDN providers versus using them as cache providers? The only minor difference that I can see is that my server would provide the original copy for every cached item, but this does not mitigate the risk of cache poisoning and the privacy risks that you also encounter with CDNs.
I think I'd prefer deploying a new CDN link to updating DNS. "No code changes or backend deployments" makes the assumption that DNS is a manual change (not terraform, etc.)
> Our systems are designed to remove HTTP referer information before logging, to avoid associating requests with any individual website using the Google Hosted Libraries.
So they at least are claiming the CDN is not used to snoop on traffic paterns.
TIL.
Best Practice does not mean "advice that I believe in." Best practices are legally binding documents issued by insurers.