Please stop using CDNs for external JavaScript libraries
shkspr.mobi
shkspr.mobi
That being said I'm amazed how much production software depends on multiple libraries that are developed and maintained by a single person as a hobby.
It's pretty common for someone building a library to do so for the utility, scratching a real itch. Not universal (people build libraries to be someone who built a library, too), but common.
It's very common for agencies to be building client projects in order to bill for it. Sometimes there's a special alignment where the client is also primarily interested in spending an allotted budget so long as there's a plausibly adequate deliverable.
And Guava in Java land, and Abseil in C++.
It's not too surprising that the stuff big companies make for their needs isn't as useful for small indie devs as stuff small indie devs make.
Virtually everything we do is in-house on top of first party platform libraries - i.e. `System.`, `Microsoft.`, etc. We exclusively use SQLite for persistence in order to reduce attack surface. Our deliverables to our customers consist of a single binary package that is signed and traceable all the way through our devops tool chain, which is also developed in-house for this express purpose of tightly enforcing software release processes.
This approach is certainly slower than vendoring out everything to the 7 winds, but there are many other advantages. Every developer knows how everything works up and down the entire vertical since its all sitting inside one happy solution just an F12 away. Being able to see a true, enterprise-wide reference count above a particular property or method is like a drug to me at this point. We are definitely over the hill and reaping dividends for building our own stack. It did take 3-4 years though. Most organizations cannot afford to do what we did.
[Edit] - But what I have found through experience is that the code that was written under these constraints seemed to be better, more secure and robust, than without them. YMMV.
Plus if all you know is this custom stack, where are you gonna go?
Dependencies in general, like 95% (or more) of the kind we see in modern package managers? No, they're mostly untested liabilities and the majority of them could be rewritten in an afternoon.
This whole discussion is a bit strange. GP clearly uses dependencies, just not as much as everyone else today. I don't understand why the fixation with polarizing the discussion into "use lots of dependencies" vs "write everything from scratch".
Dependancies are inevitable.
I'm OK with depending on glibc and _almost_ as OK with depending on openssl. And of course I inevitable rely on gcc/clang/CPU microcode/transistor photolithography...
But those examples are a world away from npm pulling in 12+ levels of random dependancies of unknown origin buried so deep it's all but impossible to audit - and the segment of "the software comm8unity" who built that particular house-of-cards and then promotes it blindly to the segment of the software community who obliviously use it to generate code that runs on expensive production platforms in business critical applications just boggles my mind.
(And makes me alternately weep and drink heavily - having to support exactly that in production because "velocity" and "moving fast and breaking things" are considered more important that quality or security by both management and clients - both of who will happily point the finger of blame elsewhere when the inevitable happens, in spite of having been repeatedly warned... :sigh: )
Believe me, as someone who cares about code, I far prefer a world where I know how every line of code in a system I'm building works. In fact, in my free time, I do just that. But my job as an in-house software dev for a non-tech company is to solve business problems with technology, and as with every business problem, there is an acceptable level of risk that you need to be OK taking. In our case in this post-COVID world, people's jobs are literally depending on us iterating quickly.
I spend a lot of time talking to "full stack developers" who came out of graphic design into front end web dev and then fell into nodejs backend dev, who's only "architecture" course was about drawing buildings and who's main complaints about security requirements are that bars across windows look ugly. A startlingly large number of them don't even know there are things they don't know. I'd say fewer than half of them could tell you what OWASP was, and fewer than 10% of them could tell you what SQLi or XSS was, and whether their sites need to consider them as attack vectors. Most of them just say "I use $frameworkDeJour, it handles all the security stuff!"
(BTW, last time I needed to supply dynamic PDFs in a web app, one of the good fullstack-via-FE-and-graphic-design devs build the PDF generation in the browser (using some random js/pdf library) so we could just feed it JSON from the backend. Made me sleep better at night doing it that way...)
I know _of_ HIPPA regulations, but not being in either the US or in healthcare records, I have very very vague notions of the HIPPA requirements. Same with PCI compliance - I know there are important rules and requirements, which I don't fully understand the details of, because I choose to use 3rd parties like Stripe to handle all my CC processing so that those requirements don't apply to me (with there exception of needing to understand the risks of webapps and problems like XSS in the context of Stripe-powered CC forms).
Within the codebase, the infrastructure and tooling is by far the most important aspect in terms of productivity and stability of daily work process.
If you take the time to position yourself accordingly, you can make the leverage (i.e. software that builds the software) work in virtually any way you'd like for it to. If it doesn't feel like you are cheating at the game, you probably didn't dream big enough on the tooling and process. Jenkins, Docker, Kubernetes, GitHub Actions, et. al. are not the pinnacle of software engineering by a long shot.
Unless you happen to be one of the very rare companies that sells source code and not built artefacts, your asset is the built artefact and your code is the expense you take on to get it.
Having less code to get the business outcome only makes sense when you see the code as a cost, not a thing of value itself.
> your asset is the built artefact
Therefore, summarized, you mean that the sourcecode needed to generate the resulting <app/service/whatever> is a liability, but that the result can be an asset (if it does generate external revenue, or internally lowers costs ,etc..)?
I personally never thought about this kind of separation - interesting.
It's easy for us as developers to think that source code is valuable - but this leads to problems like never removing code "in case it's needed", or with an in-house dev team developing systems you could get off the shelf.
If code is an asset, then it makes a lot more sense to write stuff yourself: you not only get the artefact, you also get the source.
If it's a liability, then it makes a lot more sense to let someone else bear the costs of that liability, especially if they have economies of scale, except where you can't get your desired outcome other than writing code.
This is of course technically incorrect. It's perhaps more accurate to say that source code requires upkeep and is expensive to maintain. That tends to draw less interest and discussion, because it's "obvious." Except as an industry we're overall pretty lousy at paying the required upkeep on code.
The moment you want to minimize something while still achieving your goals you know it's a liability. Do you want the same profits or solutions with a smaller team? Then the team is a liability. If team were an asset you'd be trying to hire a bigger team without any work for them to do. If you had the chance to double egg production with constant demand, you'd eat or kill half your hens. They're a liability.
That's only true if you assign an appropriate meaning to the word derogatory.
Why not just throw it away, then?
If you can't, it's (IMO) because you're stuck with the burden of having to write and maintain that code.
If it's an asset, why not just write more?
Companies do write more.
So in your analogy, what's the software equivalent of replacing the drain pipe?
Its an asset. Like many (virtually all, other than pure financial) assets, it has associated expenses; maintenance, depreciation, and similar expenses are the norm for non-financial assets.
> Unless you happen to be one of the very rare companies that sells source code and not built artefacts, your asset is the built artefact and your code is the expense you take on to get it.
No, things that are instrumental to producing product are still assets, not just the things that you sell. That's true if its machines on your factory floor, if its the actual real estate of the factory, or vehicles that you use to deliver goods. And all of these assets, like a codebase, have associated expenses.
The whole "code is a liability, not an asset" line is something from people who might understand code, but definitely don't understand assets and liabilities.
Its more just a mental footnote that the contract with the customer is what is valuable, and the code is either supporting that value or destroying it. So if you can have the same contract with the customer for less code that's the outcome to strive for.
More lines of code is not equal to better. What you need is better lines of code, which usually means less of them.
Let a dev loose on a codebase and they could add value, but they could also be subtracting it. Hmm maybe developers are the liability.
Perhaps the amount of code is the wrong metric to optimize. Perhaps it is readability, simplicity, maintainability that we really want and those are harder to put into numbers than LOC.
Airplanes have many functional parts and many not-so-functional parts. There are parts of the airplane that will, if removed, prevent it from working in various critical ways.
But from another perspective, all of those parts, decorative, functional, or essential, are liabilities, dragging your airplane back toward the ground when you want it to stay up in the air. The fact that a particular piece is important doesn't mean it's less of a problem having it; it means you have to suck it up and work around the problem.
Source code is like this. The mere fact of its existence causes problems. Some of it you can't do without. But it's causing problems anyway, and if you can do without it, you want to.
So you delete your source code after you ship an app? That doesn't make sense. The source code is one of every software company's greatest assets which is why we have created so many tools to keep track of it like version control, automated testing etc. Otherwise there would be no point in clean code, code documentation etc, just write something that solves the problem and done! No need to even check in the code, just build a binary locally and ship!
I understand the point, but saying that code is a liability and not an asset is false no matter how you look at it. Source code solves a lot of problems you can't solve with binaries, it lets you adapt to change much better. So instead of saying that code is a liability, use the business saying "Focus on your core business", meaning don't write code for things that isn't your core business.
I didn't say that. I said it was a liability. "Don't write code for things that aren't your core business" doesn't tell you that you should try to minimize the amount of code that addresses your core business. But you should.
Likewise, code could be many lines because it is well formatted and robust. Alternatively it could be long because there's a lot of repetition and bloat. It could be short because it is very targeted to what it needs to be, or it could be very short because the developer used a lot of unreadable code-golf tricks.
In any complex system, you have optimization problems. Very rarely is the answer to an optimization problem simply to minimize an isolated variable. You can not say a lighter airplane is better than a heavier one without asking why it's lighter; likewise a smaller codebase with fewer lines can not be called better unless you know why it is smaller.
Just because a plane could get off the ground without something does not mean that cutting that thing to save weight will move the plane closer to its optimal design. Likewise just because the code base could be reduced does not mean that actually makes it more maintainable and better performing.
If a business owns a car that they use to make a profit, guess where the car is in the balance sheet?
> Having the biggest codebase(car) is not a benefit.
Sure, lines of code is not how you measure the value of code as an asset, just like tons of gross weight isn't how you measure the value of vehicles as an asset.
That doesn't mean code, and vehicles, aren't assets.
“Lines of code are a measure associated with maintenance cost, not asset value” is a reasonable statement. “Code is a liability, not an assset” is not.
To get technical about it, code whose discounted future maintenance costs exceed the discounted revenue it is expected to bring in or helps to bring in is a liability. Some maintenance costs are invisible, such as having team members leave; leading to recruiting expenses to seek a replacement and the associated ramp time when they're hired. Companies are often not equipped to assess the cost side in any way that approaches reality and just stick the cost of the "code" on the balance sheet as capex. Having an asset on the balance sheet after doing this doesn't mean you have an asset in reality.
> The whole "code is a liability, not an asset" line is something from people who might understand code, but definitely don't understand assets and liabilities.
Code can be a net liability. Unless you are looking purely at the asset side of the balance sheet and ignoring liabilities, which tends not to be very useful.
The reason is very obvious - the most expensive part of most projects is the people. Why pay them to write code that you can just download for free? That's a tremendous waste of money.
Most of us just do what we're told and we're in no position to question the "company's priorities" -- nevermind the fact that very often those priorities actually align with the techie's vision.
Here's an unrecognized truth:
The cost of forking a third-party library and maintaining it in-house solely for your own use is no higher than the cost of relying on the third-party's unforked version. Depending on specifics, it can actually be lower.
Note that this is a truth; the only real variable is whether it's acknowledged to be true or not. Anyone who disputes it has their thumb on the scale in one way or another, consciously or unconsciously.
Specifically, I think hard forking is a bad idea for any sort of library that needs to be regularly updated for compatibility or security reasons.
If you're depending on some random person on the internet to update software which underlies your whole stack, then when the next imagetragick drops you can't update until they get around to fixing it. Since you won't have developers familiar with the code, fixing it won't likely be feasible for you. That's a lot of risk.
(I actually included the appropriate hedging to clarify that my comments are scope-locked to that topic and to prevent digressions like this, but I edited it out because it made the comment too hard to read. Goes to show...)
Well, not just client-side JS; server-side, too, or anywhere that NPM is used, but even more than that: e.g. other package managers that were influenced by or work similarly to NPM and encourage a similar package-driven development style, e.g. Rust's crates.io or the Go community's comfortability with importing by URL. It applies for many of those cases, too, it just wasn't the focus of my comment.
Could you elaborate on the actual argument for why this is the case. On the surface it seems like the opposite of what you are claiming can just as easily be true depending on the situation. For example take a widely used utility lib such as lodash or jQuery. In your scenario there are two options:
1. Use lodash via a package manager and rely on the lodash team to fix bugs, write tests, and add new useful utilities over time.
2. Fork lodash and take on the maintenance burden yourself. You are responsible for keeping up with security vulnerabilities and making patches. You are responsible for writing high quality tests for the parts of your code that diverge for the original.
Think about how much ramp up time is required for new hires to become familiar with a company's codebase. Why would you ever want to devote that amount of time to maintaining code that for the vast majority of businesses is already "good enough". For some companies the engineers working on the open source project may even be more competent than the resources available in house. Sure there may be edge cases were performance or security is absolutely paramount, and in those this approach may make sense, but not for the majority of generic CRUD apps.
I would have a very tough time trying convince a competent manager that it is beneficial to devote so many man hours to this task rather than to business logic, new app features, etc.
C# 8.0 / .NET Core 3.x / AspNetCore / Blazor / SQLite
Of these, SQLite is arguably the most stellar example of what open source software can provide to the world.
Everything else in our stack consists of in-house primitives built upon these foundational components.
I'd love to understand why you chose that.
Feel free to hit my up at the address in my profile if you don't want to talk here.
Effectively, our client's operations cannot rely on cloud services and all of the related last mile connectivity into their infrastructure. If AWS/Azure/et.al. go down, many of our customers are still able to continue operating without difficulty.
When I go looking for a dependency, I check the license and I have a quick read of the code.
Last time I did that was for autocomplete. I checked the most popular six options; in the space of 2 hours, I found obvious-from-reading-code bugs in all 6.
None of them had a CLA signed by contributors, so there's really no evidence their code is genuinely available under the license they claim to offer.
I wrote my own. It took about 3 hours initially plus 2-3 hours ironing out edge cases found over the following weeks. It only added 700 bytes to my bundle.
Total time spent: 1 days work. Smaller code, loads fast, free from license issues, does exactly what I want.
It's not NIH syndrome, usually. It's about having control over the whole software supply chain for security, reliability, licensing compliance and general quality.
1) NIH. Almost always the problems that need solved are interesting, and engineers are naturally chomping at the bit to solve them. Added bonus you can potentially make a name for yourself. This happens way more than it should, in cases that don't meet the other two ways I saw. Solving problems that have already been solved very effectively and efficiently, in a mature low friction fashion.
2) It doesn't scale to needs. A lot of software just doesn't scale to the requirements of the platform. It's hard to understate just how much traffic and work a lot of Amazon infrastructure has to handle. Most software doesn't scale that well because it's not run in so big an environment. We're using some well known commercial software at my current employers (because it works, has a good reputation, and did everything we need), that is experiencing major scaling issues because we're literally orders of magnitude larger than any of their other customers. We're seeing stuff they've never had to deal with before. We're not even close to Amazon's scale for this particular type of software.
3) Need to control the entire software stack, have the ability to drastically modify it to meet the changing demands placed on it. A lot of public software is written to meet one need, and it rarely changes that drastically over time. That's not what the consumers want, even though needs change over time. Change your software too much and you'll lose your existing users that fundamentally need what the software is providing. You can see the boom and fall of it all with so many projects. Take a look at what's happened with Chef and Puppet, for a quick off-the-top-of-my-head example.
That's why I wrote "usually".
Maybe, though that's usually the exact rationalization given for NIH syndrome. I mean, “NIH syndrome” is never the stated reason for anything.
(1) Need to interact with other internal systems.
(2) Pure NIH.
(3) Need to scale further than outside solutions.
1 and 3 are closely related. There are quite a few legitimate category-3 internal services for provisioning, configuration, service discovery, monitoring, upgrades at various levels, fault remediation, etc. Any other production service would have to interact with most or all of them. It's often easier to build something local than to add all of those "touch points" to an open-source project. I know because I did both while I was there.
But pure NIH is very close behind as a reason. Despite all protestations to the contrary, engineers get far more "impact" for creating new things than for fixing old ones. It's hard to get somebody to do X when their bonuses and raises are better served by doing !X. This ends up amplifying, instead of attenuating, the natural impulse of all engineers everywhere to build new things because it's more fun. People always make up other reasons, and perhaps even believe those reasons themselves, but nine times out of ten those reasons are pure delusion.
In general, yes, and that's a problem. Luckily in some teams it's much better.
> but nine times out of ten those reasons are pure delusion.
If you were in a company/team with such level of true NIH syndrome it's good you left.
Dependencies are painful to pull in and only pulled in when a dev needs a specific version. The internal repos end up being a missing version nightmare where people cobble together whatever works with what's available. Where feature and security upgrades go ignored, left to rot like the brains of the devs who struggle to keep up with what's available in the real world.
Many of those networks I've worked on are becoming more permeable at the edges because the cost of the air gap outweighs any benefits.
You can actually read the whole thing: https://www.motionpictures.org/wp-content/uploads/2020/07/MP...
I've had to deal with legacy projects that include multiple versions of the same libraries, all of them being shaded/relocated so they don't conflict. The result, however, is bloated binaries that takes 30 minutes to build.
Seems arbitrary, but I bet the rationale is fascinating. Could you go into more detail? I'd love hearing a little about your work experience, the industry, process, etc.
Your intuition is right: they're afraid of the content leaking.
For cloud providers, this results in ... amusingly long documents. Here's GCP's at 110 pages [1], while the AWS folks were clever and used landscape mode for theirs [2] so that it's only 59 pages :).
[1] https://cloud.google.com/files/gcp-mpaa-compliancemapping.pd...
[2] https://d1.awsstatic.com/whitepapers/compliance/AWS_Alignmen...
It'll be interesting to see in the future when content is cheap to produce. I predict a complete shift away from this.
> It'll be interesting to see in the future when content is cheap to produce. I predict a complete shift away from this.
It’s actually one of the most obvious forms of friction, causing an increase in the cost of doing business. A lot of VFX houses take these rules to imply that they must segment their networks and keep workstations completely unable to reach the internet (coming full circle to the airgapped comment at the start). Pretend that your entire development workflow is like being on a plane when the WiFi is down. That’s modern VFX software engineer life :(.
Google themselves do this with gstatic.net and ytimg.com etc
> Google themselves do this with gstatic.net and ytimg.com etc
Most probably not. The point of cookieless domains is that you can use a very simple web server to serve content (no need to handle user sessions, files are pre-compresses and cached, etc.) and it lowers incoming bandwidth a lot. If you have a lot of requests (images, css, js) the cookie information adds up quickly.
Opening video thumbnails from ytimg.com will still be cached for youtube.com as before. The only thing that will change is for embedded videos on 3rd party websites as those won't be able to use caches ytimg.com thumbails from elsewhere.
The current way seems like needless DNS spam to me...
Given that generally people have slower upload than download, shaving off a few bytes from requests is worth it.
I also recall that browsers [used to (?)] limit concurrent requests per domain which this helps work around
If Site A loads a specific JavaScript file for users with an administrator account, Site B can check to see if the JavaScript file is in your cache, and infer that you must have an administrator account if the file is there.
The attack can happen with different types of resources (such as images).
Furthermore a CDN can't track you as simple as you might think, it often would require thinks which need explicitly opt-in agreements on a per website basis to be legal.
Furthermore due to technical limitations you can only get that permission from the user after the CDN was already used.
CDNs can still track aggregated information to some degree but they can't legally act like a tracker cookie.
I suppose with HTTP2 some of the benefits of serving JS through CDNs are gone anyway, so I guess it's time to stop using them.
If you don't use any of those sites, you're considered higher risk/fraudulent user/bot.
Here's an example of a very short and easy way to see if someone is probably gay: https://pastebin.com/raw/CFaTet0K
On chrome, I consistently get 1-5 back after it's been cached, and 100+ on a clean visit. On Firefox with resistFingerprinting, I get 0 always.
> Here's an example of a very short and easy way to see if someone is probably gay
Ok, but now the resource is in my cache, so from now on they will think I'm gay?
This resource is just generic, so probably not, but if you actually visited grindr's site without adblocking heavily, they load googletagmanager and a significant number of other tracking services, which will almost certainly associate your advertising profile and identifiers as 'gay'
I also can't believe they send/sell your information to 3 pages worth of third party monetization providers/adtech companies for something that is this critically sensitive.
const start = window.performance.now();
const t = await fetch("https://example.com/asset_that_may_be_cached.jpg");
const end = window.performance.now();
if (end - start < 10/*ms*/) {
console.log("cached");
} else {
console.log("not cached");
}> In that case, the browser would always load the asset (it is not cached).
Agreed, if the cache is partitioned per domain AND the current domain has not requested the resource on a prior load. If the cache is global, then the asset will be loaded from cache if it is present: https://developer.mozilla.org/en-US/docs/Web/API/Request/cac...
> So the rule would be that only stuff that is directly in the <head> may be cached (or stuff that is on the same domain).
You could be more precise here: with a domain-partitioned cache, all resources regardless of domain loaded by any previous request on the same domain could be cached. So if I load HN twice and HN uses https://example.com/image.jpg on both pages, then the second request will use the cached asset.
Ah right, the thread is becoming long :)
> So if I load HN twice and HN uses https://example.com/image.jpg on both pages, then the second request will use the cached asset.
Good point!
Someone below mentioned doing requests for a large image that requires authentication. Short response time means the user isn't logged in (they got a 403), long response time means they downloaded the image and are logged in.
There are, of course, other vectors to consider, but I can't think of any that could be abused by third parties. If anything, isolating caches would make it easier for the CDN themselves to carry out the attack you mentioned, as they would be receiving all the requests in one batch.
I'd be able to tell if you visited Fox news recently, correct?
But three specific files can already be pretty unique. I chart.js with two specific plugins in my toy project, and I'm willing to bet that no one else on the world uses the exact same set and version configuration.
[1] - https://www.webdigi.co.uk/demos/how-to-detect-visitors-logge...
Plus the trend now is to use webpack and have all of your deps bundled in and served from the same server.
https://developers.google.com/web/updates/2020/10/http-cache...
And by going that route you make sure that all pieces of your website have the same availability guarantees, the same performance profile, and the same security guarantees that the content was not manipulated by a 3rd party.
You can already guarantee the security of the file by using the integrity attribute on the <script> tag. And the performance of your CDN is probably worse than the Google CDN (not to mention that you lose out on the shared cache).
> And the performance of your CDN is probably worse than the Google CDN
What means probably? Other CDNs (Akamai, CloudFront, Cloudflare, etc) are also fast.
And by pushing one piece of your website on a different CDN you force your users browser to create an additional HTTPS connection which takes additional round-trips, instead of being able to leverage one connection for all assets. This alone might as well outweigh the performance differences between CDNs.
Also the "shared cache" benefit might go away, if I read the other answers in this topic correctly.
I mean, wouldn't that take care of a whole class of attack vectors and make cross-origin requests possible without having to worry about CSRF?
While privacy sensitive users may consider this a feature in case of e.g. google.com and youtube.com, the average user is more likely to consider it an annoyance, and worse, it is likely to break some obscure portal somewhere that is never going to be updated, so if one browser does it and another doesn't, the solution will be a hastily hacked note "this doesn't work in X, use Y instead" added to the portal. And no browser vendor wants to be X.
[1] The workaround of using the public suffix list for such purposes is being discouraged by the public suffix list maintainers themselves IIRC, so the "right" thing to do would be breaking Wikipedia.
Edit: If done naively on an origin basis right now, it would break the Internet. You couldn't use _any_ site/app that has login/account management on a separate host name. You couldn't log into your Google account with such a browser anymore (because accounts.google.com != mail.google.com). Countless web sites that require logins would fail, both company-internal portals and public sites.
1) User logs in at google.com/login and sets google.com cookies. 2) Server generates a nonce and redirects to youtube.com/login?auth=$NONCE 3) youtube.com checks the $NONCE and sets youtube.com cookies 4) youtube.com redirects back to google.com.
Firefox's container tabs can maintain isolation despite this since even this redirect will stay within a container. However there is a usability penalty since the user has to open links for sites in the right container (and automatically opening certain sites in certain containers will enable cross-container stapling again).
webapps.stackexchange.com/questions/30254/why-does-gmail-login-go-through-youtube-com
And on a side note, very unhappy about how the entry to be a developer has lower significantly over the last 10 years or so.
Security is a concern, use SRI.
Reliability can be mitigated with fail over logic to a backup.
The part missed is bandwidth. Using a CDN means your web server doesn't have to serve out static files that you are paying per a GB to serve. Small sites it's not much but it does add up. It's a Content Delivery Network not a Cache Delivery Network.
- reliability
- delivery speed through closeness to user (having nodes all around the world)
- cost
- ease of use
- handling of high loads for you / making static content less affectedly by accidental or intentional DoS situations
That multiple domains might use the same url and might share the cache was always just a lucky bonus. Given that the other side needs to use the exact same version of the exact same library with the exact same build options accessed through the exact same url to profit from cach sharing it never was reliable at all.
I mean how fast does the JS landscape change?
Given how cross domain caching can be used to track users across domains safari and Firefox disabled it a while ago as far as I know, and chrome will do so soone.
it all looks like https://cdn.example.com/foo/bar.js?v=129a1d14ad3
QUIC puts everything in UDP, so theoretically its a never ending firehose of data for a download with the occasional "hey, I missing packet 3, 12, 18, please resend". Mimicking TCP but putting the app in control versus the kernel.
> QUIC congestion control has been written based on many years of TCP experience, so it is little surprise that the two have mechanisms that bear resemblance. It’s based on the CWND (congestion window, the limit of how many bytes you can send to the network) and the SSTHRESH (slow start threshold, sets a limit when slow start will stop).
https://blog.cloudflare.com/cubic-and-hystart-support-in-qui...
Per stream and per connection flow control windows, which kind of indicate how much data the peer may send on a given connection before it gets a window update. Those windows also indicate how much the server is willing to store in its receive buffers, since the updates are likely sent when those buffers are drained.
A congestion window, which indicates how many low-level packets and data in them can be in-flight without being acknowledged. Those also account for retransmissions, and packets which do not necessarily contain stream data.
"Speed:
You probably shouldn’t be using multi-megabyte libraries. Have some respect for your users’ download limits. But if you are truly worried about speed, surely your whole site should be behind a CDN – not just a few JS libraries?"
Also even before QUIC HTTP/2 fixed a lot of the problems with distance as you no longer need to wait for separate handshakes for multiple files to be streamed. QUIC will still give a few advantages but again those advantages would be good to have on your whole site not just a few libraries.
But to your question though even un-cacheable content can be "those parts of". There are products from CDNs like https://blog.cloudflare.com/argo/ which combine CDN cache tiering with higher tier network transport to origin servers for all cache misses (or uncacheable content). Again though, it depends on if it's critical to your site's performance or not. If you don't have a bunch of uncacheable content, that content doesn't need the absolute best transport, or the time/money could improve some other part of the site speed more then it's not critical to your site's performance.
The final nail in the coffin for me is that CDNs are a shared resource, if your CDN is getting heavy traffic or otherwise suffering, it becomes your slow point, while your site is fine. I just don't see any upsides worth the tradeoffs.
If you have scale that demands some serious content distribution, that is different, I would argue you shouldn't be relying on public shared CDNs then even moreso. Pay for a CDN service or roll your own.
Because PoPs are closer and transfer speed ramp rate scales to latency, the further away a server is, the longer it takes the download to ramp up to full speed. This is especially relevant when talking about smaller resources like javascript, css, and small or optimized images, and webfonts.
From some quick research, it doesn't seem like the script tag has built-in support for this. One could imagine something like multiple src attributes (used as a search order for the first valid file), but that doesn't seem to exist. So it seems like the web page has to do it manually.
Which I guess means you have to have some javascript (probably inline, so you know it's loaded and for performance?) to check and fix the loading of your other javascript.
If it's really that manual, it sounds like it adds cost to implementing this correctly. in other words, it might be one of those scenarios where correctness is achievable, but it's a whole lot simpler to just not do it that way.
Link for reference. .Net Core has this built in as a tag helper too! https://www.hanselman.com/blog/cdns-fail-but-your-scripts-do...
Quic can't defeat physics. Performance will still lineary degrade with distance to (edge) servers, and therefore CDNs will stay important.
What Quic however will do is reduce the time-to-first-byte on an intial connection by 1RTT due to one less handshake - which can be e.g. a 30ms win. After the connection is established it aims to yield more consistent performance than e.g. HTTP/2 over TCP. But packets will still require the same time to go from the browser to an edge location, and therefore the minimum latency for a certain distance is the same.
If you want the benefit of a CDN, you need to put your own code up there. And if you’re doing that, you might as well host your own copy of the libraries too, so the browser won’t have to talk to two different CDNs.
Incidentally, a few years ago when people were loading third party scripts over HTTP, I demoed a fun hack where, if you control a user's DNS, you could redirect queries to popular CDNs to a proxy that injects keylogger code and tells the browser to cache it indefinitely. Because at the time almost every site included either jQuery or Google Analytics, you'd have a persistent keylogger even after the user switched to a more secure connection. How far we've come!
I think this is the code: https://github.com/paulgb/cachebeacon
It basically just runs two servers:
- A DNS server that resolved a list of domains to its own IP.
- An HTTP server that looked at the HTTP host and proxied requests to the upstream server. If the content type or extension indicated that the response was JavaScript, it would add the payload and set cache headers to cache as long as possible.
- A special HTTP endpoint on the proxy server to capture data sent back from the payload.
That's the only reason you should ever need to not load any 3p javascript including google-analytics.
If you're like me and want to avoid loading popular javascript libraries from CDNs but want those webpages to work, get the Decentraleyes plugin: https://en.wikipedia.org/wiki/Decentraleyes
If it’s an important part of the site, it might make the failure more obvious in newer browsers.. but small libraries used on only some pages might not be noticed quickly... so you’d probably also want to test all of your resources regularly.. and even then the time between those test runs may allow some users to be compromised.
So using it is a good idea but it’s not a fix for the actual problem.
It is basic available for every browser except IE and Opera Mini, so I think it is user's problem to use an old browser that don't support a wide supported security feature.
And your response is: "that's their problem" ??
I hope you're not in charge of any important or large sites or anything that handles financial data (ecommerce, etc)... because this isn't a good attitude when it comes to security.
It's perhaps worth accepting there's no silver bullet here but a combination of initiatives like SRI is still worthwhile to help reduce the attack surface for the majority of users?
SRI is the equivalent of just enabling 1.2. You haven't disabled access to browsers that dont support SRI.
You 2nd sentence sounds remarkably similar to my first post that maple responded to: SRI can help mitigate the damage, but it cant fix it.
You seem confused about the difference between mitigation and fixing.
Mitigation: the action of reducing the severity, seriousness, or painfulness of something.
Key work there is reducing. A fix actually eliminates the issue.. like enabling 1.2 + disabling 1.1 eliminates the potential for communicating insecurely.
It's important to understand the difference because anything short of actually fixing the issue leaves some users exposed to the vulnerability.
I kept reading this article looking for an actual decent reason to not use a third party CDN, and I never found one. In fact, the right answer is really that you should always set up your CSP and subresource integrity rules to completely prevent these kinds of attacks, whether from an unintended script injection from your own domain or a 3rd party.
A CSP would have stopped this attack. The exfil server was baways[.]com.
Ticketmaster, on the other hand, did have a 3rd party JS that was compromised: https://www.riskiq.com/blog/external-threat-management/magec...
https://www.localcdn.org/ (fork of decentraleyes, with many more resources)
edit: will be sticking with the original, looks like the fork maintainer made no effort to work upstream first, which is a very bad look for what is essentially a piece of security software. https://gitlab.com/nobody42/localcdn/-/issues/5
> will be sticking with the original
It's your choice. The fork is better. The maintainer seems a bit more active (more updates) and extremely pro-privacy (I concluded this from his home page and extension settings)
It even opens donation pages locally, instead of opening the author's website. He says ''I think it is better if your public IP address is rarely listed in any server log files.''
You must enable the rulesets. It's very easy and a one time job. To generate them, go to LocalCDN settings and select your adblocker.
https://codeberg.org/nobody/LocalCDN/issues/51#issuecomment-...
I haven't gone searching for a PR yet and didn't think to do so beforehand (all made more complicated by both projects' repos having moved locations at least once recently).
Initial commit[1] in the LocalCDN repo is Feb 2020; I don't see any PRs on either of the Decentraleyes repos in early 2020 or late 2019. Of course, it's still possible the author reached out
[1] For some reason, the project does not continue the git history from Decentraleyes. For me this is a red flag (much easier to sneak in a change this way) and I will continue using Decentraleyes.
I realize there are benefits, but are the benefits so extreme that they merit all the hype around CDNs? So many developers talk about them like the web would crawl to a halt if they were stopped, but I think that they have their own slippery-slope of problems that has resulted in folks just hand-waving away the expense of web assets since it's a hidden problem from the developer. I doubt actual bytes transferred and latency are affected in a significant way as folk that promote CDNs claim.
It'd be nice to see actual comparisons in a real world scenario. Keep in mind, web site responsiveness is not just linked to download time/size, but also asset processing. If your page is blocking because the JS is still being parsed, the time you saved downloading it is moot.
If you're that stingy, why aren't you just putting the entire website behind CloudFlare and calling it a day?
Given, its for use in embedded webviews of a video game, and it powers like 50 different atomic interfaces via react.
When I look at AWS pricing, us-east-1 at < 10TB, I see: EC2 data transfer out to the internet is 9¢/GB, and Cloudfront is 8.5¢/GB for the lowest price class. That's a slight savings, but at 6% I can't justify the effort to switch over on cost alone.
Should I be looking at a different CDN service?
I’ve never used BunnyCDN, but they charge a flat $0.01/GB for North American traffic, and I’ve heard some good things about them.
DigitalOcean, Vultr, Linode, and some other cloud providers charge $0.01/GB without a CDN, just using their regular servers, but obviously a CDN is more than just a way to save money — it’s a way to lower latency and improve user experience.
The mega clouds (AWS, Azure, and GCP) seem to significantly over-charge for egress bandwidth as a nice profit mechanism, just because they can.
My unpopular opinion is that mega clouds are overrated. They’re fine, but they have a lot of weird gotchas that most people have just accepted as “how the cloud works.”
The other thing to remember is old browsers used to cap the number of connections per domain HTTP 1.1 only supports serial requests, so there were benefits to hosting on multiple domains.
Even today, the big benefit of a CDN domain is that CDNs can host static resources faster and cheaper than your webservers. Yes, you can forward requests from the CDN to your webservers, but it's also one point point of failure. What's interesting is that with a modern, JS-only site, the split becomes API and static JS.
It's not just geographic latency you're addressing with a CDN, you're also reducing the number of network hops. It's not uncommon to experience higher latency going from SF to San Jose datacenter just because you're on a "wrong" ISP. A good CDN usually has a POP on the same network as you.
Fortunately, these days that just means creating a free Cloudflare account.
In your case, are all CDNs equal? Do I just have to throw my content to the biggest provider?
Don't get me wrong. Disenfranchisng non-Western visitors is the last thing I want to do, but the issue is not CDN or not, it's caching content closer to people whose ISPs don't provide sufficient service outside their own borders.
I dislike that CDNs are the only way around this and I feel like it is centralizing Internet access in an unfavorable way.
The big data centers in this region are in Singapore and I expect all CDNs have a presence there, so probably. I haven't exactly done any benchmarking though.
> the issue is not CDN or not, it's caching content closer to people whose ISPs don't provide sufficient service outside their own borders.
How would you do this without a CDN (on a small budget)?
Scripts and CSS are (or should be) small compared to the images.
Yet, with the images, you can use mod_pagespeed on the server to replace images with picture source sets, with the server able to detect the bandwidth of the client and their device to serve them highly optimised images. So that means images that look glorious on a 4K screen with a good connection and images that are a bit jaggy for the person on their phone with only 3G.
I am sure that CDNs can do this too and that you can get mod pagespeed to work with CDNs but there is so much that can be done on a server without having to sign up for extra services and their overheads.
https://www.privateinternetaccess.com/blog/swedish-police-we...
uBlock Origin allows specifically blocking external fonts.
> I don't visit websites to marvel at their beauty
Using the right font is rarely about beauty.
Yes, it's often aesthetically pleasing. But good fonts can have lots of practical effects for people, especially those who are disabled or aging.
It's fine for you not to care about fonts and to block them, and it's fine for people to host fonts themselves (if the license allows), but it's misinformed to complain that all non-standard fonts are a meaningless aesthetic contrivance.
In any event, the point about accessibility is wrong because it's more common for designers to use custom fonts that are harder to read than easier.
I am running my own recursive DNS resolver and some sites pull crap from such a sheer variety of sources that before it all gets cached, the load times remind me of dialup 20 years ago.
if you're ok with that then by all means, use some ugly font. ime most authors aren't though...
I see a similar problem with Google Analytics or Google Fonts, where you're sacrificing user privacy and agency in exchange for developer convenience. In a slightly more privacy-centric era (perhaps in a few years), I think people will consider this practice unethical, as the web developer is sending their visitors off to fetch and run javascript code from a third party's computer. The security, usability, and privacy problems noted by the article are not worth the few minutes saved by the developer.
I'd also be very careful including a JS library from a CDN that gets auto-updated to the latest version, this might break your site without you noticing.
We did this with our own JS library (Klaro - https://github.com/kiprotect/klaro), for which we offered a CDN version hosted at https://cdn.kiprotect/klaro/latest/klaro.js. We stopped doing that with version 0.7 as we realized we could not introduce any breaking changes without risking to break the websites of all users that were using the "latest" version. So what we do know is that we only automatically update minor releases, i.e. we have a "0.7" tag in the CDN that will receive patch releases (0.7.1, 0.7.2 , ...) and which users can safely embed in their websites.
That said we always recommend users to self-host their scripts as well and to use integrity tags whenever possible, it's usually better from a security and privacy perspective as well.
Like most JS libraries Klaro is quite small (45 kB compressed), so the loading time is dominated by the connection roundtrip time, which again is dominated by the TLS handshake time. For our main servers in Germany we have a roundtrip ping of around 20 ms (for most connections from Germany), and with the TLS negotiation it takes about 60-100 ms to serve the script file from our server. If the server connection is already open that time reduces to 20-40 ms. So the benefit of hosting small scripts with your main website content is that your browser doesn't have to do another TLS handshake with the CDN, which for small scripts can dominate the transfer time. If you then use modern HTTP & TLS versions you can reduce the transfer time quite a bit.
We deliver integrity tags for our CDN files as well btw (https://github.com/kiprotect/klaro/blob/master/releases.yml), that doesn't solve the auto-updating problem though.
<script type='text/javascript' src='https://stats.wp.com/e-202041.js' async='async' defer='defer'></script>
Just ensuring the lib gets deployed to the right place, in the right version, and with the right caching-headers, has cost me hours of my life in some instances. When I let a CDN do it, it was minutes.
Those days are over. I don't load static libs any more, so the question doesn't even come up. Except maybe for fonts.
> You probably shouldn’t be using multi-megabyte libraries. Have some respect for your users’ download limits.
Can you imagine what would happen if BE languages had to deal with this problem? Want to pull in 30MB `numpy`? Too bad, find something smaller.
Understanding this problem illuminates why there are so many one-liner npm modules instead of monolithic utility libraries.
Other than embedded developers, who else has to deal with code size?
[1] https://addons.mozilla.org/en-US/firefox/addon/decentraleyes... [2] https://chrome.google.com/webstore/detail/decentraleyes/ldpo...
> Cacheing. I get the superficial appeal of this. But there are dozens of popular Javascript CDNs available. What are the chances that your user has visited a site which uses the exact same CDN as your site?
Quite high - if you're using the Google CDN I'd imagine.
> Speed. You probably shouldn’t be using multi-megabyte libraries. Have some respect for your users’ download limits. But if you are truly worried about speed, surely your whole site should be behind a CDN – not just a few JS libraries?
(a) Some websites/web apps do actually need to load a lot of JS to function properly (e.g. Google Maps). (b) Putting the HTML of a website into a CDN introduces extra complexity around making the html stateless (cookies, user data, etc.) + the issues that come from having to cache bust the CDN every time some dynamic content on the page is updated
> Versioning. There are some CDN’s which let you include the latest version of a library. But then you have to deal with breaking changes with little warning. So most people only include a specific version of the JS they want. And, of course, if you’re using v1.2 and another site is using v1.2.1 the browser can’t take advantage of cacheing.
Yeah, don't use library/latest.js. But then this basically boils down to the same argument as point 1 ("Caching").
> Reliability. Is your CDN reliable? You hope so! But if a user’s network blocks a CDN or interrupts the download, you’re now serving your site without Javascript. That isn’t necessarily a bad thing – you do progressive enhancement, right? But it isn’t ideal.
Fair enough - depends on your CDN. The Google CDN is probably more reliable than whichever one you're going to pay for instead (not every site can/should be put on Cloudflare for free-ish).
> Privacy
Yup - this is the real valid issue IMO
> Security. British Airways’ payments page was hacked by compromised 3rd party Javascript. A malicious user changed the code on site which wasn’t in BA’s control – then BA served it up to its customers.
If OP had actually read his own link, they would have seen that the British Airways JS in question was loaded _from their own CMS_ - not an external CDN. Ironically, in this case - it probably would have been harder to hack the CDN.
Plus, always use SRI.
> Quite high - if you're using the Google CDN I'd imagine.
Quite low in reality as there's no critical mass – most sites don't use the same versions of common resources.
Google fonts is one of the better candidates for shared caching working but resources just don't live in the browser cache for very long (Facebook, Yahoo studies)
Does it have to be all one or the other?
I feel like this kinda hand waves away an advantage.
Why should I trust this blog post on the "Caching" point, for example? It's got no data and no references.
"What are the chances that your user has visited a site which uses the exact same CDN as your site?" ... hey, you can measure that.
It’s trendy these days to require client JS for every old webpage, even ones that work just like webpages did in the pre-JS-all-the-things days.
This breaks graceful degradation, uses more power and bandwidth, takes longer to load, and is a security risk, for little or no benefit.
Please stop doing it.
See my comment above, if you have "x.js" in your cache and it's your first time on this site, it means that you previously visited another site that contains "x.js".
On the server you see that user requested the HTML for the page and other resources, but never requested X.js.
You could keep a count of requests, per user per file.
> The first difficulty of implementing cross-origin, content addressable caching on the Web platform is that it may leak information about the user’s past browsing. For example, site "A" could load a script with a given hash, then later when the user visits site "B", it could attempt to observe the time it takes to load that same resource and infer whether the user had previously been at "A".
[0] https://hillbrad.github.io/sri-addressable-caching/sri-addre...
https://developer.mozilla.org/en-US/docs/Web/Security/Subres...
This has been discussed before and it is unlikely it will be implemented as it is almost impossible to eliminate finger-printing and privacy issues if this cache is implemented plus it might lead to security problems:
https://github.com/w3c/webappsec-subresource-integrity/issue...
Whether it's for caching or security purposes, a hash of the contents is a hash of the contents.
The link you shared is about tracking whether a given document is already cached. That's the same problem as with normal caching. At least the top N libraries could be downloaded with a browser update and be equal for everyone.
They had an issue where they were responding with 302s redirecting to malicious adware:
https://twitter.com/unpkg/status/852660203275276289
They promised a post-mortem which (to my knowledge) never happened:
https://twitter.com/unpkg/status/852666784792444928
I removed unpkg from my infrastructure and won't be using any of these services again. The minimal gains aren't worth the risk in my estimation.
(Matomo is your friend.)
S3 is not a CDN, and you should not be serving assets from it directly
S3 is convenient, but assets are only served from a single region, and sometimes traffic from other geographic regions can be very slow. I'm in Europe and personally have seen excruciatingly slow transfers from us-* S3 buckets (10s of kbps, over my 1gbit connection). Even if you only expect files to be requested without being edge-cached, just sticking Cloudfront on top can delivery files much faster.
I just love this person's attitude. So many blog posts try to beat you over the head with some opinion that doesn't feel like such a big deal. The disclaimer on this actually gave the author a lot of credibility in my mind.
Great post, and I will consider my opinion changed on this topic.
I can presumably trust the code distributed by the 1st party, I can mostly trust code distributed by known CDNs, I cannot trust a randomized subdomain.
On another note of using CDN's, Deno is of interest. It's gaining popularity because it's using CDN's as part of it's pipeline dependency process. You don't have to waste a huge amount of time installing dependencies and node packages that you don't care for, just import and your done
If you're building a specific application, then I'm less likely to rely on a CDN... I think applications should largely deliver all of their own resources.
If you're in financial or security (military) applications, there may be legal, resource and other requirements that supercede what you can do.
It's not a one size fits all... that said, modern web application development centers around resources via npm, and largely delivered in your appliation bundle, so it's less of a concern in the common case these days imho.
It's caching, not cacheing.
- realistically, nobody is going to be testing if the fallback functionality works as the site grows so you're going to get weird breakages if the fallback misbehaves (e.g. your local file is a different version).
- big CDNs go down so rarely it's not worth thinking about (and if it is serve it from your own site)
- it's easy nowadays to put your whole site behind a CDN so there's little point (the article debunks the caching advantages)
I think people are just copy/pasting snippets that do this from somewhere without thinking.
For the 2-3 most popular, almost 100%?
Let’s take jquery for example;
Speed - if everyone loaded jquery from the official CDN, considering the usage of jquery is high, you’re very likely to get a speed improvement.
Faster website = happier customer (and happier search engine, which could mean more customers)
That to me is enough.
But... I do agree that if everyone uses that CDN and it gets compromised everyone is in trouble.
That to me is enough reason not to use it. The article doesn’t do this one any justice IMO + it’s a personal decision because major CDNs don’t get compromised every day, so convenience might win here.
[append]
By the way, the article title is not what the article intends it to be. He wants to say "please stop using shared CDNs for external JavaScript dependencies". There's nothing wrong with CDNs.
Perhaps what’s really needed is some kind of browser based package versioning and management system? Where your CDNs jQuery 1.10 is seen as a mirror of mine, with perhaps even some signing and authority of libraries not directly hosted on the host server? (I’m probably talking about something that already exists or has been discussed I’m sure...)
Recently, Sandstorm added the ability to block apps from loading external client-side scripts, and I hope more app platforms and hosting platforms adopt similar so that doing it goes out of style.
Problem: You can use a home-grown tracking solution, but, if the site has to be sold, how do you provide independent third party verification of your traffic details? Having Google Analytics helps you to do that.
Solution: Install and configure AWStats or Webalizer. Many lower tier hosts have these built into their offerings. Provide those statistics to interested buyers.
If your server is somewhat fast, supports HTTP2 and whatnot, clients will be able to pull these assets in crazy fast. If you're hosting a ginormous website a CDN might become interesting to improve worldwide latency.
I even saw people arguing using a CDN 'saves' a DNS request as the CDN is probably already cached. This is of course utter bullshit as this request is already made when doing the page request.
There is an argument for being more efficient with cdn redirects. The ideal cdn'ified site ensures that all requests that are always served from origin goes directly to origin, and all requests that are cachable go through a cdn, this means seperate domains or subdomains.
I use webpack etc for big projects.
1 million * 400 kilobytes = 400 gigabytes
For AWS the entry level tier per gigabyte is $0.15/GB.
That is a whopping $60 USD per MONTH that you can save by running your scripts on a CDN.
So, yeah.
CDN providers are not charities: they sell traffic metadata.
These numbers are certainly not on the low end.
https://www.ldoceonline.com/dictionary/for-the-sake-of-argum...
Also it's a bit like a proof. By assuming worse numbers, and arriving at $60, your better numbers will only prove the point even more.
It's like saying should I use a dry cleaner? Well it's only $50 to get them to iron my shirts, so it's worth it for my time saved. Then someone says "$50! You are being ripped off, if your "doing it right", you'd get it done for $10", to which I say, well if it's worth me paying $50, then it is worth me paying $10, so my original argument "use a dry cleaner" is made stronger.
My go to test for my QA team is to have them add a zero internet test case. Test your product in localhost with the network switched off.
HTTP/2 changed everything. Now the fastest way to serve your assets is to push them down the pipe that has already been set up.
Another issue with modern web apps is, they update almost every day.
Imagine bundle and reload the whole js file with only small changes is not a great idea
In a world where all websites have to be perfect, we would have far fewer websites.
The problem with WordPress is a lot bigger with the Themes/Plugins than with the platform itself.