As developers, we try everyday to squeeze till the last byte and optimize things. We all know how performance is important.
So why download for every website the same asset: React, jQuery, libraries, CSS utils, you-name-it? What a waste!
As developers, we try everyday to squeeze till the last byte and optimize things. We all know how performance is important.
So why download for every website the same asset: React, jQuery, libraries, CSS utils, you-name-it? What a waste!
I'm not really sure why this isn't already in place.
edit: The reason I'm not sure why because it sure seems to me that multiple threads on the post are all suggesting basically the same simple idea, the ability to serve a file by it's hash (either just by it's hash alone, or by a url + the hash). Personally, I think whatever form these url's take i think it ought to be backward compatible which I think is possible.
Maybe you're right. But still, today we need very complex build tools and silly server hacks like setting expire headers in 10 years.
I wish instead it was as simple as deploying my 3kb app.js and tell the browser: "Hello there, here's my manifest with all the dependecies I need to run my app. Thank you.".
And that's exactly what Nix and Guix do. Elegant solutions to age old problems.
And I also have the feeling that surely, just this exist, I just don't know about it ^^
Linking things by hash, other than being so very ugly, consider this case: you have this asset that changes every minute, and is several MB. If a users users your site too much, you will flush out out all other cached stuff with old versions of that file that will never get referenced again. That just strikes me as extremely wasteful, that is, you get a short term boost but even worse performance overall. If other sites do it too much, it will mean your own stuff will not even be cached when visitors come back.
Mozilla already implements it: https://hacks.mozilla.org/2015/09/subresource-integrity-in-f...
That sounds like IPFS, [1].
Everything we need is already in place, except for a tweak in the caching strategy of the browsers[1]. With Subresource Integrity [2] you provide a cryptographic hash for the file you include, e.g.
<script src="https://example.com/example-framework.js"
integrity="sha384-oqVuAfXRKap7fdgcCY5uykM6+R9GqQ8K/uxy9rx7HNQlGYl1kPzQho1wx4JwY8wC"
crossorigin="anonymous"></script>
As it is, browsers first download the file and then verify it. But you can also switch this around and build a content addressable cache in the browser where it retrieves files by their hash and only issue a request to the network as a fallback option, should the file not already be in cache. Combine this with a CDN which also serves their files via https://domain.com/$hash.js [3] and you have everything you need for a pretty nice browserify alternative, without any new web standardization necessary.[1] And lot's of optimization to minimize cache eviction and handle privacy concerns, but that are different questions.
[2] https://developer.mozilla.org/en-US/docs/Web/Security/Subres...
[3] Imagine if some CDN would work together with NPM, so every package in NPM would already be present in the CDN.
I think that using the integrity attribute is great because if it happens it's going to have to work through a lot of tricky implementation details (e.g. things like origin laundering) of moving to an internet of content by hash rather than content by location.
However beyond just having an integrity attribute added to html I am interested in the question of how do we encode an immutable url as well as the content-hash for what it points to (as well as additionally required attributes) into a `canonical hash-url` (i.e. encode all these attributes) that is backward compatible with all the current browsers / devices, and which browsers can use in the future to locate an item by hash and/or by location.
The driving reason for this encoding is make sharing of links to resources more resilient, and backwards compatible. Eventually the browsers could parse apart the `canonical hash-url`s and use their own stores for serving the data, but not until the issues (and likely other unthought of ones) listed in the sri addressable caching document you linked are worked through.
One thing that I've found incredibly disappointing about SRI is that it requires CORS. There's some more information here: https://github.com/w3c/webappsec/issues/418 but it essentially means that you can't SRI-pin content on a sketchy/untrustworthy CDN without them putting in work to enable CORS (which, if they're sketchy and untrustworthy, they probably won't do).
The attack that the authors lay out for SRI requiring CORS is legitimate, but incredibly silly - a site could use SRI as an oracle to check the hash value of cross-domain content. You could theoretically use this to brute force secrets on pages, but this is kind of silly because SRI only works with CSS and JavaScript anyway.
And unfortunately, there are too many hosts that make the attack you mention credibly silly:
It is not uncommon that the JavaScript served by home routers contains dynamically inserted credentials. And the JSON response from your API is valid JavaScript.
To be completely honest: Only reach out if you have solutions for any of the problems or can reduce what you want down to something that is solvable with these problems in mind.
If your solution does not live on the web, you'll have a hard time finding allies in the standards bodies that work on the web :)
You'll have a hard time convincing spec editors and browser vendors already. The working group mailing list is https://lists.w3.org/Archives/Public/public-webappsec/
If you have minimal edits to the spec, we can take it straight to Github. SRI spec contains a link to the repo.
The problem is that www.victim.com/evil.js doesn't exist, and never did, but your browser won't know that if it's in the cache -- this gives you a way of faking files existing on other servers at the URL of your choice, and as long as they're in the cache you'll get away with it.
[1] https://news.ycombinator.com/item?id=10310594 [2] https://news.ycombinator.com/item?id=10311555 [3] https://news.ycombinator.com/item?id=10312333
What sort of failure modes are there when something inevitably goes wrong, either through malice, incompetence or sheer plain accident, and you end up serving different content from the same hash?
The probability of collision is still negligible.
If you're drawing at random from a pool of size k you need approximately sqrt(k) draws until you reach a ~50% chance of a collision[0].
With 256 bits, there are 2^256 possibilities, so following the rule-of-thumb you'd need 2^128 draws until you had a 50% chance of a collision.
2^128 > # of atoms in the universe.
If you adjust your risk tolerance you'll have different numbers come out, but the chance of a collision in any realistic scenario is negligible.
The Subresource Integrity spec for strong content hashes could improve cache hits by allowing a different URL to be used if the hash matches but everyone wants to avoid that turning into a massive security / privacy hit — see e.g. https://hillbrad.github.io/sri-addressable-caching/sri-addre...
(That's not an exaggeration: I've measured uncached latency for DNS + connection for a CDN host in the high end of that range, especially over cellular connections or outside of the U.S.)
International connectivity out of Korea is pretty congested, especially at peak hours. Tokyo is 30-40ms away in the morning but randomly jumps to 150ms+ in the evening. It can take more than 1 second for the DNS lookup, TCP and SSL handshakes, let alone the actual transfer. So unless the CDN in question has a physical presence in Korea, it is almost always faster to load assets directly from the web server.
I suspect that many regions outside of US/EU are in a similar situation. Using a POP in another country 2000km away does jack shit for local websites, and only harms companies that fall for aggressive CDN marketing.
Presumably there's a reason that you'd _want_ them to use 1.12.4 over 1.0.0 (or you'd want to at least check that they didn't break anything that you rely on in 1.13 when that comes out)?
You'll also miss other features: version range alone could save tons of bandwidth.
but centralizing in the browser is?
Packages inevitably create a tangle of complex problems that page/document model avoids. That is one of the main reasons web apps replaced normal apps in so many domains. Not to mention that a package system will give a huge competitive advantage to already popular scripts, which will inevitably lead to technological stagnation. And it will remove the biggest disincentive to avoid bloat.
If someone loads a bad JS library, it doesn't matter where it came from.