HTTP Immutable Responses
tools.ietf.org
tools.ietf.org
https://www.ietf.org/mail-archive/web/httpbisa/current/msg25...
https://bitsup.blogspot.co.uk/2016/05/cache-control-immutabl...
Rough summary:
> At Facebook, ... we've noticed that despite our nearly infinite expiration dates we see 10-20% of requests (depending on browser) for static resource being conditional revalidation. We believe this happens because UAs perform revalidation of requests if a user refreshes the page.
> A user who refreshes their Facebook page isn't looking for new versions of our _javascript_. Really they want updated content from our site. However UAs refresh all subresoruces of a page when the user refreshes a web page. This is designed to serve cases such as a weather site that says <img src="" ...
> Without an additional header, web sites are unable to control UA's behavior when the user uses the refresh button. UA's are rightfully hesitant in any solution that alters the long standing semantics of the refresh button (for example, not refreshing subresources).
cache-control: private with either sliding or concrete expiration time already handles this.
Browsers mistakenly continue checking for new copies when they shouldn't within the expiration time. Fixing poor implementations with more standards never works well.
Immutable makes it clear that the server won't update the resource in place and will handle updates by generating a new one so the browser can happily avoid checking those resources on page refresh.
It's up to the server to use proper headers. Why say a file is ok to cache for years if it actually isnt? If the same url will change content then use shorter cache times or require active revalidation and/or etag checks - or just use the typical cache busting querystring parameters.
This "immutable" flag is unnecessary.
From the cache control rfc:
When a response is "fresh" in the cache, it can be used to satisfy
subsequent requests without contacting the origin server, thereby
improving efficiency.
From the immutable rfc: Clients SHOULD NOT issue a conditional request during the response's
freshness lifetime (e.g., upon a reload) unless explicitly overridden
by the user (e.g., a force reload).If the implementation is faulty, what is another spec going to solve? Again, there is no need for an "immutable" flag because existing cache headers already express everything that's necessary.
Also, I will certainly want to clear out my browser's cache on a regular basis. I do not want it keeping immutable things just because they shouldn't ever change.
This header won't make browsers cache data any differently. It skips a step when the cache is being read from.
That said, I am ultimately for this. I think. There is plenty of data showing that this is a low hanging fruit to hit.
Again, though, I am ultimately for this. I just remain skeptical of any panacea.
Should we optimize the web for clueless server operators?
The attack might work the other way around: the attacker buys a bunch of domain names, serves "sleeper" malicious JS files with this on common paths (say, the paths used by Wordpress and other common CMSs), then releases the domain. When the new owner installs a CMS and start serving their site, the browser loads the malicious JS instead, which is now running under the new site's Origin (security context).
No, because a malicious JS file by itself can't do much. The attack vector is the malicious JS running on the new site, with permissions to steal session cookies and interact with the application. That's why caching without verification is important: to make sure the browser uses the cached malicious JS instead of fetching the new one.
Another solution would be to use an unpredictable versioning scheme so the attacker can't anticipate the name of the resources.
Nothing is 'immutable' if you look at it hard enough. The Bible's teachings have changed over time. On the right scale, the Earth is a temporary novelty and questions about proton decay become relevant.
Tl;dr: set your expiration dates appropriately and you won't have a problem.
Even if the second owner follows the advice of setting expiration dates appropriately, nothing helps them.
I could readily see a fun novel about a nation state or isp using this trick to seed malicious code to a lot of public sites. Imagine setting it for index.html on any site with a source that merely bootstraps to the intended site. Along with the malicious code.
I'm not trying to say it is everyday. But it is easier than you think.
Browsers are removing bad authorities. Which is good. There are a lot of them, though. Which is a larger attack surface than many folks acknowledge.
Other situations I imagine would benefit from this are web crypto and HIPAA compliance.
Forever is such a long time.
Besides, aren't there already ways to say "cache this resource for <acceptable timeframe>"?
You never worry about when to expire your cache entries if the key changes every time the item does. It's nice to finally see cache-busting coming out of the woods.
So this is a new thing, properly called immutable responses, to tell the client that they really can treat this (in various ways) as a completely immutable response.
The example of a 1 year cache time seems a little extreme, though. I think a month would be better.
But the immutable keyword gives an additional clue to the agent about semantics, to inform caching.
For example: Say many websites start adopting immutable, version all of their resources, and ~25% percent of their resources are updated monthly. I could easily see the browser cache growing to an excessive size over a year before things begin being purged from the cache. Excessively long cache times are sloppy. That's why I think it matters.
They can also evict things before the freshness lifetime is up, and do.
The freshness lifetime is one of the things a browser can use to decide how long to keep something in its cache, but not the only one.
https://bugs.chromium.org/p/chromium/issues/detail?id=611416...
Instead, when the user attempts a full-page reload, Chrome will revalidate the resource in the URL bar, but not subresources (they will come from the browser cache, if the cache-control header checks out):
https://blog.chromium.org/2017/01/reload-reloaded-faster-and...
"Clients SHOULD ignore the immutable extension from resources that are not part of an authenticated context such as HTTPS. Authenticated resources are less vulnerable to cache poisoning."
This must NOT read SHOULD. It must read MUST! Otherwise, your computer will be subject to an executable planting vulnerability.
I'm surprised they don't have a much larger list of security considerations. There are many other issues that can happen.
There is a temporal difference. An attacker may wish to plant something today, in a coffee shop, that would execute in a protected environment, tomorrow. Immutability of caching can only help an attacker.
Yes, there are other ways for an attacker do this, but there is no reason to add to more ways! That's why the web is in its current state.
I see how you put "executable planting vulnerability" in quotes. Sometimes (in the current marketing), these are called APTs, but they have been around forever. e.g. Think of dll planting in Windows and the millions of attacks and three or four new API sets from Microsoft that resulted from that single ability to plant a dll in the search path.
This type of persistence can also be called incubation.
find ./ -type f -exec sed -i s/?sitever=X/?sitever=$VERSION/g' {} \;What's wrong with this approach?
The last company I worked for did that and everything was much slower and fragile than it would've been had we deployed packaged artifacts instead.
You can use an ETag of a checksum, instead of a checksum in the filename. Now a user-agent can just check with an if-modified-since and get a response quickly. But it's still got to check regularly.
That or guaranteed changed URLs when content changes are about the only way it's ever going to work in the HTTP client-server architecture, I think. Maybe there's a creative way to come up with guaranteed-changed-url that isn't as inconvenient in development for you, but most people find the current practices a pretty good spot I think.
Add a HTTP header on resource responses that is 1:1 with the deployment. When this header changes on any response, treat any associated resources in the cache as needing an If-Modified checkin.
Then add that header to a non-cached dynamic page, like the user's home page. When the browser checks this page and sees a deployment change, it'll know to check the static assets.
Clients SHOULD NOT issue a conditional request during the response's
freshness lifetime (e.g., upon a reload) unless explicitly overridden
by the user (e.g., a force reload).
The server is still going to serve up the resource when requested. This behavior is for the client-side of things (browser).