ETag and HTTP Caching
rednafi.com
rednafi.com
— Which HTTP code I should return for my API? I already used 404, 403, but I need another one. Damn, HTTP is so old and it makes no sense.
— You can't use HTTP codes like that Bob, they're not a free choice. They're for the protocol, not for your app.
— Let's look at the list. Hm... "412 Precondition Failed". Hey, it sounds nice. It fits to my use case. I'm gonna document it. It means the account is out of balance.
— What is this garbage? Please read the spec. This is going to make our API gateways, CDNs, everything go crazy. Can't let you move on with this PR.
— Look. I documented it, made an enum with the code, it's clean. I'm an experienced REST developer.
— It... it doesn't work like that Bob. Please, read the spec.
— Hey, got enough approvals, "412 Account Out Of Balance" it is! It passes the tests.
For each dev that knows proper HTTP, there's 10.000 Bobs.
https://developer.mozilla.org/en-US/docs/Web/HTTP/Status/412
304 means "you're good, your cached version satisfies the conditions"
412 means "you're not good, your cached version does not satisfy the conditions"
412 usually applies to modifications, but it could be for reading too, in the case of ranged requests (getting a specific range of bytes from a large representation). See the "Range" header.
These are all interconected. The headers, the codes, etc. They are very useful for caching and can save a lot of bandwidth. Browsers and CDNs use them extensively. Server-to-server communcation could use them as well, but I haven't seen popular implementations (let's say, a web framework that provides abstraction over these mechanisms).
For example, 404 implies "No indication is given of whether the condition is temporary or permanent". Cache invalidation headers don't apply to this code because a 404 means there's something about the _resource_ and not only the _representation_ that could not be found. That client cannot cache the 404 result, not even for a fraction of a second.
404's brother 410 implies "This condition is expected to be considered permanent.". A client that gets a 410 can cache that result, never needing to reach the server again. It means it's gone forever. That client can decide to never look up that URI again.
Very often, "400 Bad Request" is the best HTTP you can use if you are not sure what to use. Then, describe what the error means using other HTTP components and/or the response body.
HTTP can be very simple. GET (ask for a representation) and POST (send a representation) as methods only. 200 (success), 400 (client error) and 500 (server error) as response codes only. It's the best way to start, then move to more elaborate protocol features as you learn.
RFC-7807, Problem Details for HTTP APIs[0]
From the introduction:
HTTP [RFC7230] status codes are sometimes not sufficient to convey
enough information about an error to be helpful. While humans behind
Web browsers can be informed about the nature of the problem with an
HTML [W3C.REC-html5-20141028] response body, non-human consumers of
so-called "HTTP APIs" are usually not.
This specification defines simple JSON [RFC7159] and XML
[W3C.REC-xml-20081126] document formats to suit this purpose. They
are designed to be reused by HTTP APIs, which can identify distinct
"problem types" specific to their needs.
HTHAt least at that point it’s clearly not HTTP anymore, and it’s better than pretending to be compliant like your Bob. But something dies inside me whenever I have to work with it.
Returning 200 for an error makes no sense. Having 400s and 500s is the simplest way to have observability over protocol behavior (think logs, error rates, etc). If you use all 200s, you'd have to re-implement observability by yourself, so you lose simplicity that you gained by ignoring those statuses.
It's the same thing with caching stuff. You could implement those outside the protocol, but then you'd be writing your own protocol (trying to be smarter than decades of engineering efforts).
- 404 might mean the URL was wrong
- 405 might mean swapping POST with PUT
Both of these are simple programming errors. You might as well have used the wrong host or protocol for that matter
Fortunately, for a lot of my API-type work, I also get to not care. I don't want some smart cache to think it knows how to cache my responses or anything and don't care about the sort of infrastructure that thinks it understands HTTP doing anything with my request.
200 {"error": "..."} is not necessarily invalid from this point of view, either. 200, the request was successfully processed and the successful result of that request as far as HTTP is concerned is an error. There doesn't seem a great need to tell HTTP there's an error, HTTP doesn't really care. Telling the browser there's an error has some marginal utility, but if it's an API and there's no browser involved, that doesn't matter much either. The 200 isn't going to fool it into thinking it should put the error into the history or whatever.
I've also learned to avoid getting too fancy with the codes. You will invoke some weird behaviors from systems you didn't even know cared about your connection. 200 {"error": "..."} may seem "wrong", but it is also generally safe. It will do what you expect.
It might be nice to live in a world where there are HTTP error codes that are suitable for everything I need, instead of a big pile of useless codes for abortive standards that never came to be and things nobody uses, and an underspecified set of codes for the things I actually want and use, but there's no point pretending that the standard is something other than it is, and as it stands now, a lot of times the HTTP result code is almost useless.
You can get more specific, and for REST style APIs the correct specific HTTP status code is usually apparent for both successful and unsuccessful requests, but 2xx/4xx/5xx is simple and should be trivial to determine for anything you are using HTTP for even if it's not REST-like.
However, while your mileage may vary, I end up getting the same complaint from the users either way. Even when my 400 contains an exact reason why the input is incorrect.
Granted, on the one hand, this can be fixed on the individual level, but on the other hand, it's the same effect writ small that when writ large makes the response codes nearly useless, so this post is maybe half cathartic grousing. I can't push caring about response codes. I can document it, I can yield detailed errors, and I can be as careful as I like, but this is a "it takes two to tango" situation and at scale, on average, the other end doesn't want to tango.
A lot of codes get real usage though. CDNs use ETags extensively to save bandwidth. It can save a lot of server time.
- https://developers.cloudflare.com/cache/reference/etag-headers/
- https://techdocs.akamai.com/property-mgr/docs/val-entity-tag
Varnish, HAProxy, Apache mod-proxy, nginx all can do similar things. Some of them can do this even if you always return 200 (by having rewrite rules and so on). It is often better to leave this kind of work to some upper abstract layer. Some of thse codes are only applicable in a layered system (502, for example, often seen when nginx can't reach a backend application), so they seem useless to developers, but they're not.For APIs, other stuff uses those codes. Tools like DataDog and NewRelic will get better if you use generic 400 and generic 500 for client and server errors respectively. You can make them work with 200s and a little configuration though, but it's extra work.
If you never needed any of this, it's better not to use it.
The non-browser web, they approach useless. Which I'm not happy about and not celebrating or advocating for. It's just how it is.
I've used SOAP and JSON-RPC, both of which (at least in many implementations) send RPCs as HTTP POST requests and receive 200 responses with any error messages in the body. They're just tunneling over HTTP. It's not necessarily wrong, although I'm convinced that leveraging the HTTP verbs and error codes with REST is a fundamentally better design for the use cases I've seen.
I think I've seen everything on that scale.
On one end: { status: 200, error: InsufficientFunds }
On the other: - "let's use 409, it perfectly fits our use case" + "but we don't have useful error codes for all the other domain errors."
Let's assume your resource is "/account/1234/withdraw-availability"
It's a hypothetical endpoint you can GET to know if you can withdraw money. You hit it, and the request is sucessful (the server understood and will inform you whether withdrawaw is available or not).
Let's assume your resource is "/account/1234/withdraw"
This other hypothetical endpoint you can POST a request for money withdraw. Returning 200 here means the server understood and processed your request, so a 200 that does not withdraw makes no sense.
The same endpoint could also return a success "201" accepted (the server understood the request, but it is not processed yet). In the body, there would be a link for "/withdrawaws/48957987593845983475/status", a resource specific to this future processing, which you can GET later (maybe 1ms later, within the same socket). This GET could also return a 200, saying that such withdrawaw was not possible (the server sucessfully understood the request and will inform you about the status of the withdrawaw).
For this modeling stuff, the Roy Fielding dissertation about REST is more enlightening than the spec. The spec is still needed though.
- The first line of an HTTP request has its own format (space-separated-ish), and mixes method and URL path.
- The URL path in that first line has its own format and weird escaping, and mixes one path with zero-or-more key value pairs.
- The headers have their own format.
- The body has its own format.
Most (all?) HTTP libraries for clients and servers abstract all that mess away into a neat object that could be easily represented in JSON (or bencode if you want something simple-ish while supporting binary data), but it's like using a nice program while knowing it's written in C[1].
Of course, JSON has its own problems, but in that case at least we only need to deal with the problems of one format that supports with proper nesting, not 4 different weird formats masquerading as one.
[1]: Disclaimer, I don't like Rust, so don't take this as a RIIR thing.
EDIT: To be clear, I don't agree with how Bob misuses HTTP in your example. I just find it sad that we're locked into this weirdly complex protocol.
- https://www.w3.org/2000/12/drm-ws/pp/connolly/slide8-0.html
- https://www.w3.org/DesignIssues/Axioms.html#opaque
For the protocol, any structure or meaning in URLs is irrelevant. It is one more example of mixing application domain with protocol level stuff.HTTP header format goes back to ARPA times. Email reused them, so did HTTP.
- https://datatracker.ietf.org/doc/html/rfc822#section-3.2
Many of these choices are there for backwards compatibility and reuse. I can totally understand why.
An HTTP body has no predefined format. It can be anything. It can be a stream (HTTP into WebSocket upgrade, for example). It is the media type that defines how the body should be interpreted.
HTTP requests are meant to be used before they are fully transmitted, and are formtted in a way to leverage socket communication. JSON, on the other hand, needs the whole document to be read before it can be safely interpreted.
HTTP has more moving parts, but it also does so much more. These two aren't even comparable, they're not in the same layer.
I understand the urge to "improve" on all of this "legacy", however, one must consider how much was built upon these standards and if there's anything real to gain by changing them.
You can totally drop accept headers if you write your own client and server implementation. HTTP works fine without them. The web as a living organism, not so much.
But hey, we don't need content type negotiation, right? XML will reign forever, mp3 is the ultimate audio format. It's not like new codecs and document types appear all the time and some kind of underlying architecture has to reserve space for that kind of change.
Using `?format=json` is not offensive. It won't mess up some cache layer like improper status code semantics, so I don't really care that much about these if I see it. I wouldn't block a PR on that.
The overall web on the other hand, is supposed to be made of many different client and server implementations. Your browser still relies on Accept headers for displaying images, detecting language, uncompress gzipped responses, resuming paused downloads, showing that JSON API in a nice UI when you open in a dedicated tab, so many things.
To me, it makes sense to leverage the same content negotiation ideas for home grown stuff, even if only one content type is being used.
It starts being a problem if you're working on microservices, and one of them uses `.json` while other uses `?format=json`, made by different teams. The standard is the obvious solution. Instead, they'll either create inefficient clients full of complexities or fight until one of the workarounds prevail. So much easier to follow the standard.
I understand that the people at the time (presumably) did it the best they could with the information and knowledge they had at the moment, while keeping backwards compatibility.
I only know a bit about HTTP/0.9 and can see how we moved from that to what we have today. I just find the current situation sad.
Like when something only supports ASCII, or only supports IPv4, or assumes I only have 1 CPU thread. Or like when some binary file is encoded in base64 before being sent over the network, only to be decoded on the receiving end. Stuff like that.
But I only feel that way thanks to having the huge benefit of hindsight and modern technology.
HTTP/3 was published in 2022.
Anything can be a footgun if misused.
ETag: W/"750"
This also means the API can just check the revision number and avoid pulling out and decompressing some of the larger payloads that are stored there and the implementation is absolutely minimal. It's a great standard.https://levelup.gitconnected.com/no-cookies-no-problem-using...
I am using that approach in my project https://github.com/claceio/clace. It removes the need for a build step while making aggressive static file caching possible.
And so you could find the same file on IPFS if anyone served it there, as the content hash in the url tells you what to look for.
Even though on my side I’ve done this all completely manually, so much so that it’s literally just me calculating the IPFS hash on my machine one time and then having symlinks with those content hashes so that /ipfs/ directory on my sites contains content that is served by their IPFS content hash, even though my server does not run an IPFS node or anything.
A very interesting side effect of this is that one time I loaded one of my pages, the web browser actually picked up on the pattern and offered to load those files over actual IPFS!
Browsers now offer to load over IPFS?
I was wondering how you can trust an IFPS gateway, does does browser verify the file is legit using some checksum? Maybw subresource integrity supports IPFS content hashing or something? How does it generate cid anyway ?
https://docs.ipfs.tech/concepts/content-addressing/#cids-are...
How would you use SRI here to verify the cid (and not an additional out-of-band hash) make sure the gateway isn’t returning some crap to, say, inject malicious JS?
Using the example file from your link, I would host it as something like
https://mywebsite.example.com/ipfs/QmPK1s3pNYLi9ERiq3BDxKa4X...
And this was enough for that particular browser I was using to recognize that this file can be attempted to be retrieved directly from IPFS
But either way, I run Brave with a local IPFS gateway on my laptop. So I think when it said that it could retrieve those files via IPFS for me, it meant in my case that those files would be retrieved from actual IPFS and not via a public IPFS gateway hosted on the web
So I wrote https://github.com/infogulch/xtemplate to scan all assets at startup to precalculate the hash for templates that use it, and if a request comes in with a query parameter ?hash=sha384-xyz and it matches then it gives it a 1 year immutable Cache-Control header automatically. If a file x.ext has a matching x.ext.gz/x.ext.zst/x.ext.br file then (after hashing the content to make sure it matches) client requests that support it are sent a compressed version streamed directly from disk with sendfile2. I call this "Optimal asset serving" (a bit bold perhaps).
It seems like generating a weak ETag would take more effort since you'd need to either ensure consistent ordering and formatting of the JSON, or generate the Etag on the content before converting it to a JSON string.
> You could make the `calculateETag` function format-agnostic, so the hash stays the same if the JSON format changes but the content does not. The current `calculateETag` implementation is susceptible to format changes, and I kept it that way to keep the code shorter.
They seem to agree, a true weak ETag implementation would probably be trickier and require more code :P I'd be fascinated to see how that might work in practice, though.
It's good to know about this option for handling in the frontend if a system returns one though.
In the example request, the server still has to do all of the work generating the page, in order to calculate the ETag and then determine whether or not the page has changed. In most situations, it's simpler to have timestamps to compare against, because that gives the server a faster way to spot unmodified data.
e.g. you get a HTTP request for some data that you know is sourced from a particular file, or a DB table. If the client sends a If-Modified-Since (or whatever the header name is), you have a good chance to be able to check the modified time of the data source before doing any complicated data processing, and are able to send back a not modified response sooner.
Obviously doesn't work if you're using ETags on dynamic resources, but works well for non-dynamic, but unpredictably frequently changing resources.
I would agree that you shouldn't be doing signifiant calculations -- like checksumming the contents of a file -- to determine the current etag value.
Also, the excessive semantics of timestamps is a negative. It adds complication which sometimes trips developers up but doesn't improve the caching.
[1]: https://en.m.wikipedia.org/wiki/Fowler%E2%80%93Noll%E2%80%93...
Hash Name, Bandwidth, Small Velocity, Quality, Comment
XXH3 (SSE2), 31.5 GB/s, 133.1, 10,
FNV64, 1.2 GB/s, 62.7, 5, Poor avalanche properties
[0]: https://github.com/Cyan4973/xxHash
Also ETag is exactly the kind of thing non-cryptographic hashes are meant for, but if you can't convince them Blake3 is a very fast modern cryptographic hash function.
If you're implementing weak validation, then you might need to preprocess the payload before running it through the hash function. For example, if your payload is JSON and you want to make it format-agnostic, then you'll need to normalize the payload and then compute the hash.
In either case, the hashing algo probably doesn't matter as much.
There was still one thing that surprised me a bit (but also makes sense). Images are fetched only once per page load in my testing. If an image with 60sec of cache is loaded, then removed by JS and added back after 2 minutes, then the browser will reuse the image from the initial load.
1) I (AWS CloudFront) supply ETag and If-None-Match header; I can see that header in responses.
2) browsers sometimes do respect that (once) I see 304 in responses, but 99% of the time they don't include ETag/If-None-Match in requests and thus I never get 304 responses (albeit AWS CloudFront, resource, data — nothing changed) and instead they perform some other caching and reload whole resource again with TTL that does not seem to come from my headers, totally disregarding ETags/If-None-Match logic.
for videos it is even worse. unless you set `preload=none` in html, Safari, Firefox, Chrome will have all different policies trying to preload all videos on screen ignoring lazyload html tags. worse of all, caching does not work well, and videos will be attempted to be loaded almost every time and ETags/If-None-Match totally ignored.
actually caching is happening, but it does not follow ETags or Caching policy headers that backends return, instead some in-browser internal caching policy being run
I’ve used that flowchart too many times.
my conclusion in the end, this was intentional by browsers. they give priority to the their internal caching policies for performance or user-experience :/