They have access to both GitHub and the raw service. I know there are usually all sorts of layers between that make interconnectivity logistically complicated, but am I wrong that at the top-level it’s that simple?
For URLs that return the latest entry there is no valid amount of time known in advance by GitHub unless they want to introduce mandatory publication delays. For URLs of specific change sets, they should never be corrected again and an infinite cache is pretty much valid unless a user overrides good git practices.
I think GitHub frequently misidentifies which scenario they are in and when they return 1 day for a current state URL users notice, while when they return 5 minutes for a permanent change set that gets a lot if traffic they lost network capacity.
Let’s leave this aside since it both has known client-side mitigations, and is not the cause of the issue that was posted.
The server side would require pushing any invalidation to (I imagine) whole tree of caches, which isn't exactly that hard if you plan for it from the start and have some way of upstream telling downstream file changes, but, well, they probably don't as I'd imagine they didn't expected people to pin their infrastructure to some binary blob on github that mutates
In fact, loading a file by name from a Git repo is rather expensive, and is definitely not the way their CDN should be keeping things in cache (gotta load ref, uncompress and parse commit object, uncompress and read tree object(s), just to get the blob's hash. Every one of those objects is both deflated and delta-encoded.)
I'm open to the idea that it's less computation to simply hash the files instead of deflate and decide, but my point was the hash is already calculated on the client when the changed file is added to the repo.
0) Cache invalidation
1) Naming things
5) Asynchronous callbacks
2) Off-by-one errors
3) Scope creep
6) Bounds checking
So this list is too bloated for the joke to work well, I think. Even before we talk about how off-by-one gets ruined this way.
I think you missed the joke.
That's three things with an off-by-one error.
I think it’s open to interpretation. And I think it is valid to apply the saying to naming of variables, functions, and many other things.
See
https://skeptics.stackexchange.com/questions/19836/has-phil-...
Links to
https://www.karlton.org/2017/12/naming-things-hard/
Links to a bunch of places.
None of the ones I looked at seem to confidently say what was originally meant by the saying.
Either way, regardless of what the guy meant when he originally used to say it, it’s allowed to apply a saying to new situations.
Someone could get the cached of it, of course. But since it content addresses what it links to, it should avoid incoherent groups of cached data.
All in all it's a "certified hard problem", there's a bunch of domain specific nuance an HN thread couldn't hope to capture.
That said, I do find folks bend way too many things into fitting this pattern; such that I do not mean to be dismissive of the criticism.
As long as name client asks for (even if it is reference) is constant, it can have cache problems. Of course it makes it simpler as youc an opt to cache it so short that invalidation is less of an issue, but that's again working around cache invalidation
Or what if your CSS change just deletes some unused classes... it'd be fine for users to keep the old version until it expires. If you rename the resource you'll be causing a lot of users to wait unnecessarily. Not a huge problem, unless you're Meta or Google.
And so on.
When people say cache invalidation is hard it's best to believe them, because it is.
Except that they‘re colliding all the time with that tiny problem that the world isn‘t pure... People expect the same page at the same URL, with different content, tomorrow, so good luck choosing a different one than the one they have bookmarked and that ranks on Google.
Caching doesn‘t get any less hard by trying to define the problem away.