GitHub is aggressively caching raw.github, breaking many use cases
github.com
github.com
(Well, if we believe the statement "github engineer here". Of course every clown could write that, too)
Edit: Found it. Need to click on one of the achievements. Then the layout changes and it appears in the lower left corner under Organizations.
I think you missed the joke.
That's three things with an off-by-one error.
I think it’s open to interpretation. And I think it is valid to apply the saying to naming of variables, functions, and many other things.
See
https://skeptics.stackexchange.com/questions/19836/has-phil-...
Links to
https://www.karlton.org/2017/12/naming-things-hard/
Links to a bunch of places.
None of the ones I looked at seem to confidently say what was originally meant by the saying.
Either way, regardless of what the guy meant when he originally used to say it, it’s allowed to apply a saying to new situations.
Someone could get the cached of it, of course. But since it content addresses what it links to, it should avoid incoherent groups of cached data.
All in all it's a "certified hard problem", there's a bunch of domain specific nuance an HN thread couldn't hope to capture.
That said, I do find folks bend way too many things into fitting this pattern; such that I do not mean to be dismissive of the criticism.
As long as name client asks for (even if it is reference) is constant, it can have cache problems. Of course it makes it simpler as youc an opt to cache it so short that invalidation is less of an issue, but that's again working around cache invalidation
Or what if your CSS change just deletes some unused classes... it'd be fine for users to keep the old version until it expires. If you rename the resource you'll be causing a lot of users to wait unnecessarily. Not a huge problem, unless you're Meta or Google.
And so on.
When people say cache invalidation is hard it's best to believe them, because it is.
Except that they‘re colliding all the time with that tiny problem that the world isn‘t pure... People expect the same page at the same URL, with different content, tomorrow, so good luck choosing a different one than the one they have bookmarked and that ranks on Google.
Caching doesn‘t get any less hard by trying to define the problem away.
They have access to both GitHub and the raw service. I know there are usually all sorts of layers between that make interconnectivity logistically complicated, but am I wrong that at the top-level it’s that simple?
For URLs that return the latest entry there is no valid amount of time known in advance by GitHub unless they want to introduce mandatory publication delays. For URLs of specific change sets, they should never be corrected again and an infinite cache is pretty much valid unless a user overrides good git practices.
I think GitHub frequently misidentifies which scenario they are in and when they return 1 day for a current state URL users notice, while when they return 5 minutes for a permanent change set that gets a lot if traffic they lost network capacity.
Let’s leave this aside since it both has known client-side mitigations, and is not the cause of the issue that was posted.
The server side would require pushing any invalidation to (I imagine) whole tree of caches, which isn't exactly that hard if you plan for it from the start and have some way of upstream telling downstream file changes, but, well, they probably don't as I'd imagine they didn't expected people to pin their infrastructure to some binary blob on github that mutates
In fact, loading a file by name from a Git repo is rather expensive, and is definitely not the way their CDN should be keeping things in cache (gotta load ref, uncompress and parse commit object, uncompress and read tree object(s), just to get the blob's hash. Every one of those objects is both deflated and delta-encoded.)
I'm open to the idea that it's less computation to simply hash the files instead of deflate and decide, but my point was the hash is already calculated on the client when the changed file is added to the repo.
0) Cache invalidation
1) Naming things
5) Asynchronous callbacks
2) Off-by-one errors
3) Scope creep
6) Bounds checking
So this list is too bloated for the joke to work well, I think. Even before we talk about how off-by-one gets ruined this way.
Source: Implemented global TTL in our own caching DNS in front of kube-dns (which is horrible if you, among other things, have node containers with no DNS caching; i still have a pcap with 20000+ queries for A in s3.amazonaws.com in a 0.2s span) before coredns was a thing.
The CPU spikes were huge, but remained hidden for a long time due to metrics resolution. But eventually it got bad enough that clients ended up not getting responses.
Not saying they’re doing that.. just explaining the cache strategy.
An explanation isn’t a recommendation for you to go out and apply it to everything.
Source: founded and operated a cdn for 5 years of my life.
You can already get consistent views of raw.github by looking up the HEAD commit and requesting your files from that commit directly, though.
I would guess GitHub is slowly cutting down on people (ab)using it for free file hosting. Files that are hit a lot probably get significantly longer cache timeouts.
With sounds more like standard git use case than file hosting abuse
Rename the file to get what you want, am I the only one finding that a very strange approach?
> GET /file/latest HTTP/1.1
< HTTP/1.1 302 Found
< Location /path/to/real/file/hash.tar.gz> This response is only cacheable if indicated by a Cache-Control or Expires header field.
Blupblub;1 Blupblub;2 --> Blupblub;2 = Blupblub
Actually a good idea ;)
https://raw.githubusercontent.com/burekasKodi/repository.bur...
I'm fine with them caching "latest" tbh.
I would expect the headers to make this clear, however.
You can't know what the actual latest is, only some cached value. The actual value may have changed while the message is in flight
This is a beyond ridiculous statement. It is a BUG that you do not get the latest version of the file when viewing raw, not an error you made that you should address by having a filename driven versioning system.
If I'm using Git and GitHub, it's specifically to NOT have to deal with v1, v1.1, final, final_for_real, final_of_the_finalest, this_time_its_really_final.
Your suggestion to work around this GitHub bug is to essentially not use Git. Ridiculous.
You can use Git just fine without GitHub.
Maybe it’s a bit questionable to use the raw feature as a content host, but GitHub has intentionally moved pretty fair from plain old git (I believe people call this a “moat”)
I’m sure there are plenty of other players competing for users that would be happy to solve the problem for free.
And I don't see where you see that I'm not a paying user? This issue affects every repository, including ones where the user pays.
You don't have to take my word for it! https://github.com/orgs/community/discussions/46691#discussi...
And I’d assume the GitHub api has a way to get the hash for the head of a branch?
I don’t know, I guess the entire point of GitHub is to be able to obtain up-to-date files, so maybe they should just improve the caching.
https://github.com/orgs/community/discussions/46691#discussi...
Intel and Motorola processors compete:
How much is 2 + 2? - asks the Motorola
5!
Wrong answer.
But I was quick wasn't I?!I was surprised to see that sometimes that file was being requested millions of times a day due to attempts to loading it too early. With some optimizations I was able to greatly reduce that by 99%.
Having seen all those requests, I understand why github aggressively caches these files.
I agree that repo owners should be allowed to invalidate those caches upon an API request - or it should happen automatically on every change.
Was this a project that interacts with github otherwise? Because this seems very .. brittle.
I'm actually more amazed GitHub will give you a response in the 2xx range at all for these use-cases.
It's not like it's serving the wrong content for a specific commit ref, it's when you request the file as on a branch, and that branch has recently changed.
The issue was fixed within minutes, but the broken filter list is still being served.
Thankfully, it turned out to be a bug.
It's been for example impossible to make files completely disappear from an open source repo without deleting the repo or contacting their support.
https://github.blog/2022-09-13-scaling-gits-garbage-collecti...