Caching Antipatterns
hidefsoftware.co.uk
hidefsoftware.co.uk
Responses to a request are usually dependent on a wide variety of different internal states: properties of the user, cookies, state in one or more databases often across multiple tables, or state in multiple dependent services.
A cached value is a snapshot of a calculation over this internal state: but it needs to be invalidated when any of the states that it is composed from changes. What used to be a simple local calculation now becomes a global concern; anything that modifies any of the underlying states needs to be aware of the global state in the cache, and evict or invalidate cached items that were dependent on the underlying state.
No longer is local reasoning enough. You need to have a global understanding, possibly across multiple subsystems.
Caching worthy of the name - caching things that are complex to calculate in time or space - inflicts serious architectural harm on a software system. It impairs maintenance and understanding of the system, and will cause bugs when people who don't have sufficient global understanding make local changes.
It's not all bad. Caching static assets, or a small well-understood layer (like a disk cache), or designing calculations in such a way that dependencies have immutable mappings from identifiers to values - there are ways to make caching work. But it needs to be done very carefully in a large system or you risk long-term damage.
> ... it needs to be invalidated when any of the states that it is composed from change ... need to have a global understanding, possibly across multiple subsystems
:)
There are ways of using caching that are relatively simple to reason about that ought to be evangelized a bit.
For example, many caching proxies expose quite a bit more functionality than just caching, such as replacing placeholders in server-side rendered pages with transclusions (by the proxy) from another URL (which can be part of the same web app). This lets you compose server responses from fragments that have different caching rules, neatly side-stepping many of the issues mentioned in the OP as well keeping global state disentangled.
For example, you can cache most of a rendered web-page with a generous TTL so that the cache only queries the web app behind it once an hour or even less often, and have the proxy insert the portion of the page that must be current, which is also cached, but with a TTL of <1 second and relying on if-modified-since to speed up the usual case where nothing has changed.
The resulting setup is easy to reason about, doesn't have any particular gotchas (unless you are doing a lot of A/B testing), and will absolutely solve the problem of your site slowing to a crawl because 50k impatient users are obsessively hitting Ctrl-R over and over on your homepage at the same time every Monday morning to check to see if that one thing has changed yet. Notably, this type of setup works without having to beef up with more server instances to handle the load, or having to accept a slightly stale homepage being served for several minutes, or even accepting a longer response time for the very first user to hit the server after the content does change and the cache is invalidated and repopulated. All you've done with the cache is eliminate wasted cycles at the cost of adding less complexity to the code than you need for feature-switching.
This approach of fast composition (transclusion or even simpler concatenation) of separately cached values avoids the problem of global state entanglement because the final transclusion or concatenation is simple and fast enough to not (or hardly) bother with caching the final result at all.
Not listed: persistent caches. A friend of mine ran a cache for a CDN. Their caching solution mushroomed in complexity to satisfy the legal obligation of taking down illegal content because their cache persisted to disk, so when servers came back up from maintenance they had to check in with a master revocation list. This made operations somewhat complicated because the system was failsafe - an old server would refuse to startup if it detected that it was older than the size of the revocation list (a 1 week KAFKA queue). Lots of horror stories about zombie resurrected content being served before this was implemented.
Somewhat alluded to: over reliance on cache. My former job on a high traffic website was behind a fantastically fast high hit rate cache for almost the entire site, necessary for the scale they were operating at. The trouble was the cache shielded the system from so many requests it became a single point of failure, so if the cache was flushed then the system would become quickly overloaded and wouldn't recover the hit rate for an hour or so. We had to invest a lot of time and energy in cache replication systems, consistent hashing, and related systems - mostly careful choice of existing systems. Shoutout to mcrouter by Facebook.
I guess you are alluding to the special case of single-page web applications prefetching resources at startup? Here the article also holds true: if you can you make your dependencies fit for purpose - making all resources small enough to be fetched as necessary - or maybe you have to admit defeat to a dependency you don't control: the user's internet connection.
Or were you thinking about a different scenario where clients are justified in caching on startup?
In e.g a game scenario we might call this "precomputing" or "preloading" but it's really equivalent to memoization and caching of any other kind (you load the whole level rather than a part, you precompute trig tables rather than wait for the first access of sin(123) and so on).
I make a thick (non-game) desktop app where quite a lot of data is computed on startup because 500ms spent during a splash screen is ok, but a 500ms delay to a user input is not.
Web or database requests - probably not - so I can't imagine any scenario involving the web.
We have frameworks that can handle 1000s of requests per second on regular hardware, that enough for a single machine to run 99.9% of the sites in the internet.
Take WordPress for example. It's so dog slow that caching is basically required, since even powerful servers can only manage 10-20 pages per second.
If WordPress was running Java or Go instead of PHP caching would be totally unnecessary
We have frameworks that can handle 1000s of requests per second on regular hardware
Well, that's the server farm for helloworld.com sorted then.