We should add a cache to reduce latency tho…
We should add a cache to reduce latency tho…
Caches usually reduce load on the thing behind the cache. They can sometimes reduce latency in ways that matter, but often don't. For example if the p99 requests all miss in cache, then the cache won't help.
P(cache_hit) * cache_read_latency + (1 - P(cache_hit)) * cache_miss_latency
Assuming your scenario of 99% cache misses, that means you need a cache whose read latency is less than 1% of the cost of a cache miss. There are plenty of ways to design such a cache so that you still get a net performance benefit, even in your greatly exaggerated scenario.
One very simple example of an incredibly cheap cache that is almost certainly going to cost less than 1% of the cache miss latency is a bloom filter. Bloom filters can be tuned to be incredibly space efficient as well.
https://en.wikipedia.org/wiki/Bloom_filter
Please don't be so dismissive of people who answer questions by giving generally good and standard advice, just because you have some silly corner case that you want to play "gotcha" on.
EDIT: I recommend reading the Google SRE book or the famous "tale at scale" article.
I used to work at Google long ago in the platforms division, in fact I worked on the BigTable cache (among other things, mostly related to performance). It would be very sad indeed if today's SRE book dismisses caching as a vital and standard optimization strategy and instead plays all kinds of gotchas with potential candidates as I believe you're doing.
With that said, you article you linked does not in any way support your claim about caching, on the contrary it hardly discusses caching at all. It's as if you just wanted to dump a document you thought I wouldn't read as a way to be dismissive.
And yeah this is exactly my point. If you present a "standard and common" solution that isn't applicable to the question I actually asked, and if you blindly apply solutions without thinking about what problem is actually being solved, then that's bad in an interview.
Average latency is usually not the thing you want to reduce because it's not representative of what any actual user is experiencing.
With that said, to the best of my knowledge it is the same as my HackerNews handle, so you are welcome to find whatever info you'd like on that, but please understand your request is very creepy and akward.
In fact, one of the key projects that is being worked on is to improve caching to reduce latency both at median and at tail.
More common than I'd like to admit.
No doubt with a follow up on how the variable would be best named.
Caching is an easy thing to blurt out in an interview setting, but not all problem spaces benefit from caching and caching often isn't the only available solution, or the most ideal one.
I honestly can't decide which would be worse: a developer who literally doesn't know what a cache is or one who installs caches with bad invalidation policies.
Usually adding caches helps by reducing load on the service behind the cache.