However, my understanding is the classic algorithms for caching have yet to lose to the new machine learning ones. Has that changed?
However, my understanding is the classic algorithms for caching have yet to lose to the new machine learning ones. Has that changed?
More importantly, I know this is claimed a lot. But I thought the last few explorations of the idea I saw did not actually see fancier algorithms win. Indeed, the best strategy from my memory was random spreading of the data with an almost random replacement strategy. I think some of the win there was just the low overhead of the bookkeeping, but it was still one of the better bets.
(This is all off the table, of course, if you know what the access pattern will be. Then, by all means, set things up accordingly.)
With Spiral, we were able to approach this top down as a classification problem.
e.g.
If you have a cached query for "Friends that liked my post", the Spiral classifier quickly learns that "Post Last Viewed At" or "Post Last Modified At" is not relevant to this via the feedback from the caching code.
Pre-spiral, this was expressed via a curated blacklist/whitelist which had to be recreated if the query characteristics changed.
That said, I think I see where I was mistaken in thinking that was an odd example. It was literally the example. Not just a random pedagogical one.
To that end, thanks for sharing! Cool stuff.