Spiral: Self-tuning services via real-time machine learning
code.fb.com
code.fb.com
However, my understanding is the classic algorithms for caching have yet to lose to the new machine learning ones. Has that changed?
More importantly, I know this is claimed a lot. But I thought the last few explorations of the idea I saw did not actually see fancier algorithms win. Indeed, the best strategy from my memory was random spreading of the data with an almost random replacement strategy. I think some of the win there was just the low overhead of the bookkeeping, but it was still one of the better bets.
(This is all off the table, of course, if you know what the access pattern will be. Then, by all means, set things up accordingly.)
With Spiral, we were able to approach this top down as a classification problem.
e.g.
If you have a cached query for "Friends that liked my post", the Spiral classifier quickly learns that "Post Last Viewed At" or "Post Last Modified At" is not relevant to this via the feedback from the caching code.
Pre-spiral, this was expressed via a curated blacklist/whitelist which had to be recreated if the query characteristics changed.
That said, I think I see where I was mistaken in thinking that was an odd example. It was literally the example. Not just a random pedagogical one.
To that end, thanks for sharing! Cool stuff.
> Today, rather than specify how to compute correct responses to requests, our engineers encode the means of providing feedback to a self-tuning system.
"encod[ing] the means of providing feedback to a self-tuning system", got it, very cool!
But don't they still have to "specify how to compute correct responses to requests"?
if (conditionA && conditionB && !conditionC) cache_it()
to
hey look, an item with featureset X is cacheable while one with featureset Y is not.
This reflects in the API which is just two calls predict() and feedback()
this simplifies the integration code and is easily debuggable even in the face of changes.
I tried something kinda similar to help with tuning data engineering jobs and pipelines for performance and costs. But, it turned out to be a fruitless activity because there were too many variables that affected performance. I’d produce some models that seemed to be marginally effective. But, after code changes, configuration changes, changes to data input sizes, the models quickly became stale and ineffective.
(Isn’t all ML about good labels and features? :-) )
Structured ML systems require you to provide (ideally) unambiguous examples of expected behavior. In case of Spiral (or any other online learning) such examples need to be generated automatically. In our experience this part took a good amount of effort: distributed systems issues (aka race conditions and transient bugs in remote systems) made automatic generation of “clean” examples difficult. Once the bulk of these problems were addressed the system began to operate very smoothly. Specifically, it adapted to changing conditions very well.
Spiral is designed to be a drop-in replacement for hand-coded heuristics. In other words, if you had a somewhat working tree if-else statements that specified your image caching policy (if size<100k and type==jpeg..), you should already have an idea for what features to use. There is a bit of work involved in translating these features into the form suitable for classifiers in Spiral. For example, if a classifier is using binary features, the file size feature would need to be quantized (123kb -> “100-200kb bucket”). While this type of work requires forethought and effort, runtime cost of running this classifier is very low.