Forecasting at Uber with RNNs
eng.uber.com
eng.uber.com
One of the interesting points, that is often overlooked in ML is model deployment. They mention tensorflow, which has a model export feature that you can use as long as your client can run the tensorflow runtime. But they don't seem to be using that b/c they said they just exported the weights and are using it go which would seem to imply you did some type of agnostic export of raw weight values. The nice part of the TF export feature is that it can be used to recreate your architecture on the client. Bu they did mention Keras too which allows you to export your architecture in a more agnostic way as it can work on many platform such as Apples new CoreML which can run Keras models.
The model has to be converted to a format for CoreML, which does not work with the Keras 2 API yet: https://pypi.python.org/pypi/coremltools
1 biased perspective I have here: Infra is often a different team from data science. They don't always do the deploying. Beyond "some sort of serving thing" the data scientists might not necessarily know about what's being deployed. This is not true at every organization and there are exceptions. This is typically true of most companies we sell to though. There are usually ML platform teams that do the "real" deployment (especially at sizable scale)
Another characteristic of production is it's "boring". "Production" is a mix of databases to track model accuracy over time, possibly microservices depending on how deployment is "done". Characteristic ways of giving feedback when a model is wrong, experiment tracking and model maintenance among other things.
A lot of these things are typically very specific to the company's infrastructure.
The "fun" and "sharable" part that people (especially ML people) is usually related to "what neural net did they use?"
The other thing to think about here: "production" isn't just "TF serving/CoreML and you're done" there's typically security concerns, different data sources,.. that are often involved as well that might be specific to a company's infrastructure. There also might be different deployment mechanisms for each potential model deployment: eg: mobile vs cloud.
Grain of salt sales pitch here: We usually see the "deployment" side of things where it's a completely different set of best practices that happen to overlap with data scientists experiments. This includes latency timing, persisting data pipelines as json, gpu resource management, kerberos auth for accessing data, managing databases and an associated schema for auditing a model in production (including data governance), connecting to an actual app/dashboard like the ELK stack,..
TLDR: The deployment model would be its own blog post.
Thanks for your interest!
* the weights * and the architecture
But the very first sentence on that page says: These APIs are particularly well-suited to loading models created in Python and executing them within a Go application
This is exactly what Uber is doing.
[0] https://tensorflow.github.io/serving/ [1] https://github.com/tensorflow/tensorflow/blob/master/tensorf...
You probably want to do experiments with multiple model variants, or teak your model and fine tune from deployed weights. To do that you need a way to recreate it from layer-level objects instead of the add/reshape operations Tensorflow and its kin store internally.
All that said, the company that outsources (and that's really what you're proposing) such a core component of their business is probably taking on way too much risk.
One thing I thought, it would be really convinient if Uber could amortize their surge pricing over the month/year. In order not to hit customers with unexpected rates and essentially offer a flat predictable fee over the whole period. Problem with that is you can't really plan the future demand to calculate how much you need to save/dip. Could an auction house help to hedge the bets?
Surge is only really solved once autonomous vehicles can be preemptively positioned near demand.
Also, I wonder if they checked how a feed-forward NN that operates on the contents of a sliding window (e.g. as in the first approach above) compares with their RNN results. I am curious about this, as it would give us a hint whether the RNN's internal state encodes something that is not a simple transformation of the window contents. If this turns out to be the case, I'd then be interested in figuring out what the internal state "means"; i.e. whether there is anything there that we humans can recognize.
[edited to increase clarity]
A feed-forward NN wouldn't do much because it doesn't hold a state variable which you need to be able to understand context in time series data. There are probably some pieces of the state that you'd be able to interpret but the majority of it would mean nothing to us.
There are undoubtedly things that machine learning is right for, however to me it seems like it's become a buzzword more than anything else.
Maybe O'Reilly should release a book on RDD?
The lifetime compensation of developers isn't just tied to how much salary they have at the moment. Getting into a dead-end and not developing your skills will definitely set you back over the long term. Like, there's a reason people will pay you a premium for working with COBOL. So there's a very rational pressure to get better total compensation out of a role by choosing resume-developing tools and technologies.
On the other end, organizations mostly just care about getting the task done. They've got the choice of doing it in a boring way and paying lots of money for developers who don't want to grow their resume, or indulging in the developers fancies and getting it done more cheaply.
tl;dr - resume driven development is a way for companies to pay for projects with "you'll get experience".
Also, some actual benchmarking would be great. Say, against Facebook's Prophet (which also deals with covariates and holiday effects).