Michelangelo: Uber’s Machine Learning Platform
eng.uber.com
eng.uber.com
But when you have that many smart people under one roof, you have to keep them busy. So you get these kinds of reinventing-a-beautiful-wheel projects. Not knocking the quality, just knocking the growing distraction and temptation to build in-house rather than use existing tech.
LivingSocial had such good quality output, but they were all over the place. They had to keep all those productive people busy after all. And the best way was to keep churning out new products.
Even considering crunch-time and potential diversification/expansion, full-time employment of non-critical engineers is obviously inefficient. And, as Uber proved, politically toxic and consequently harmful to the productivity of the business.
More pressing is the existential tension underlying this quandary. That our economy will generally and increasingly forgo the skilled and unskilled labor provided by our society.
Framed concretely: one day, we may define a "10x engineer" as an engineer who can send ten other engineers to the soup kitchen.
I'm still looking for a well-thought-out solution.
That is, suppose I'm willing to do the following:
* Send said service a firehose of data
* Write code for feature selection, data cleanup, transformations, etc. -- i.e. the sort of code you'd expect to find in a data scientist's Jupyter notebook.
* Specify the type of ML model and parameters I'd like to experiment with.
Is there a service that handles data warehousing, serving up predictions via a REST API, deployment of new models, reporting model performance over time, and general scaling issues? I know it's possible to set up a large portion of this using AWS services, but ideally, it'd be nice to have something with a Heroku-like ease-of-use, but for ML.
* You can send your firehose to BigQuery, and write Javascript UDFs to transform the data.
* You can use a service like DataRobot to automatically train a bunch of models.
* You can deploy the trained models using a serverless platform like Lambda.
Nothing that's one-size-fits-all, but watch out for trends in that direction as more ML workflows get standardized.
The estimated time is always less than the actual time, but every minutes the estimated time re-adjusts. By the time it's actually delivered it's about 15 to 20 minutes more than the original estimate.
Note: I'm just complaining because the article uses User Eats as an example and as a service I use often I wonder if this ML Platform is used at all...
It takes 50 mins to get your delivery but to make you order they need to tell you its going to take 30 mins. ML also tells them that the 20 minute discrepancy won't stop you ordering next time.
Cycle continues.
But.. there are existing solutions in this area. Notable Luigi[1] (from Spotify) and Airflow[2] (from AirBNB) both seem to have a lot of overlap with this.
I'm most familiar with Luigi, and it does many of the things that are listed here.
Some of the model visualisations look pretty nice, and don't come out-of-the-box in the other platforms.
So I'm not really sure what makes this unique.
And feature extraction and modelling is my main usecase. It works really well, across Spark, Scikit and R, saving data in HDFS.
Some examples:
https://blog.dominodatalab.com/luigi-pipelines-domino/
http://blog.richardweiss.org/2016/10/13/kaggle-with-luigi.ht...
A quote:
We provide containers and scheduling to run regular jobs to compute features which can be made private to a project or published to the Feature Store (see below) and shared across teams, while batch jobs run on a schedule or a trigger and are integrated with data quality monitoring tools to quickly detect regressions in the pipeline–either due to local or upstream code or data issues.
That certainly sounds like you write code to run inside the containers. The deep integration with the standardized feature store etc sounds nice, but not radically different.
we created a DSL (domain specific language) that modelers use to select, transform, and combine the features that are sent to the model at training and prediction times. The DSL is implemented as sub-set of Scala.
So this is pretty much the equivalent of Spark Dataset API, or maybe the RFormula stuff in Spark ML[1] except in Scala right?
[1] https://spark.apache.org/docs/latest/ml-features.html#rformu...
Pachyderm.io (I'm one of the founders)
>For important model types, we provide sophisticated visualization tools to help modelers understand why a model behaves as it does, as well as to help debug it if necessary
How is this done - what serialisation format (pickle, pmml,etc) is being used? And more importantly, how are you tracing into the model ? It looks like this is a Spark based framework
Sure different frameworks handle different uses better than others, but it doesn't change the fact that in a given genre (machine learning, web frontend, ORM, etc.) that there is a lot of duplication going on.
If you want to look cool and impress people, make a tool that doesn't exist, develop a new algorithm, or contribute a great feature to an existing framework.
Do you simply have very high expectation of everything in this world?
I just picked this thread as a starting point Uber may have a legit need for a particular framework.
My main point is there is no need to make frameworks for frameworking's sake. Purposeful OSS collaboration is just as good, if not a better way to gain influence than a PR splash.
Of course that assumes that the goal is to make cool things, which seems loosely connect with wooing investors at best. (Again in general, not picking on Uber.)
Not only does it seem wrong to judge things people post here that way, anyone out there doing good work knows that the path to doing meaningful work begins with "useless projects."
It's not for me to say where you apply your time and from what you get your enjoyment from, but I do believe there are objectively more important problems that could be solved. There are plenty of problems out there that need to be solved, and engineers here have the capability to solve them -- I think it's fair to say we're more well equipped than most to solve some of the largest problems humans face.
Here's a quick list of problems that are worth solving (in my opinion), vs. things that are figuratively useless:
- Rising global temperatures will disrupt agriculture and food supplies, how will we farm when desertification consumes the croplands?
- Thousands of people walk by homeless people on the street every day, if a small portion of those people stopped and engaged with the homeless, how many lives could be improved?
- Children in third world countries struggle to receive proper education, how can we reach them and improve their education?
- Governments around the world are failing to represent the needs and desires of their people, what can be done to help governments succeed?
- Growing levels of automation are replacing a staggering number of jobs, and if the growth continues, the great depression will be ahead of us, not behind us.
- Global financial markets are controlled by companies which you and I have no say in, yet they hold most of the worlds wealth.
In contrast:
- I want X programming language that does Y because Z doesn't do Y, in all likelihood, I will be the only one who ever uses this language, but it'll be a personal accomplishment.
- I don't want to leave the house or stop working for half an hour to make a meal, how can I improve/speed up the time it takes for me to get a meal from my favourite restaurant?
- Making X sucks, I want a robot to do it
- Bad AI for X
- IOT for X
- Framework for X
- Uber for X
How often do I see the latter vs. the former? Imagine if you saw FOSS projects attempting to solve problems on the first list as often as you saw projects attempting to solve problems on the second list. I truly believe the world would be a better place if that was the case.
I said earlier I'm guilty of this myself. I build consumer facing supporting software for entertainment media. I like what I do, but I don't think I'm contributing much to society, and that makes me uneasy.
Maybe revolutionizing the world is not the point. Maybe the people who create toy languages gain a deeper understanding of some issue from pursuing the project. Maybe sharing it helps others discover that deeper issue or hidden complexities and elegant solutions. Or maybe they are just an inspiration and serve as a reminder that everything that exists is built by humans. Everything can be understood. There is no magic.
In my opinion, it's absolutely fine to create anything as a hobby if it makes you happy, whether that's woodworking, drone photography, or even reading or watching something that makes you think deeply. It's impossible for us to judge a priori which of these contributions will be useful to society in the long run. So if the people doing it are happy, why stop them?
Are you really that surprised that you see more of the latter than the former?
The truth is, when you are working on a project you care about, nobody really knows where it will lead you. Big things start out small and usually with humble ambitions. Usually the things that end up impacting the world start out as being considered "useless" and "toys."
I say, work on interesting problems that you feel good about working on but most importantly enjoy working on. Feeling like the thing you are working on is an important problem to solve is a necessary but insufficient condition to ensuring you will get through the tough parts.
Richard Hamming once wrote that he asked his peers the question: "Are you working on the most important problems in your field? If not, why not?" And I think this is a valid question and not a leading one at all. But it also is fair to answer "No" to this question if you are able to clearly answer "why not?" to yourself in a way you can live with.
On the one hand I completely understand your point, and from the perspective of one 'culture' it feels like a waste of time and instead we should focus on 'getting things done' and 'making an impact' (in the shape of a successful business or whatnot).
But on the other hand, one of the main reasons I love HN is that there are still plenty of posts that are just about hacking/tinkering, without concern for 'usefulness' and 'purpose'. In fact, more effort put into something as silly as possible often makes the whole thing even more delightful.
Personally I try to focus on the latter as long as I can afford it. While I really appreciate the startup advice and 'useful' stuff, a few years ago I developed a burnout that, in hindsight, was caused in part by the fact that I stopped being able to just enjoy things for their own sake.
I noticed this same thing happen to friends of mine (without the burnout, usually): in our twenties we often wanted to do cool stuff together because it was cool, even though we really needed money! But then, in our late twenties, it all became about whether we could make money from an idea or whether the idea was 'useful' by some other metric, despite the fact that we actually had enough money finally to not be in survival mode anymore.
For me one of the best improvements to my life (and stress levels) has been to try and bring more of that 'silliness' in my life. And judging by the people I admire, that seems like a good approach to life in general. And already now I notice how my 'silly' explorations actually end up providing tons of benefit for my 'career'.
Indeed, this is what makes most of us engineers. But if I separate myself from it and look at it from the outside, we "waste" a ton of time doing things that... well... don't matter. The exceptional few make advances most of us dream of, and they contribute a lot to the advancements of engineering. Many of us just tinker for the sake of it. I respect people want to do what makes them happy, but I find the lack of interest in solving problems bigger than what can fit inside a hard drive a bit... depressing. It feels like resignation.
> On the one hand I completely understand your point, and from the perspective of one 'culture' it feels like a waste of time and instead we should focus on 'getting things done' and 'making an impact' (in the shape of a successful business or whatnot).
I don't think it needs to be a business, just something that serves as a testament to a legacy that says "I left the planet in better shape than I found it."
The world could benefit a lot from more stoicism, and I think we'll need it in the future that's to come.
The first assumption is that it is better to advance engineering, or somehow 'leave a mark' (fame?), or 'make the world a better place', than to simply live a happy life, all else being equal. While personally I do enjoy it to see others enjoying or being grateful for the fruits of my labor, I don't actually consider this a good incentive. I'm a huge fan of the 'wu-wei' concept in taoism though.
And while I'm still actively figuring out how to apply all this to my life, I generally find that some (and possibly most) of the most intensely happy moments in my life were rather... insular and independent of (rational) 'context'.
The second assumption is that the kinds of things we engineers do are a net benefit to the world. I've begun to doubt that, despite my love for all things engineering. If I were to do anything that has a large enough impact to 'matter', there's always a chance that this might end up having unintended consequences that also matter. I can think of many pursuits that seem more unambiguously 'good'.
All that said, it's possible that you are either 'wired' to want to make the world a better place, or that your life took a path where this is who you've becoming. By my own logic, I can't really disagree with you wanting to pursue that! Just offering a different point of view.
This tool may not be great but can inspire similar tools or may be helpful for developers who want to learn how to develop a software similar to this one, or attack developers to contribute the project and eventually may become a standard tool as UI for ML pipelines etc.
This is a platform, so this is not at at the SKL/Tensorflow/Caffe level, this is something which says: put your data in our data warehouse in this format, tell me what model you want to run on it, go have a coffee, and by the time you're back I'll magically run your model on a big cluster quickly; if you want I can also try out different hyperparams/models, and I'll give you standard plots/metrics to evaluate the thing. You can also use 1000s of existing feature vectors/metrics that other people have put in the DWH. You can also clone an existing model and tweak that. If you're happy with the model, you can deploy, either on my API by querying and I'll tell you the prediction, or I can publish the model in binary form to a datastore and you can just import it in your iOS code in 3 lines. You can also schedule nightly re-trains, I'll notify you if something big changes in the accuracy metrics, etc.
The point is, this sort of ML platform allows engineers at big companies to leverage existing data/metrics/f.vectors, models, infra, etc. and allows them to move 10-100x faster than startups/hobbyists. When I say 10-100x, I actually mean it, I'm also running ML models at home, and the amount of time I waste on this glue stuff is huge (starting from data prep to f---ing around with tensor ranks).
I'm sure there are already startups working on platforms like this for everybody, we'll just have to wait for one of those to become good.
Does anyone knows of any such open-source product actually usable today?
This would solve a real need.
(Also I'd love to work on such a thing)
H2O, Amazon ML Platform, Azure ML, Google Cloud
Actually, in this case I'd bet on the big cloud providers to eventually deliver sth nice (or acquire and integrate).
Looks like H2O would fit the bill.
Anyone has experience with pipeline.ai?
It also has feedback loops to fine tune the model on real-time data and has an option to store all the predictions made by model into Redshift or S3 to be visualized on a tool like Tableau.
The tool itself is not open-source but it can be installed on the user's AWS, Azure, or GCP account to keep their models and data private.
It's still in alpha testing (bunch of pilots) and I'll probably do a "Show HN" post when it is ready for a public beta
Yann LeCun said once a deep neural network predicts many targets, say 1000 classes in ImageNet, it is possible for the model to learn quite generic features. So it makes sense to pre-train on a large amount of data and a reasonable number of targets, and then share the learned feature extractors with others.
Could this be a business? Or a community? Thoughts?
Not a whole lot of general use in having a feature extractor tuned to Uber's data.
"Specifically, there were no systems in place to build reliable, uniform, and reproducible pipelines for creating and managing training and prediction data at scale. Prior to Michelangelo, it was not possible to train models larger than what would fit on data scientists’ desktop machines, and there was neither a standard place to store the results of training experiments nor an easy way to compare one experiment to another."
Seems like h2o.ai fits a lot of that bill.
My team's experience with MLlib has been bad. Especially compared to H2O Sparkling Water.
Anyone else on here find MLlib to not be as good as advertised? I was surprised to see Uber using it.
I'm curious what tooling they're using to ensure data quality, in particular time series data.
The key question is? Is Uber going to open source it? If not, why bother writing articles about the specifics of their platform?
From looking at this from the outside the only things in this platform that may be open-sourceable would be the job scheduling and visualizations - and there are already variants of open source tooling which could be repurposed for those tasks (or may even be powering those components).
The main purpose of this post seems to be Uber's way of standardizing their workflows + a little extra glue (which they're calling their platform). It still provides a lot of value. Also, Uber does have a few cool open source projects: https://uber.github.io/ (but could admittedly have more).
With that said, the title of the post is not "Meet our workflow", It's "Meet X: bla bla". In the world of software, one expects to actually see a product named X that bla blas.
This is a really valuable thing to acknowledge. There is some sharing of company philosophies, but seldom do I see companies fully "open sourcing" their workflows and strategies. Perhaps because the people at the top see that as the real value their company brings - that knowledge. Nevertheless, it's extremely valuable and I wish I could see more things like that. Basecamp's book "Getting Real" is close to what that might look like, I think.
I don't think you would find an incredible amount of use from an open-sourced Michelangelo. The biggest advantage that Michelangelo has for Uber is that it is easy to integrate into all of Uber's other tools.
Depending on what your machine learning needs are, you could get pretty far with just Spark + MLLib, and wouldn't need any of the customization that Michelangelo has on top.
If you're a tiny startup, then Spark + MLLib is more than enough. Even that would be overkill if your data fits on a single machine.
But if you're at a young, but quickly-growing company with:
- terabytes of data
- tens of thousands of features extracted from the data
- dozens or hundreds of unique machine learning models being tweaked over time
then hopefully a blog post like this is helpful. It shows off various effective patterns for solving machine learning patterns at scale. Presumably, you'll want to build your own internal system with its own set of hooks, but the best practices and lessons learned should be roughly the same.
I don't think Uber is anything like IBM Watson, mostly because Uber uses machine learning to solve business problems for its own products--not for other companies' products.
Most large companies that use machine learning in production will have something similar to Uber's Michelangelo. For example, Facebook has FBLearner Flow[1]. Catherine Dong recently wrote a TechCrunch article describing this broader industry trend[2].
[1]: https://code.facebook.com/posts/1072626246134461/introducing...
[2]: https://techcrunch.com/2017/08/08/the-evolution-of-machine-l...