Google Stakes Its Future on TensorFlow
technologyreview.com
technologyreview.com
This is exactly the same strategy that Google is running with Apache Beam / Cloud Dataflow - make an API that can run anywhere, then make it trivial to use that API on Google Cloud. I think TensorFlow is doing better than Beam because TensorFlow is more of an advance over the previous state of the art than Beam was, and because Google has managed to grow a community around TensorFlow.
TensorFlow is Python/C++. Nice! Beam is Python/Java. Not nice.
I agree with your post.
But speaking about the community part, I am still irritated that there's no official support for running TensorFlow on ARM / Raspbian / Raspberry Pi. Maybe I'm an outlier, but to me it looks like such an obvious choice. Here's a platform that lots of people are using for automation and robotics and computer vision, and it could be their first contact with machine learning, and you could basically lure them all in if you gave them an official release.
Sure, there's someone on GitHub who provides unofficial packages, and props to him for the awesome work...
https://github.com/samjabrahams/tensorflow-on-raspberry-pi
...but I'd really like to see Google stepping in here and doing the obvious.
I've opened a feature request a while ago. Google closed the ticket, no explanation:
Remote involves sending the inputs/receiving the outputs and having the computation be done on a desktop, which even at its most basic, is going to be more powerful than an embedded. Local involves no communication, but has to use slower computational system. Since the typical input/output is on the order of a few MB, data transfer latency is likely to be low. Considering deep nets require significant number crunching for evaluation, the advantage by using a faster machine might make up for the data transfer losses. Thus while there might be benefits here, they aren't going to spectacular.
Training is a different matter, of course. You want to do training on a fast GPU or something.
ML moves fast, and to keep up, you need to be using a flexible yet powerful framework or you will run into major slowdowns in implementation time. Tensorflow is that framework, and, by the way, it also happens to have the best support for actually deploying your models across datacenters.
Migrating your data into some cloud provider is irrelevant and consumes maybe 3% of the total effort involved. Data is maybe 10%, and the rest is developer/research time.
> “The head of Google’s cloud business, Diane Greene, said in April that she expects to take the top spot within five years...”
Namely, full support of symbolic differentiation and all operations and optimisers ported. Having to define everything in python and then load the graph definition into a poorly supported C++/Java wrapper is incredibly tedious.
Surely, hiring 5 engineers per language team and having them port and maintain versions would not be a big stretch to the budget.
Or is this an intentional strategy where the lack of language support forces users to deploy their models on Google cloud workflows in lieu of being able to easily integrate them into their big data pipelines (mostly JVM based)?
Most ML researchers and leading-edge folks want Python, and there's always a lot of demand for making things faster/easier/etc. Adding more first-class languages is on the "should do" list, but it adds a long-lived support burden to keep the language bindings in sync with the core. There are some fairly deep questions about how to make it easy to have good, first-class language support across many languages without multiplying the work by a factor of N. For example, Andrew Myers is spending some time on the team and has been working on beefing up the Java support, but as you can probably infer from this PR, it's ... complicated. The discussion in this PR is a pretty good glimpse into some of the underlying issues surrounding first-class language support.
https://github.com/tensorflow/tensorflow/pull/11251
So - your concern is absolutely founded, and is taken very seriously, but is treated cautiously for bigger-picture reasons. I'd expect continued progress on this front, but I wouldn't expect it on very short time horizons.
From my limited insight (we collaborate with Google Brain), it was my impression that the core problem around such questions is that most research engineers with the expertise to do this are (and I understand this very well as a researcher) more interested in doing research than do full time language porting, and I would expect that to be the case for almost anyone with the expertise.