Google Open-Sourcing TensorFlow Shows AI's Future Is Data, Not Code
wired.com
wired.com
Also, as far as I understand, TensorFlow is not technically a set of proprietary algorithms. It's basically a framework for ML.
Ramping up at Google is definitely a stressful process, but I don't think open sourcing technology puts much of a dent in that. Even if you show up your first day at work knowing every tool Google uses, on your second day there will be some new tool and something else will be deprecated. By a couple of years, damn near every piece of software you were familiar with will have been replaced by something different. You are in a constant state of learning.
This is, I think, one of the reasons Google places such a premium on algorithms and data structures in hiring. They are some of the few things that don't change often and being familiar with fundamental concepts makes it much easier to quickly pick up a new tool that uses them.
I wonder if those same concepts translate over to onboarding people who might have a weaker programming background, such as mathematicians / theorists who may have a more difficult time making the switch. For example if they've been using the same toolset their entire careers.
This last point isn't in response to your comment, but more a response to the ideas presented in the article. I don't necessarily buy the idea that the data is incentive enough to researchers at the top of their fields to leave what they're doing to go work at Google. Surely they have access to plenty of public datasets large enough to accomplish what they want to accomplish. So the requirement for switching tools may be a much bigger hurdle when trying to recruit for ML. Maybe I'm wrong in assuming they are targets for employment at Google.
So, I think that they open sourced TensorFlow more as an advertisement/inducement to hire more ML people.
Machine learning frameworks are now commodity software. I give my point about this in a blog post titled "TensorFlow is commodity software" http://blog.eriksen.com.br/en/tensorflow-commodity-software (HN link: https://news.ycombinator.com/item?id=10574444)
You last phrase is right too: Google built a massive infrastructure because it had Google File System, BigTable and MapReduce. Hadoop, inspired by the published papers, turned into a more sophisticated software (an ex-Google engineer said that, but other Googlers denying his claims) and grown into a big ecosystem of products and services. But while Hadoop was matching the feature set of Google's proprietary implementations, Google was getting even bigger.
And to your second point http://www.bloomberg.com/news/articles/2015-10-29/apple-s-se...
Also important to note TensorFlow is probably not the complete package of what they use at Google.
The biggest issue with our learning algorithms is that they are incredibly complicated and require high levels of mathematical understanding. The number of people driving forward machine-learning is small simply because it is such a difficult subject. There are many more people aggregating large and interesting collections of data. I think by releasing TensorFlow Google is encouraging data-collection built around their software; making it easier for a majority of people to benefit from machine learning while ensuring the continuation of their own product, code, data-collection, and ecosystem.
The analogy to deep learning [one of the key processes in creating artificial intelligence] is that the rocket engine is the deep learning models and the fuel is the huge amounts of data we can feed to these algorithms." http://www.wired.com/brandlab/2015/05/andrew-ng-deep-learnin...
Now if you just want to train multiple neural networks (or other classifiers) on different datasets (to have different strengths) then you can keep them separate and build a composite system that lets each network "vote" on an answer to a given problem; the decision of the overall system is a weighted some of the components. See [0].
E.g., once some colleagues and I gave a paper at an AAAI IAAI conference at Stanford. All the good work was just engineering. For our work, basically just some code, later I found and published some applied math that did much better.
Maybe using them has nothing to do with standing on the shoulders of giants but much more with standing on the shoulders of the local maximum that is achievable by throwing insane amounts of data at dumb algorithms.
If you could choose between access to algorithms and data structures that exactly mimic the human brain and a data set that contains everything all humans taken together know, what would you choose?
In end, you need both the algorithms and the data to do the work, and choosing between the two leaves you still wanting the other.
But I think we need to question why data seems to have this outsized value compared to algorithms. I don't think it is some sort of information theoretical invariant. It's a relationship between the specific algorithms and the specific sort of data we have.
Evolution did the processing and saved the results as the architecture of the visual cortex. I suppose that, in a way, it's a work in progress.
It's quite likely we ll have intelligent agents before we can model the brain (even better, we will let them do that for us)
"TensorFlow runs on CPUs or GPUs, and on desktop, server, or mobile computing platforms. Want to play around with a machine learning idea on your laptop without need of any special hardware? TensorFlow has you covered. Ready to scale-up and train that model faster on GPUs with no code changes? TensorFlow has you covered. Want to deploy that trained model on mobile as part of your product? TensorFlow has you covered. Changed your mind and want to run the model as a service in the cloud? Containerize with Docker and TensorFlow just works."
Even if the software is trivial, and is not, is still necessary a lot of specialized, high skilled, work to make the whole AI deal work...
Not to mention that to collect data you need well crafted software...
We interpret it into meaning something without even being aware of all the data that we recieve.
Humans are pattern recognizing feedback loops who simulate a reality. There is nothing what so ever that indicates AI wont be able to do that.
The apocryphal "nobody will ever use more than 640kb" should have taught us to never say never. Of course machines will understand data. Dirty data, full of errors, unfiltered and non curated. Just like we do.
Jokes aside, it if perfectly natural that if we ever manage to understand how biological computers work, we might be able to make them think just like us.
It doesn't need to be now, it can take a few hundred years more, assuming we don't destroy ourselves until then.