Microsoft forms new 5,000-person AI division
geekwire.com
geekwire.com
So OpenAI is spending a billion dollars over the next several years.
Microsoft is spending a billion dollars per year.
Google, etc. do the math.
There's literally billions of dollars now being spent on moving deep learning forward. Pretty amazing when I think back to 2011 and there were machine learning conferences where literally no one I spoke to had heard of deep learning.
People worrying about a second AI winter are like the people that have been worrying about an "internet bubble" since 2004. It's fine to be worried, there will be bubbles, but this time it's different and there are many reasons for that. There is no "internet industry" anymore; it's become more segmented and just plain bigger. Similarly, there will be no "AI industry"; it will branch out, and there are more potential applications of "understanding data and automating decisions" than there were in the 80s.
> something that they are branding AI - I might wait and see how much of it is on real AI - what ever that is.
Yeah, I wouldn't get hung up on that, they're calling it AI because that's easier to explain to reporters than deep learning. The difference is deep learning techniques are already being used in released products and companies are looking to do more of that, so there is a very real definition and set of goals associated with what these groups are doing, it's not just, "hey everyone, let's make an ai!"
Real AI is a fool's goal. Once we achieve that, we will simply be obsolete, I'm in no hurry to get there.
Yeah, can't wait for the new Clippy.
"I asked Cortana to query our clickstream data and she said that men over the age of 50 like to click on Viagra ads."
This is not AI. Although there is some real research going on, it seems that everyone is jumping on the bandwagon and using AI as a buzzword.
I'll panic when the headline is "Microsoft Forms New Zero Person AI Division" :)
I've never been able to figure out how my employer says I cost them $280 per hour.
If you are client-facing consultant, your billable rate is paying:
- Your salary
- Your benefits
- Your expenses (travel, per diem, equipment, software, anything else)
- And then... Salary, benefits, expenses of all your NON-client-facing teammates: Salespeople, marketing, management, client relationship, administrative, HR, etc.
So as a consultant, you are the product, and as such have to be sold at a price that nets enough profit to pay for rest of the organization.
Normally when they say "you cost $150/hour", what they mean is, "you, and all the support framework behind you, needs to be billed at $150/hour to pay off".
Hope that helps. If you're in-house developer and not client-facing, it's harder to understand and justify cost figures, but typically would be nowhere near that rate...
Or maybe just really, really bad at your job?
You need one of each these people for every N developers (different value of N for different classes of service). Therefore the cost of having you includes the 1/N cost of each of these people.
Billable hour = salary cost + benefits + G&A + overhead
Where:
- Salary cost = gross salary / 1800 (or whatever # hours/yr your company uses)
- Benefits = proportional cost of vacation, 401K, etc
- G&A = General expenses and admin (office, equipment, free food, etc)
- Overhead = amount needed to cover everyone else who isn't billable (HR, Finance, IT, assistants, President, etc)
If you do the math, you'll see $200-300 is a standard rate in high cost markets (NYC, SFO) for someone with a six-digit figure. More junior positions are ~$150, while very senior roles (Partner, Director, etc) would be in the $300-500 range.
MSR is already a pretty big group at Microsoft and does all sorts of work that is completely unrelated to AI. Now, I am confident the teams within MSR will have to start thinking about AI, and how their work relates to AI, but this initiative is much more of an interdisciplinary effort than a 5000 person deep learning team.
Don't they always say that?
The warnings of "Second AI Winter" are in response to the glut start-ups whose sole products are ML/AI platforms.
It's wiser to invest in real solutions that happen to use AI models.
Not a single mention of that anywhere.
So really, any AI-based learning is going to be "deep" from now on, simply because we've reached the point where we can handle large neural nets and complex datasets, right?
It is an substantial amount of money in the rest of the world.
Anecdotally, pay I've seen from Google and FB are higher than Microsoft/Amazon, but of course that's not saying much. Still, the first source I can find matches what I've seen: https://blog.step.com/2016/04/08/an-open-source-project-for-...
Also, Google and Facebook have offices in Seattle, and from what I understand the pay between their HQ and Seattle offices are the same.
Now, Google also has an office in Kirkland. That would be an interesting place to compare. But it's also not particularly large.
You'll and most people too be replaced with a machine and have no job, pretty amazing.
My equivalent to your dearth of people in the know in 2011, was studying AL (Artificial Life) in 1990s, and nobody heard of it. It included the study of ANNs, GAs, GP and AI in general (which I prefer to call CI - Computational Intelligence nowadays). The book that started it all for me [0].
There were no immediate applications aside from expert systems here and there, and the fuzzy logic appliance controllers coming out of Japan. However, now, with self-driving cars, recommender systems, image recognition (face-matching surveillance post 9/11) have big pay-offs or budgets to spend on it.
VR is having its second renaissance. I thought the first time it would have taken off even with the clunky glasses and headsets, since gaming was already such a huge money industry.
And now with modeling becoming prominent again, AL paradigms are being modified, created and repurposed for all sorts of cool things. I play with NetLogo since it is a fun environment for that [1].
My only regret is that I was in at the beginning, but left it to pursue other things, and so I am not at the level of practice to get one of those 200K jobs. I still kept studying it all these years though, and I have coded my own bits and pieces, but mainly for conceptual pieces, art and music, not practical applications. I play with Darknet [2] now, since C was my second language after Assembly (6502 and then x86_32), and it is great fun, and fast. A very understandable and manageable platform.
I am hoping it all leads someday to keeping me alive longer to enjoy studying some more, because as I get older that's really what I enjoy most aside from family!
[0] http://www.springer.com/gp/book/9780387976143
IBM was the first company with a massive investment in AI (during this cycle) - they are now trying to push Watson as a product. Did it pay off so far? I am not quite sure.
It seems like deep learning only makes sense if you have enough data to feed the algorithm with. The kind of data that only big companies can produce or harvest.
Sure you can produce some video and audio data and you can spider the web a little bit, but that doesn't even come close to the resources that these big corporations have and the 'depth' of learning that they can achieve.
So I'm not even trying.
Or should I ?
Is there any place for solo/indy developers in this field ?
Yes you have less data, but your users are probably more similar to each other and less scattered all over the place.
While I feel this is true for recommendation engines, it might not be true for other AI applications.
But this seems more like doing research in hopes of coming up with something new that big companies will be interested in. Or, if that doesn't happen, learning enough so that they can hire you as a researcher or consultant.
Option 1: If you're working on image recognition or anything similar, it's easy to get a ton of data. There are many corpora of image data available, probably on the order of hundreds of terabytes, with tags or at least some structured data. If you're going to "spider the web" you should use Common Crawl instead, and can access all of Blekko's data for the cost of data transfer via AWS. Same for text data, use the Google N-gram corpus.
Knowing that all of that exists, I would still recommend Option 2: outsourcing your AI needs unless you're a researcher or have a decent budget for AI development. Go search "ai api" and pick one that matches what you want to do. Match your skill and risk tolerance; nobody got fired for using IBM, but you may get a better experience from a startup with the possibility of flaming out in a year. You'll get the leverage of whatever company is spending those dollars on your behalf, and you'll be able to concentrate on the user experience instead of the science-ey part.
Option 3 is "if you can't beat 'em, join 'em". Go work for MS / GOOG / FB / IBM building something, and get access to those resources for yourself. Then at some point in the future you'll know their API interface, and you can go back to option #2 with much better data.
Most are typically priced per API call with generous free tiers. I expect these will get cheaper. I'm working on a search engine for lectures (https://www.findlectures.com) - for ~30k API requests I've been able to do everything free.
A lot of interesting large data sets are hosted in the "cloud" too for you to use for research, so you can get at them that way.
Any resources you would recommend for a good introduction to AI/deep learning?
So, I can build products that aren't big enough to interest Google, but include a bunch of tech developed by Google (and others) and leverages their APIs to provide a unique service that is feasible for me to build, and will be useful to a wide variety of people. I can offer it for very little money (one person's side project), and hopefully have some fun learning about deep learning.
I'm so new to it that I'm not even thinking about advancing the state of the art or doing novel work. But, in a couple of years, who knows. Just tinkering with things in a new industry often provides pathways to cool stuff because so many doors are opening all the time. This is like being involved in the Internet in the early-to-mid 90s. You probably won't become the next Google, but the odds of finding a highly profitable smaller niche seems pretty high.
Also, there's going to be a ton of acquisitions in the AI/deep learning space over the next decade. Every company that even does a little tech will "need" an AI story to keep their investors happy. Your tiny thing could be one of those acqui-hires, or maybe not.
Then again, if you have an interest in other stuff, and really don't feel excited about it...probably not worth forcing yourself to get into it. Life is short, you should do stuff that's fun, even if you have to ring the cash register now and then.
But, there's a bunch of ideas I've brainstormed around using things like sentiment analysis and other kinds of very simple-to-use AI concepts for automating tedious stuff. Things like automatically triaging support requests based on how angry the customer sounds, or based on keywords and an analysis of earlier requests; off-the-shelf NLP algorithms can do this today (and Google uses it that way for their own support tools, but doesn't make it widely available in that form, though Inbox has some of that kind of tech working in it). All you need is training data and some familiarity with Python.
My brainstorming exercise goes something like this: Append "with spooky powers" to a bunch of common things until one of them seems cool and useful to me. So, "forum notifications bot with spooky powers", "IRC bot with spooky powers", "twitter bot with spooky powers", "customer relationship management with spooky powers", "analytics with spooky powers", "server monitoring with spooky powers", "log analysis with spooky powers", etc. And I try to think of what I would use such a thing for, if it existed. Then, I sit down and see if I can make it real. The Yahoo NSFW image detection announcement reminded me of ideas I had and tinkered with a decade ago when I worked on a content filtering system for schools...the difference is that now we have the horsepower, the data sets, and the algorithms to actually make it work (but, I haven't worked in that field in a decade and never really liked being a purveyor of censorship tools, even if only for children, anyway).
Anyway, the possibilities are kinda huge and wide open. Many of these ideas will fizzle out, even the ones that look promising, but as with the Internet a lot of millionaires are going to be made by people saying, "It's like X, but with AI." just as people used to say, "It's like X, but on the Internet."
The big companies have an advantage in hardware and research. But they dont care about niche applications of their tech, because prizes worth less than $1B don't matter at their scale. That's where I try to focus on.
The key challenge is data. Too many AI startups get stuck in the "give us your data and we'll do some awesome stuff." That almost never works. [This](http://mattturck.com/2016/09/29/building-an-ai-startup/) talk does a really good job explaining why. The trick is figuring out how to get enough initial data to deliver value upfront.
I suspect there will be very few "pure AI" startups, and a ton of regular old tech startups that figure out how AI fits into their business faster than their competitors or figure out how to use it for a business advantage or to deliver a service that couldn't exist in that way before AI. With the early "like X but on the Internet" startups, the ones that succeeded in the biggest way (Amazon, for example) were the ones that built a great X that leveraged the internet to make it an order of magnitude better X.
So, Amazon was the best book store because they got everything right about being a regular bookstore (good prices, good service, efficient sales channel, solid relationships with publishers) and had damned near every book and could serve customers everywhere; a thing that is only possible on the Internet.
So, the best "X except with AI" company will be a great X company, and then AI will allow them to do some kind of force multiplier to push them to the top of the heap. That means we need to look for opportunities that currently require a lot of resources (say, people, or vehicles, or ) and can have AI added to it to make it produce 10x value given the same inputs. Even 2x value could be a big enough difference to beat your competitors at market, but the real out-of-the-park success stories probably need an order-of-magnitude boost from AI, even if it starts out slower because AI is still clumsy and most of the small companies are starting with tiny data sets (relatively speaking).
Anyway, mostly I think it's cool to play with. I think I see some ways to provide value and make some money with it, but it'll be as much an experiment as a business plan in the short term.
We already have a disturbing quantity and variety of user metric apps for the web. Realtime feedback just requires someone to bring engineering discipline to bear on the space and produce a functional and efficient version. Instead of a bunch of code monkey asshats making my beautiful, fast web app soul crushingly slow because my bosses said yes to one more tracking addon.
At the time, I thought that sounded like an amazing exit (OK, it still sounds like an amazing exit), and wasn't clear how Google could get that value out in a reasonable timeframe.
I was so, so, wrong. The amount of value hidden and public in that acquisition is astounding. Whoever put it together deserves a massive bonus, ideally in Alphabet stock.
MS putting $1bn a year in on AI is a catch-up game. They may do very well at it, but make no mistake -- we are only seeing the public side of the value Google is generating. I don't imagine we'll ever see blogposts about how they're tuning adwords using AI, for instance. But you can bet the same sort of gains they are seeing with translations, audio generation, game playing they are seeing in the ad space.
edit: the Podcast was Talking Machines (about machine learning) http://www.thetalkingmachines.com/
I don't know how reputable PageFair is, but they estimate more than 20% of smartphone users are adblocking now[2].
[1] http://bits.blogs.nytimes.com/2015/08/10/study-of-ad-blockin...
[2] https://pagefair.com/blog/2016/mobile-adblocking-report/
Tell us, how the valuation can be justified in terms of dollars.
Which 'AI' products are useful in the portfolio today, that makes products useful to you and I, that are derived from DeepMind?
The fact is - 'AI' is really short-hand for Multi-layer Neural nets - and they are applying those things in some very specific areas such as voice recognition and image recognition.
I think there will be many more places where we can do this - but I think it's going to be a very long time before we get to 'true AI'.
A) First, I would ask what they likelihood they have in making through trials, what is the real 'cure rate', what is the operational cost of the therapy, what kind of cancer it cures (obscure?).
B) They are not curing cancer. They have some intangible AI technology that doesn't necessarily or may not ever do anything.
As far as 'cutting the cooling bills on infrastructure' - there's no reason to think that that was an issue they were looking at solving, and used the tech as an 'example case' - but that some other, normal technology couldn't have solved th e problem just as well.
B) To confirm the ridiculous hype around this etheric technology, somebody, here on HN is equating this 'black box' to 'curing cancer'. Please.
It seems kind of unlikely that Google is doing that in this case, (if so, the rev rate on new AI tech is insane, and, well, I'm waiting for our AI overlords politely), but I would be very surprised if they were pushing out news about their most bleeding-edge tech.
No - it's definitely not.
Go and walk down the street in any part of the world and ask how many people know about this acquisition? Nobody will.
Even the vast majority of technical people will have never heard of this acquisition.
There is no 'PR' at all, really, from this acquisition.
I don't think any acquisition in history was worth it's 'PR' value.
Unless there is mass market news on it, in which households are getting to know about it, then possibly, but even then, I can't think of a single acquisition that was worth it.
Zuckerberg bought some tech that enables you to wear 'masks' on your face while on video - when he made the announcement, it was fun (he was wearing and Iron Man mask) - and it got picked up globally. If they paid less than $2M for that company, maybe it was worth it for the 'PR value'. But even then it's shaky.
Or frankly they probably already do this, and are just increasing that sell-through rate via AI.
This time around, the technology is seeing real applications today, so the valuations are more grounded in reality. So while people are definitely investing based on the future potential, the worst case scenario -- that we're going to hit a wall next week where no further progress can be made in machine learning research -- wouldn't be as devastating as the last AI winter.
At this point there's a lot of work to do, and money to be made, applying the current state of the art even if no further progress can be made.
It took decades before people figured out Aluminum was actually useful for things for example. It was a chemical curiosity for the later-half of the 1800s (hmm, this is a cheap metal that is found everywhere. But its weaker than steel, what should we use it for?)
Just because you discover something useful doesn't mean you figure out what to do with it.
Napoleon III had his fancy-dinner utensils made from aluminum, for those occasions when gold did not seem lavish enough. And then cheap manufacturing was invented, and the rest is history.
Disclosure: I was part of the initial IBM Watson team, left recently.
1) It took Ballmer to leave for them to really get serious about innovating again.
2) ...because under Ballmer, they were able to make billions of dollars without innovating.
Bill Gates is probably pretty happy about #2 even if he's not thrilled about #1.
Or was there some issue with the group where they had to sacrifice the good with the bad due to lack of overall profitability?
Good people have been leaving MS for a very long time. But big company corporate politics are really toxic in the US. MS is actually one of the better ones. Intel, Amazon, Netflix are much worse. Facebook and Google are on the decline. Yahoo.... yeah
As for this AI thing, my bet is that it is mostly fluff. Most people will ignore it until the next re-org comes along. There was a similar thing with Big Data at MSR a number of years ago and look what happened there.
Something like that.
IMHO, there is some value there.
But, IMHO they would be better off just drawing all they can, including AI/ML but much more from the QA section of research libraries. There they will find oceans of material, where in comparison AI/ML look like farm ponds, in pure and applied math as math but also operations research, statistics, optimization, control theory, applied probability, stochastic processes, mathematical finance, mathematical parts of high end electronic engineering, signal processing, experimental design, quantitative methods in business, and much more.
But the pattern does repeat: Microsoft releases an AI which fails. Tesla's autopilot cannot "see" white object on white background. Apparently, Google also had a crash which is recently being claimed as human error. My guess is that this list is not going to stop here.
Suppose I ask you to build me a teleporting machine. You try, and like the movie Spaceballs, my torso and up comes out aligned wrong. This is now declared part of the iterative learning process, except that the cost borne by the corporations for the failure is quite minuscule compared to the cost borne by the affected party (risk asymmetry).
So while people talk about the huge advancements in AI, shouldn't we be quite skeptical especially at this point? Since none of us have seen the alternate parallel universes, and considering
a) the resources being thrown at the problem
b) the risk asymmetry involved
c) the privacy intrusion involved in the data collection (you knew I would bring it up, didn't you?) and not to mention
d) the inability of anyone to demand any kind of transparency from these AI pioneers
I can as well ask, are we as a society paying too high a cost for this progress? Could we really not do any better than this?
So as Elon musk said recently: "Whatever this thing is you are trying to create.. What would be the utility delta compared to the current state of the art times how many people it would affect?"
The wonderful thing about this AI algorithms is that we can rate them on their efficacy, they might be a black box, but the input and output are always known. If we see that google crashes 10x more cars, we wouldn't use their AI.
This is fundamentally what is being debated here. While the current paradigm can seem fairly poor, let us consider a few things which are true for the human driver.
1. He/she puts himself ALSO at risk, as opposed to the self driving system (remember it is theoretically possible for the self-driving car to not have any occupants at all. It is potentially only a matter of time before it happily wades through stand-still traffic to go and buy grocery for you).
2. He/she is not, in the process of being/becoming a good driver, also taking away personal freedoms of other people - which is effectively what is happening when the megacorps collect any and every piece of data they encounter. In a recent article in the Economist, we hear about a system which augments the autonomous cars by mapping roads in extremely high resolution. [1] Remembering all the work Google does to occlude sensitive information from its maps, imagine how much more effort has to go into this system to have it occlude personal details completely. Now imagine this data (which is currently being collected by a third-party company) landing in the hands of Google/Tesla/Uber etc. who are going to combine it with other human oriented information (e.g. Bob always leaves his office at 5.00PM, and always swerves sharply to avoid the pothole at so and so corner street, let us add that info to our system and improve it).
3. If you think the above scenario is ridiculous, then the next thing you would probably ask for is accountability. In other words, at some point, you are going to ask these companies to open up their data collection processes and algorithms to the world. This is exactly what would happen if the entire thing were a completely OSS-based process. There isn't an equivalent problem for the human driver, because you have sufficient faith in a human's need for self-preservation that you will not demand a real-time thought reading machine which will warn oncoming traffic if the human driver is having an onset of road rage.
> So as Elon musk said recently: "Whatever this thing is you are trying to create.. What would be the utility delta compared to the current state of the art times how many people it would affect?"
This is also being debated. There are such things as side effects, and some of them are invisible. The current state of the art (i.e. the inefficiency, or rather the inadequacy, of humans to perform these tasks) does not, as a side effect, also rob society of its peace of mind. Imagine if, for every piece of information which is collected, you also had a tiny pebble placed somewhere in your neighborhood. Soon, by the time these systems have reached the utility delta that you are happy with, we might have a mountain the size of Everest. Will we? I don't really know. Because it is invisible. Some people would still be OK with it. But most people, hopefully, would want to see the size of the hill. Is it a molehill or is it really a mountain? The lack of accountability surrounding these questions is actually quite shocking to me. [2]
[1] http://www.economist.com/news/science-and-technology/2169692...
[2] Not to mention the other cascading side effects of the data collection process itself, such as your personal data, which you don't even know how it was collected, being collated to be made sense of and sold to the highest bidder
I left Microsoft for Google in 2010 and TGIF Q&A was one of the things I appreciated the most about Google culture (despite the occasional screwball live question). I think any company could benefit from a similar tradition.
What's old is new.
[1] http://www.businessinsider.com/cycorp-ai-2014-7 [2] https://vimeo.com/158956032
Or 5000 employees and 5000 servers each with 8 GPUs.
Oh yeah, and one word: Fusion-IO.
I assume AI development is a niche field. And you would want smaller dedicated teams of brilliant researchers and practitioners focusing on a single problem.
I can't imagine the overhead in maintaining and operating such a large division. I hope they know what they are doing.
The bulk of 'practical AI' don't come from massive products, like 'Office' - but from very specific neural algorithms, applied to very specific things - done by a few people.
That said - this kind of research may benefit from a lot of concurrent research.
Also - there are a couple of strategic issues:
1) Prestige - it's important to be recognized as 'a leader' to maintain brand cachet among tech talent
2) Talent Hoarding - Google, FB and MS are each big enough to tilt the landscape in any specific field. It's actually economically viable for them to pay the best talent to sit in a lab and fiddle, even while accomplishing little, over letting the talent go to competitors.
(Joke; Microsoft Research are actually highly respected, it's just that like Xerox they have some trouble turning it into products)
Clippy is an interesting case because they had a good product and then they downgraded it for reasons.
It's kind of frustrating that they could have introduced something brilliant but chose not to.
[1] https://blog.jetbrains.com/dotnet/2014/04/01/clippy-for-resh...
[2] https://blogs.msdn.microsoft.com/zainnab/2012/02/28/visual-s...