Apple’s Secrecy Hurts Its AI Software Development
bloomberg.com
bloomberg.com
As a strategy I have observed that people with too much secrecy lose out on innovations that people outside the process could illuminate, but as a strategy it has a history of helping the company employing it stay ahead of the game. Look at Google's secrecy about its data centers, had it gone public with that stuff when it was done, perhaps 2 trillion kwH of power could have been used to do something else if other people had adopted those things. But instead it just helped Google grow faster and do more with their stuff. Is it bad for Google? No. Bad for the rest of us? Perhaps but we don't know one way or the other.
So when Apple's efforts fail to produce a competitive product, then I think you can say its hurting itself, but until then I'm sure they see it as a competitive advantage.
They have done some "somewhat" good things. (I.e. webkit) However, at the end they haven't followed through with it. (I.e. parts of Webkit were broken, they've refused to work with the community etc) At least with Google you're getting the benefits of the Google services and libraries that may not be their own products. (I.e. Guava, Dager, Guice, GWT, Angular, etc)
They addressed this right in the article. It does hurt Apple, because the best people won't go work there. That's why they are so far behind on every AI thing they do. Siri and maps are both way worse than Google's offering, as is the predictive typing. Probably because they can't get the best people working on the problem.
I use an iPhone but it's this lack of AI that makes me seriously consider Android every time it's time to upgrade.
However, while the article claims the "best people" won't go to work there, that is a bit hard to substantiate isn't it? It's seems apparent that the best people who insist on publishing their results won't go there, but are the "best people in AI research" and "people who insist on publishing" the same set of people? Or is it possible that some members of the set "best people in AI" are completely happy to work in a dark lab with no papers published but an essentially unlimited budget?
There is another organization which operates on the same principle, the NSA. They hire some fraction of the world's best mathematicians. Those mathematicians never publish their work, and yet still work there. And every now and then when we get to see behind the curtain, or long after the fact, we discover some really great work has gone on behind that veil of secrecy. Has it hurt them?
The author asserts that it hurts Apple, but I don't find their argument compelling. And there is the sticky issue of defining "hurt". Clearly it isn't hurting them financially, they are killing it. And clearly the feature set parity is close enough that you're still carrying around an iPhone. So how do we define hurt here? As compared to what?
I'm not convinced Apple is hurt by this strategy yet, perhaps when you post here "I really wanted to replace my iPhone with another iPhone but the features of Google's AI or Microsoft's Cortana compelled me to buy a different phone." Then we can start talking about hurt. But so far I'm not seeing it.
Given the secrecy we can't really know how effective Apple's AI group has been at meeting its goals, nor can we tell how quickly they are moving with respect to the state of the art in the industry.
That leaves us with profit from the outside looking in.
I'll go out on a limb and suggest that if Google's margins continue to erode over the next five years as they have the previous five, it will will become that much harder to sustain their AI efforts. So at one level company profits can be a way to estimate a company's future ability to compete in the space.
So honestly if they allowed Google to better integrate on the iPhone, it may not hurt them at all since I'd keep buying their hardware.
Having actually been involved in the AI scene, I can tell you anecdotally that all the very best people refuse to work at Apple (or the NSA, to your point), because they know they can get the same budget and data access at Google, but still be part of the community. Most of them in fact do work at Google, or at other startups nearby that work closely with Google.
Or facebook (Yann LeCun) or Baidu (Andrew Ng) :)
I completely agree that if AI features become the compelling and ranking feature of phones, computers, and tablets. And if Apple begins to lose financially because they are paying a "tax" to license someone else's AI features to stay competitive, and unable to develop a compelling experience on their own, then you can say their internal AI group either doesn't exist, or is not staffed with competitive researchers.
[1] Although I'm not sure if it does everything over the internet or if it's moved any of that to the device now that we have more powerful devices. I'm guessing they haven't, but it wouldn't surprise me if they've changed aspects of this.
[2] IIRC Apple said that they take routes, chop off the beginning and end (so you can't identify the addresses), and upload those using a unique anonymous identifier (i.e. UUID) and no device-specific info, possibly with information on the actual driving (e.g. traffic) but I'm not sure, and they use that info to help improve the product. But that sounds like pretty thorough anonymizing (no way to associate routes with each other or with a device, and the routes don't include the potentially private source/destination) so I'm happy to give them that, and even then I assume it only uploads this if the user has already opted in to sending Apple data to improve their products (which it asks about during initial device setup, and there's a switch in System Preferences for this).
Actually since they know your IP address ,it's not anonymous - it's easy to connect to your identity.
Also , in general the way they gather location(from another post) [1] , is they gather list of cellular and wifi hotspots around you(and i wonder about whether they gather signal strength).And since for example , wifi signals are heavily blocked by walls, they are very space limited in many situations.
So maybe(IDK, the field of anonymization is tricky) combining this data with the data sent when asking for a route is enough get pretty close location ? altough sure, some data is definetly is lost in this process and that's a good thing.
[1]http://thenextweb.com/apple/2011/04/27/finally-apple-speaks-...
Only if they record that information.
When a company is going to lengths to deliberately throw away identifying information, it seems kind of ludicrous to suggest that they're still going to try and track you by using a mechanism that is significantly worse than the information they already had access to and chose not to use.
Or maybe not. we don't know.
Apple has shown repeatedly that they believe privacy is very important (Tim Cook has even gone on record as saying he believes privacy is a fundamental right), and they exert a fair amount of effort in trying to preserve that privacy (e.g. all of the data they refuse to collect, and all of the anonymizing they do on the data they do collect, which still requires the user to opt-in before that anonymized data is sent to Apple). The only way in which your argument makes sense is if you think Apple doesn't actually care about privacy, but only cares about the appearance of caring about privacy, and therefore would anonymize data it collects while still leaving loopholes for it to try and recover some of that data anyway. But that contradicts all of the evidence.
I realize that Apple gets plenty of applicants, but when you're building something that requires some very specific sets of knowledge, it's harder to find them if you don't give them any way at all to find you. The developers who can make Pages/Numbers/Keynote/iCloud great are out there, but they certainly don't work at Apple.
They should probably take an example from research.microsoft.com.
I would agree that publishing results is better for several reasons (ethics, pragmatism, probably more thorough evaluation, etc.), but the definition doesn't really deal with the public at all.
TLDR; Reproducibility != public reproducibility.
There are tons of secret research behind any company (and governments), and it's still "science".
(And conversely, lots of published scientific papers are actually not reproducible, but for most of them nobody bothers -- even if other scientists cite them as accurate).
For example, say I measure gravitational acceleration on Earth by dropping a feather repeatedly. I get a number and I publish it. If other scientists just repeat my exact experiment, they will get a similar number and my result will be confirmed. Yay science!
But if scientists decide to check my number by dropping a wide variety of other objects, they will illustrate the flaws in my original experiment and everyone will get a clearer picture of gravity.
So it's not about just reproducing an experiment; scientists seek to test and expand upon each others' results.
That's because a machine learning system has been built up from known principles and hardware; whereas the human brain is not fully understood, so researchers must work their way down from observed evidence to induce theories.
When physicists were struggling to explain radiation, they didn't call it a bug in the atom. It was their understanding, not the system they studied, that needed fixing.
It's my understanding these efforts continue, you just don't hear a whole lot (if anything) about it publicly.
Check it out, you'll see Apple quietly listed on several university websites, CMU for example: http://www.cmu.edu/corporate/partnerships/cic/tenants.html
Here's an oldie but a goodie: http://people.mbi.ohio-state.edu/datta.53/philtalk.pdf note the talk refers readers to the patent. At the time, reading Google's patents was the best strategy to learn what it was doing.
What would be the incentive to make the result public? Certainly if the discoverer were a grad student at Berkeley they may throw a brief summary onto Arxiv the next day to get some peer validation. But if it were the ML team at Uber's new secretive lab in Pittsburgh? Would they not have a compelling interest in keeping it under the kimono?
There are several possible outcomes for this game to play out. But for right now let's bask in the warmth of our open source learning toolkit abundance and concentrate on the real question of the day: now that we can teach machines, are there better ways to use their time then ascertaining what makes a top selfie?
Also, going from 96-99 may not be really revolutionary, what really matters is how you got there e.g. was new research needed to detect new unicorn classes or scenes, or if it was simply solved by great engineering, more data, etc.
Not if you can easily verify the claim is true with sample images (in the situation given). Then you don't need to care for any "peer review".
In the same vein, if I come up with a data structure implementation that is 100x faster than common techniques, it just has to pass a suite of tests. No peer review needed. If some Intel engineer finds a revolutionary, 100% speed improving cache prediction method, ditto. And for lots of other cases. As soon as you check it works, you can start patenting and marketing pronto.
So it does not matter how Google research worked earlier or if other areas of Google research continue to wait for a 5 year timer.
I know a few professional computer science researchers, people who churn out new and better algorithms (some radically so) in multiple domains on a surprisingly consistent basis for private companies. At least half their job is actually the easily understandable explanation and proof of why the algorithm must have the claimed properties. Their job isn't to produce academic papers but to solve algorithm problems for software engineers writing real code, so a lot of focus is on reduction to practice and clear explanation of the mechanics.
Apple aren't in the business of doing research for the sake of provable science though. They do research to make things they can sell. The goal is very different.
The article agrees, and suggests that because of this, Apple will miss out on the top talent that prefers to publish
The secret private work that these people are doing is a chain around their neck which stops them from being real scientists. Some are willing to give that up for money and work at a lab that doesn't publish. Some are willing to give that up out of a sense of patriotic duty and work in intelligence. It's fantastic that there are private research labs like MSR that pay well and DO publish so there isn't a forced choice between money and being a real scientist for the lucky few who work there.
You don't get super-rich like this, but you get the things that people really want. Respect of your peers, feeling like you are doing something useful, and enough money to run a lab and live a decent middle-class lifestyle, so you can get on with the work you love. That's a good life.
(Also you might want to stay in touch with your grad students who are driven to get rich. Can't hurt.)
It's a lot of hard work, and the benefits to doing it in collaboration with the wider scientific community far outweigh the added value from being able to keep some marginal contribution secret. The top-secret, non-collaborating research lab always falls behind, with perhaps the rare exception of certain large and well-funded government organizations.
I do hate patents, and I think they are often given to ideas that seems simple, trivial, or iterative, rather than ground breaking. Stuff that probably would have been invented by someone else a few years, or even months, later. But for your hypothetical "revolutionary AI breakthrough", that's the ideal use case for them.
I'd class some of Apple's main peers as: Google, Microsoft, Facebook, Samsung, Baidu, and, to a lesser extent, Amazon. These are all large consumer technology companies that are trying to develop quite intimate relationships with consumers, whether through phones or services.
Among these companies, Apple's secrecy with regards to AI does make it an apparent outlier. All its peers are, to one extent or another, publishing much more aggressively than it in an apparent attempt to woo some of the most accomplished grad students in the academic community into considering going into industry. (Not mentioned in article, but relevant for this: Samsung has started publishing a few papers, and I'm seeing lots of collaboration between Samsung-affiliated researchers and SK academics pop up on Arxiv. Huawei is coming up as well, via its "Noah's Ark Lab" and some other R&D centers.)
Tl,dr; most companies are v private and publishing so openly is the outlier, but among Apple's peers/significant competitors, there is a tendency towards open publication and interaction with the academic community.
All the quotes in the story are from professors, who are going to be more aware of those public folks than the ones that hire into more secretive companies. And, I would guess that professors are more likely to prefer (and therefore promote) a public approach to CS research.
Certainly a closed environment will select for people who like that (or are OK with it), but that is not the same thing as selecting for talent, skill, or experience. It doesn't get more closed than the NSA, and they seem to do ok hiring top CS and math talent.
And perhaps they are "the biggest company on the planet" because they don't throw lots of their money on half-baked stuff like AI, but proceed conservatively with what they can reliably offer.
I'm surprised nobody talks about topics such as simulated annealing and genetic algorithms.
AI is not just limited to deep neural networks.
There's some really bad issues with trying to do anything like this in total secrecy:
- We couldn't tell investors how the strategy worked. This is bad for building trust. When things go well, people will invest anyway. But they are quick to leave you when they don't.
- We couldn't hire anyone without bolting down everything. I ended up having to do all the coding while making sure nothing went into the repo that was the secret sauce, having to send a new binary to a consultant (whom everyone knew) every time the sauce changed, and so forth. As a tiny team we had a lot of fiddling with the network to make sure nobody with the wrong credentials could get things they weren't supposed to.
- It affected how we tried to hire people for other things, like new strategies. When you think everything you have is special, you tend to devalue what other people do. You also end up thinking you can take their secret sauce. There's at least one well known fund that thinks it can just interview people and get their ideas. It's pretty obvious when it's happening.
- Because truly original ideas are hard to find, you stop looking and you think that's it. Or, as it happens, you read some paper, slightly modify it, and you announce your revolutionary new idea to the team. Unfortunately, people seemed to buy this. They couldn't just accept that a trading strategy is a whole bunch of ideas that come and go over time. And certainly not the idea that a tiny little tweak that makes everything work is probably a dangerous thing.
- You end up not comparing your findings to the real world. Nobody was reading what was happening in the machine learning world (the guys are essentially using high school math to implement the model). Attempts at introducing things that were completely ordinary in software development were met with derision. Basically if something wasn't more secret sauce it wasn't interesting.
- The crux of this thinking is you tend to think implementation is easy and idea generation is separate and difficult. Probably most people will find the opposite to be true. In truth, if you look at the financial code, it would fit on a screen. The other tens of kLoC was was implementation: back office checking, execution, writing things to database, connections to all the exchanges, and so on.
If you look at the wider industry, you can talk to people at large funds where they don't actually care that lots of people are looking at the strategy. They know what the techniques are, that know how to build code collaboratively.
Cool comment, thanks.
I always assumed that the secret world of spies, secret organizations and their technology development was more inefficient than the secret corporate worlds? Because (a) CIA et al have the same problems and (b) they aren't even limited by having to earn money -- they can even stamp embarrassing foul ups as state secrets!
Eventually the admitted there wasn't anything special there - they simply looked at a graph of the data and made a judgement call.
It wasn't a big deal as it wasn't why they had been acquired, but they really put up every barrier possible - including the "how can we trust you" line. Had to get escalated all the way to execs of a $600M business to get someone to tell them to tell me exactly what they were doing...
This is exactly why people shouldn't be afraid at somebody else looking at the code. What is important here is not specifically the code that is implemented, but the ideas, model testing, analytical thinking that went into the writing of those few lines. The problem with many researchers in this area is to think that it would be easy to go from an idea or simple implementation into a full investment system. Most of these ideas have already been published, it is just that few people know how to use them to make money.
For instance, anyone can pick up a paper about fundamentals (or the Bible, aka Graham and Dodd) and see why it might work. To make money there's a lot of dirty work. Getting some data, making sure it's clean, thinking about biases, getting some lines in to do the trading, coding up a trading engine, writing back office code, getting someone to stare at it all day, and so on.
It is impossible for Apple to keep their self-driving car project secret (if there is one) and get some real-world experience at the same time.
How do you know they're not already testing a car in public and they can keep it a secret?
Or do you think nobody would notice a permit issued to "Ladron de Manzanas Inc" or whatever name nobody ever heard of?
Not to mention the accidents. How would Apple hide all those human-cause SDC accidents?
Occam's Razor is telling me there are no Apple SDCs on the streets, but hey, maybe they invented an invisible, aethereal SDC registered under a shell corporation that nobody became curious about :)
The plot thickens. Add a couple rocket launchers and you have a Hollywood blockbuster.
Then they'd only need to deal with the FAA, and no one is looking in that direction.
http://cyberlaw.stanford.edu/wiki/index.php/Automated_Drivin...
His linked-in profile says he works as a headhunter for apple.
Apples secrecy has become a joke, they're like that guy from sesame street selling you letters.
That is an exceptional talent retention mechanism that I don't think enough people consider when they talk about how the three-letter-agencies maintain their large talent pools. Once you are in in is very difficult to get out.
AFAIK most TLAs invest a lot in training since they're going to really struggle with recruiting people with the specialised knowledge they need; which is fine when that's a reasonable path for your org, but as a company, training AI experts seems really tough.