The Machine Learning Race Is Really a Data Race
sloanreview.mit.edu
sloanreview.mit.edu
And yeah, Google started to throw their hat in the game with Analytics 360 and an enormously larger training base. Amazon's another major player.
Weirdly enough though, people do still blindly pay my previous employers to figure stuff out because easy answers are always actionable, even if they're wrong. It's just crazy because the CEO explained to me that lying about the service to potential customers and investors was necessary because "Faking it til you make it" was a sound business principle in his mind like 1980's Michael J Fox was his primary sources of business info.
Long story short, don't waste your time with these little companies purporting ML holy grails. They're probably just lying to you, whether intentionally or not. ML is a game for the big boys with access to market level aggregates. The models that last company came up with were wildly inaccurate.
Also you can have problems where more and more data doesn't make an impact - the stock market being a clear example.
I currently believe this is how most ad tech runs internally. Scammers scamming scammers.
So I'd say if you have a very specific area that you're investigating you have a very good chance of beating larger players that don't specialize as much as you can. Of course competing against Google in self-driving cars or machine translation might be a bad idea, but even in those areas there are small startups that produce impressive results (e.g. DeepL: https://www.deepl.com/en/translator). Also, big companies regularly exaggerate their capabilities as well (sometimes more than startups), just have a look at how IBM markets their Watson AI/ML solutions, and what they deliver in reality.
So personally I'd say it has never been that easy to build relevant and interesting ML/AI based solutions as a small team, and it is possible to beat large players if you have the right approach and the right (very narrow) problem.
In sectors like automotive, where every brand is competing to try and build the best predictors using only their own data, there is a huge opportunity available for the first two companies to share data with each other. Doubling your data quantity brings a significant improvement to any model, and would put them ahead of any competition. That advantage only grows the more players you add to the sharing pool.
I believe that if humanity is really going to harness machine learning, the concept of a bulk data commons is an inevitable requirement.
This makes decentralized personal data control, homomorphic encryption, and similar technologies incredibly important.
Call me a cynic, but we can't even get companies to do the bare minimum to protect data that is centralized, and you expect them to do all that?
I'm not kidding when I say that your last sentence is a bigger, harder and more important task than the machine learning it is supposed to facilitate.
Just because Google et al have the underlying technology, doesn't mean they're the best at creating a product that uses that tech.
Certainly gives them a leg-up to get started, but it's not the end-all be-all.
If you're being bought out at X to prevent you from reaching 10X or 1,000X, it's a loss.
Is it an offer you can or cannot refuse?
I was sure someone will say this. I think it only applies for advertising, commerce and insurance. When it comes to learning ML models on images, text and games, big corporations don't have a unique trove of data. They only have the data advantage when it comes to personal data.
The more important advantage big corporations have is hiring the best and hiring more ML scientists and engineers. The job demand far outstrips the offer.
But wait, there's more! How many startups have ready access to usage statistics for every mobile app (Google, Apple)? How many startups have virtually the entire corpus of available ebooks on hand for mining (Google, Amazon)? What about traffic patterns (Google), or virtual reality game(play) data (Facebook)? Telemetry data from the most widely used operating systems and web browsers (Google, Apple, Microsoft)? Video content with official subtitles (Netflix, Amazon, Google, Apple, ...). And don't forget login and activity data through social login (Facebook, Google) and image training via reCAPTCHA (Google).
I can go on and on here. You should reconsider how large the set of training material is that be acquired by mining massive amounts of "personal" data. Of course, this is completely orthogonal to the other major advantage large companies have; namely, they have vast sums of money to throw at compute and storage resources.
The Big Tech companies like to understate the size of their data advantage because it tempts competitors into doing stupid things that won't work in the market. Don't be fooled though - the majority of useful data is locked up inside proprietary silos.
I think the labeled datasets we create for them with e.g. reCAPTCHA is infinitely more useful for training.
For example, one company I worked with manufactured, installed, and did maintenance certain electrical / power and machine/mechanical products. They had archives full, dating back many decades, of error reports and such.
Luckily for them, the equipment problems 50 years ago were no different from problems today, so the data was still highly relevant. The company went on to digitize all this data, structure it in a usable way, and then build (i.e hire in ML consultants) models on that data - which they used for predictive work - in this case, anomaly detection, early warning systems, etc.
The net results were happier customers, less faults, less warranty expenses, etc. which in turn made them more competitive.
This is just one example, but there are many others. Lots of data that the big tech giants do not have access to, but which is (or can be) of great value.
The section about finding faster, less error-prone ways to apply existing insights sounds more like automation than AI. There’s certainly overlap, but they’re two different things.
Small datasets still have massive predictive potential; we just need better algorithms. (As an extreme example, suppose I give you the first 30 digits of pi or e and ask you to predict what comes next. Despite being a small amount of data of low algorithmic complexity, machine learning cannot currently handle this type of problem.)
ML is a great tool that is creating very real and tangible value, but it still has ways to go. Just adding more computational capabilities and more data will only bring marginal improvements.
I was just saying to our partner (as well as to my wife!) how lucky we are to work on Healthcare solutions. We have access to data about medications, opioids, patients and physicians behaviors that so very little of others (who can have any clue about data analytics) has.
The realization of sitting on a goldmine of impossible-to-access-data + capability of developing cutting edge analytics solutions that could change the world is the best place to be.
Can you provide some references for this?
[0]https://arxiv.org/abs/1706.01427 [1]https://arxiv.org/abs/1711.08028 [2]https://arxiv.org/abs/1806.01830 [3]https://arxiv.org/abs/1806.01822
It also helps to be able to show you can answer some of these questions in principle with your model. It gives you hope that it might be able to cover real world images.
What's more important is data quality (e.g. structured and unbiased), and an incumbent with the right approach can have a strong impact.
This seems to hold fairly well, with a few caveats. For example, without enough data you ar wasting your t8me with some techniques. Also “data quality” most often needs to include quality labeling in practice (but you may have meant that inclusively).
For example, let's say you want to build a speech rec engine. You need 15K hours of data to build/validate a model. How would you get that? You could farm it out to some people on mechanical turk and get 15K hours of audio transcribed. With enough money, you could duplicate the transcriptions enough times to actually be pretty sure about the quality of the data set. If you're clever and have a large enough dataset, segmentation generally gets you decent quality. The big gains come in when you have features not present. For example, google realized when you build a speech rec engine, you can include video data and image processing to actually use the way people move their mouths to significantly increase the quality of an automated transcription.
That's fine with today's mindset, but over time it seems like we'll want to go beyond big data and think about how it is that a child can become quite capable without seeing millions of instances of people crossing a road, etc... We have to make our machines make models and 'want' to gather data to test them. By having the machines do the work of today's data scientist, they may well put themselves out of business, and quickly everyone else.
Business strategy is fundamentally about trade-offs, though, so we do need a caveat. Naturally, more choices and more insights can, at times, be a weakness. You always have to prioritize, and as data volume increases it doesn't seem the ability for quality prioritization grows in parallel (or at least, necessarily does).
I've been studying their moves carefully and have no doubt this is where they're headed. While I don't work for Amazon I think AWS in particular is uniquely positioned for these next couple of moves to a level that may make it challenging for others to compete. Interesting times...
The ideal approach would somehow allow this anonymized access to multiple large databases simultaneously. I don't know how you'd do that but if you claim Etherium would help, you'd arrive at buzzword nirvana even if you were wrong.
What moves specifically?
Neuromorphic computing could have the potential I think, if we can build hardware that's good enough.
I think the key fields in AI will become simulator based learning (RL) and graph processing neural nets, because graphs can express any kind of highly dimensional data and are useful for reasoning tasks. They marry the symbolic and connectionist approaches. These two subdomains have had rapid evolution over the last couple of years. They also solve the data problem - in simulation you can produce as much data as you want, and graphs have combinatorial generalisation, thus they work on new configurations without retraining.
I could show someone a single photograph of an animal they have never seen before and they will recognize it forever, from different angles, in different lighting, in black and white, probably even from a silhouette.
The more I dig in to ML the more it seems like it's cheating, like a mathematical trick. At least the way we are using it.
I can't help but feel like it's a game of Pachinko with pixels instead of balls, and neural weights instead of pins and holes.
I'm no expert though so take my opinion with a grain of salt.
What are models? They're a way to describe the data. So this is where a load of philosophical stuff like Occam's Razor comes in and favours things like having fewer degrees of freedom, lower errors, etc.
What is data? It's what makes one model a more likely explanation than another.
You can make up any number of ways to describe some phenomenon, but without data there's no way to tell which of them is better. Or rather, you will fall back on some model with fewer specifics, because of those considerations we mentioned earlier.
So getting smarter (ie more complex) with models can't help on its own.
If we can devise ML algorithms capable of generalizing from substantially smaller datasets, while data will still be relevant, it will be a less important bottleneck in the process than it is now.
This is a result of the ML paradigm essentially starting new with whatever data set it is trying to "solve" but overall, the paradigm doesn't have to be that.
A human learning to play an Atari game that involves an agent jumping over a pit has to learn only the mechanics of that game, and that can be done in minutes. Learning that game from scratch, on the other hand, requires also learning interpreting vision and the whole concept that the world has objects that may move around - which takes months of learning for human brain. So comparing sample efficiency is an apples to elephants comparison if we disregard the ability to reuse/transfer knowledge from related tasks that all humans learn during childhood.
Plus investing is data is much more predictable: the outcome is always going to get better, though the margin will be diminishing, but better is better.
While modeling is not, hiring 100 machine learning 'experts' will not solve the problem 100 times better than 10 of them. On the other side, 100 labelers are surely going to provide 10 times of throughput provided the same upscale.
So by the metric of model performance (the only thing that matters here), what you're saying is that hiring more labelers is actually not linear.
However the labellers are scalable in terms of their throughput and coverage of the data, you can always find bad examples or holes in your current data plane, and that is the time when labelers, not the scientists, are going to rescue you
The entities that own the best data mining process will win. This includes data collection + ETL/storage + model training/deployment.
Obtaining a vertical monopoly on this process is the goal.
So how do you compare the amounts of data needed?
Then show it a single example of a new animal it has never seen, and see how well it would do at telling you if challenge images are or are not that animal.
It's not even going to be close to what a human would manage.
It's not because of the data, it's because a human understands what he's seeing.
I may have all sorts of images in my mind, allowing me to make the distinction?
most people aren't like this but on average engineers who are truly innovatively thinking about these problems and creating solutions are. it is just the way it is
Human babies have bigger head-to-body ratio compared to all other species due to our brain being bigger. Our babies have to be born earlier or otherwise they cannot make it out of the mother alive.
Outside of that, we develop pretty quickly. As you pointed out, everything in us is developing in parallel which is quite impressive.
Alpha Zero got to be so strong at chess through millions of games of self-play. I would like to see how it would fare with only one hundred games, against a human chess beginner with one hundred games under his/her belt.
Only when people will realize that transparency is better than privacy, will they start putting all their quantified self data on there.
This will become a Decentralized Artificial Intelligence network, from which consciousness and AGI will emerge.