Amazon AI
aws.amazon.com
aws.amazon.com
Lex: https://news.ycombinator.com/item?id=13072813
Polly: https://news.ycombinator.com/item?id=13072944
Rekognition: https://news.ycombinator.com/item?id=13072956
Would it be possible to run a service like this in a truly distributed fashion, SETI@Home and Folding@Home style?
How is permission provided? I mean I don't foresee people calling up Amazon and saying hey dataset 2001-G yeah you have permission to use that to train on!
So I guess I would expect Amazon to ask you when you submit your training data, and how do they ask you. Is it opt-in or opt-out and so forth. I would expect them to make it easy to share the data with them, and difficult to opt out.
The problem in most cases is a lack of openly available training data. No open source library helps you with that. It's the data itself that needs to be open-sourced. Companies like Google and Amazon naturally have a huge amount of relevant data but they aren't exactly eager to share that data with the world (which by the way would have huge privacy implications as well).
AWS wins not because they have created an impenetrable black box, but because the economics of reverse engineering (and maintaining) are not cost effective for the majority of users.
(For applications where the problem you're trying to solve is relatively similar to the ones Apple/Google/Amazon are trying to solve - conversational AI)
At best you can partly do this, and even that is debatable.
Here's a simple example: say Google builds a linear regression for the incremental likelihood to click a social post based on the distance of poster from you in a social network. By merely revealing the (single) coefficient in this model, they've revealed something about the underlying data set that they may not have intended.
It gets exponentially (I use this term deliberately) more difficult with multivariate nonlinear models. I have no doubt that substantial information about a data set could be reverse engineered from complex ML models, even if we don't yet know how. But even in the case of basic multivariate regression, the model directly encodes both the mean values and the covariances of all the variables. So that's a fair amount of disclosure already.
In some sense all models are (usually lossy) data compression, and it's just a matter of understanding the compression logic to go backwards from facts about the model to facts about the source data.
Now these do use advanced AI algorithms, Bayesian, ... etc. But it's one factor out of hundreds, and they are nowhere near the top.
Now at Google and Facebook, the rules are like the rules of a firewall. Some new type of spam comes out, it gets past the filters, they look at it, they write a new rule. After 10 years of this with 50 people doing that fulltime you have thousands of rules and it doesn't look like a set of rules anymore. This is why those companies don't get blown out of the water every time a PhD realizes how to make the basic algorithm perform 5% better.
They don't really use the newest algorithms at all. With gmail it's obvious since some of the tricks it knows are things a machine learning algorithm could never ever figure out, for instance how to look up flight times, or package status. No amount of training on mails will ever yield that. Furthermore, they have custom "cards" and other small custom UIs that get triggered by their "AI" classification.
Of the remaining 5%, 4.95% is still the same principle, but basic algorithms can adapt it a bit : human designed rules, but a linear or logistic regression over the data of the past 5 (minutes, hours, days, weeks, months) can make the rule 10% stricter or looser.
The only real "advanced" AI is speech-to-text and in Google's case image search.
The idea of lowering the bar for rolling out these kinds of tasks seems great, but I'd miss being able to think about the underlying models.
And as far as my searches can dig up, it hasn't been discussed on HN yet.
There are a number of duplicates as well.
Lex: https://news.ycombinator.com/item?id=13072813
Polly: https://news.ycombinator.com/item?id=13072944
Rekognition: https://news.ycombinator.com/item?id=13072956