Machine Learning Fairy Dust
stdout.be
stdout.be
Machine learning, on the other hand, isn't innocuous. In order to use the Prediction API, you need a large corpus of data, which will just further incentivize web sites to ignore the privacy implications of their actions. Machine learning is far too abstract and too much of an "umbrella" term for it to be anything but careless to refer to it as some sort of panacea.
If you thought that Facebook's "Beacon" was a slap in the face to online privacy, just wait until you see what the feature holds. Once machine learning libraries with extremely robust, completely unsupervised classifiers become more abundant, we're going to see an exponential increase in the market for data. Banner advertisements will be replaced with much more terrifying 'targeted' ads, and we will enter into an age where we are judged not by the empirical evidence of our actions, but the inferences made from people who behave like us.
Can someone please explain to me why ads for stuff I might actually be willing to buy (as opposed to hyper-annoying junk thrown at me every day) terrify so many people?
Not that I am ambivalent to privacy issues; just playing devil's advocate here.
Gradually, this notion became slowly dispelled when stores actually started leveraging this information to provide, for example, suggestions. It's all fine and good here.
Finally came targeted ads. The part that people find terrifying is when the suggestions are "following them around!". This is creepy in multiple aspects:
The first is that people don't quite understand how this happens (of course, we do). How is this information showing up in different websites? Did the store just let these other sites handle my information? In their eyes, the boundaries for who is allowed to my personal information seem to blur. This also breaks the paradigm of location that the user has in their mind. "If I don't go to site X that is, I won't see anything about site X." It looks like everyone knows everything.
The second is that it's creepy-through-analogy. The fact that it's going wherever you're going and nagging you constantly is weird. When I walk into a retail store, I'm usually asked if I want any help, and I decline politely. If, however, the salesperson keeps approaching me and trying to sell me things that I don't want, I get the fuck out. With targeted ads, the average user can't do that! The creepy sales guy is following you out to the street and into your home. This usually ends in extreme frustration.
Finally, there are also some nuances that targeted ads miss. Targeted adds are actually not that targeted, they're just there to grab at the "low hanging fruit" customers that are on the verge of making a decision. Just because I looked at a dildo once because I thought it was funny doesn't mean I want to be bombarded with the world's finest penis emulators for the next 3 days.
I do actually care but those statements were not a reflection of my cares, it was of theirs. While this is not equivalent to a comprehensive study of average computer using people across the world, they sure as heck cared but didn't even begin to know where to start, in contrast to what you are hoping will happen.
During this time, I let my girlfriend borrow my laptop one night. Each site that she visited was riddled with ads for engagement rings.
Next thing I know, she's giving me the twenty questions about why I'm looking at engagement rings.
Some things need to be kept private even if they are benign.
Until it was too late. Until your spam filter thought it wasn't spam. And on that day, the privacy geeks didth retort: "told ya so."
For example, a supermarket chain might use an aggregate of purchase histories (i.e., products most often purchased together) to influence product placement in the store. They won't re-position products for each customer as they walk in the door, though... Online, this can, and does, happen.
It might not seem like a big deal when it's useful. However, there will be many instances of people trying, and failing, at creating a useful product. I think that's a valid point of distinction: when people fail at adopting new technologies right now, the results are mostly harmless. If they fail with machine learning, there could be some major privacy concerns.
I would add to that list "create a forum." Maybe that's part of "social." In marketing I hear it all the freaking time- you get a half-ass mediocre idea and it always includes some type of "forum" your customers will recruit themselves into somehow, and start to form a community. Most of these people have never been on a forum so I can't blame them for not knowing how it works, but it is a challenge.
The google prediction api takes care of the code for algorithmic computation. While that's handy, it's only one step of a much larger process. The scale of that process is something that many people don't fully understanding about machine learning (yet).
When AJAX first came out, not everyone knew how to do it - but now, everyone can drop in jQuery and do all sorts of complex things relatively easily.
However, if you do need some real magic to be done, and your product really won't work without it, then things get trickier; bad statistics, or at least statistics not really used correctly, is really common in the innards of these kinds of products.
Personally, I think that the benefits of a product should be so evident that people don't care if machine learning was used or not. The pitch shouldn't be "This aggregator is awesome because of machine learning", but "This aggregrator is awesome (oh and we used machine learning)"
I think it is very simple(outside of the secret sauce part) to people who know it, so they don't feel the need to explain it. People who haven't sat down and thought about it see it as magical.
As an example(lifted from Programming Collaborative Intelligence by Segaran ) say you want to recommend movies to people. You have them rate movies. Then you take people in pairs and compare movies they have both rated to produce a distance between those two people. When you want to recommend a movie to Joe, you take the people who are closest to Joe and then find a movie that they rate highly that Joe has not rated, and suggest that to Joe. The secret sauce is in coming up with the distance function.
The problem being (of course) that people forget how hard a problem machine learning can be.
Real-world ML is so full of black magic and hackery that it's the LAST thing I'd try to sell as a web service.