Google Prediction API
code.google.com
code.google.com
It's quite a bit higher-level than what Google is offering here, with all the benefits and drawbacks that entails.
You can find the full documentation on our developer site at http://developer.directededge.com/, and we offer a free developer account for non-commercial purposes: http://www.directededge.com/signup-developer.html
If anybody is giving the Google Prediction API a whirl for recommendations, we'd love to hear about your findings!
Maybe I'm slow, but I don't see how google's prediction API is a good replacement for collaborative filtering type recommendation engines.
However, this is far from a silver bullet to ML problems. It can be quite dangerous, for example, to send off a bunch of data to google and immediately trust their analysis without knowing the underlying application of their algorithm. As a researcher in this area, what I would LOVE is if I could create my own algorithm, send it to Google, upload enormous amounts of data, and get back a result. Because right now it's difficult to scale complex algorithms to datasets in the GB-TB-PB range. Mahout is taking a valid stab at this problem though.
You can have accuracy and scalability. It's just that normally you run these types of things in memory on one machine, not on tens to hundreds of thousands of machines. So these algorithms have to be decoupled and that isn't easily done (but I'm sure Google figured it out a long time ago).
Off-the-shelf recommendation systems are a bit trickier though, yeah.
The following data would be a good start:
1. Text of comment
2. How many points the comment has
3. How many points the article has
4. Time article was posted
5. Time comment was posted
I'd also be interested to see what kind of user bias there is. If you don't provide user names, you could see what kind of rating a comment should have based on its content, and what rating it actually has because certain users are generally loved (pg) or hated (jasonmcalacanis) by the community.
- grellas can post a one liner on a law issue
- DarkShikari can post about video codecs
- tptacek can post about security
- patio11 can post about bingo
- edw519 can post a grocery list
6. The same data for parent and children of the comment
Every comment's rating is heavily swayed by its position within the thread. A comment replying to something at the bottom won't get voted on. A comment replying to something at the top, with just as much merit, will almost always get somewhere between 0.5x to 1.0x as many votes as its parent. A child comment somewhere within a low-voted comment's descendants saying "why am I getting downvoted?" has the potential to call off the latch-on-downvoting effect, and thus give the comment a chance to spike upward. A comment in reply to something no one bothered reading will be ignored despite its merits. Etc.
Of course, your "content worth rating" is orthogonal to all of those concerns, and would probably help in changing the patterns mentioned above, which are mostly bad for the comment futures market. :)
Sorry to sound so negative, but I just earned a PhD in Machine Learning. How would you feel if you were replaced by an API? :-(
EDIT: In seriousness, though, this is a good thing for ML PhDs like you (and eventually me in a couple more years). Building the progress that's already been made in the field into easy-to-use APIs frees us up to work on the next steps.
Now if they just released an Integer Linear Programming API that would give us access to their compute cloud...... :)
Actually I'm not too worried about that. More worried that the guys at google are waaaay smarter than me :-)
By submitting, posting, displaying, or transmitting Data on or through the Service, you give Google permission to process your Data for the sole purpose of enabling Google to provide you with the Service in accordance with its privacy policy. You hereby grant Google all licenses to your Data necessary to process the Data and provide you with the Service in accordance with its privacy policy. As a part of the Service and through provided interfaces, Google may allow you to remotely access, view, and download results of the processing of your Data. (via http://code.google.com/apis/predict/docs/terms.html)
I imagine that they might claim the right to use your data anonymously to improve their algorithms, much like they do for your personal data in their other apps. I mean, what better way to refine their supervised learning algorithms than via an endless supply of training sets? But I hate wading through legalese, anyone have any insights?
4.2. Google claims no ownership or control over any of your Data. You retain copyright and any other rights you already hold in the Data, and you are responsible for protecting those rights, as appropriate.
And they definitely need that. In order to process requests, Google has to make a bunch of copies of your data, create models, etc. Furthermore Google will need to keep copies so that it can use the model it generated for future requests. If the data that you have uploaded is confidential, proprietary, etc, then this requires copyright permission. (Particularly since in the previous clause they made it clear that you retain full copyright.)
For Google's other apps, their privacy policy lets them confidentialize and then use data about you to improve their services. Don't see anything that stops them from doing stuff with your data as part of "providing you with the service", without giving up your ownership of it.
Yes, I'm speculating. But it just struck me as a reason for google to offer this service. And honestly, if I were a user I might be ok with them using my data anonymously to improve said service that I am using.
Also, from http://code.google.com/apis/predict/docs/developer-guide.htm..., it is clear that they perform accuracy analysis using the training data. That is, there is no "testing" vs "training" dataset distinction at this point; there is just cross-validation of the training set.
If they just create a test set from the training set, and omit that from the training, what's the difference? The main thing is that you don't want to include the test set in the training step, and I assume they're doing that.
Not doing something because 'google might do it' is not a very good reason.
There are reasons this hasn't been done before. I don't think it's a viable startup idea. Think about the capital you would need to even get something like this off the ground. Doing the AI API 'right' would require a supercomputer, if you take it far enough.
A guy can dream though. Maybe doing high-frequency trading would get you most of the way there.
IMHO, go for the niche: elsewhere in this thread Directed Edge was mentioned, and they are a good example of a business tackling a well-defined problem with exciting tools and techniques.
by implementing a web app that anyone could use (mobile or not), i also envisioned a sort of community/market place where people could post their data and do simple stuff and/or have experts try to tackle it for a service fee. and whatever algorithms came out of that would be made available for future data that has similar features. i recently came across a similar site. can't remember the name.
anyway, i know this is not necessarily a viable startup idea. and if it is, it's ultimately all about execution. i'm still dreaming though and was psyched google launched their predict api.
With NLTK you can build classifiers, decision trees, and train/predict with bayesian classifiers similarly to Google's Prediction API examples. It's pretty easy to get started, and it's code that you run locally, so there is no network traffic.
I use it on http://www.protopub.com for classifying rss feed stories based on user feedback, so Protopub can recommend future stories that you might like. NLTK is far easier than rolling your own classifiers, but even that is not too difficult. See the O'Reilly book Programming Collective Intelligence.
I made Protopub to scratch the itch I think a LOT of us have. I am about a month away from a v1.0, and that's when I'll announce it on HN. Until then, I'm tweaking AI algorithms, fixing UI bugs, and making sure the back-end can handle the more than moderate traffic that HN will send. The few users I get from posts like this are enough to do some basic testing.
Anyway, I have some appreciation for the difficulties you must have encountered, and it doesn't please me but it will please you to know that at least from them you won't be having much competition.
Is it ok to start using your service? (not from an industrial espionage point of view but because it is useful!)
Right now, Protopub is an experiment, but it also serves as a beacon to other likeminded hackers in NYC, where I live, that I am interested in meeting others who want to create unique and technically savvy projects. It has done a good job of doing exactly that so far.
edit: hm, protopub.com proxies all requests ?
You can also run those distributively without much problems.
Original paper: http://www.cs.washington.edu/homes/pedrod/papers/kdd00.pdf
You should also look on http://www.cs.washington.edu/dm/vfml/
"Automatically selects from several available machine learning techniques"
So not only does it learn, it's learning which learning techniques work best for different problems.
In principle you can do this with any data set and any set of discrete outcomes.
In general, though, you should expect that the resulting classifier won't give you much insight on why it came up with the answers that it did. Plus it frequently is less accurate than a trained human. But it is much, much cheaper.
Incidentally, Google Translate does this and starts guessing the source language as you start typing. I found it interesting that when you type a single character, w is guessed as Polish, i is Norwegian, s is Czech, e is Portuguese...
http://code.google.com/apis/predict/docs/developer-guide.htm...
It's pretty straightforward.
Linked to from this page: http://code.google.com/apis/predict/docs/getting-started.htm...
I can understand the necessity of this, but that'll be some serious lock-in.
wine,5,1,1 no-wine,0,0,10
This data is meaningless to anyone who doesn't know what the columns mean, and you don't have to tell that to Google.
And you can of course replace the labels with any other text you want without affecting the algorithm.
Not saying you need to be so paranoid, just that non-readability of data might be comparable to obfuscating your javascript to keep people from prying into your code.
At least to a sufficiently knowledgeable corporation.
In fact from their TOS (http://code.google.com/apis/predict/docs/terms.html):
4.2. Google claims no ownership or control over any of your Data. You retain copyright and any other rights you already hold in the Data, and you are responsible for protecting those rights, as appropriate.