Show HN: Readability-like API Using Machine Learning
diffbot.com
diffbot.com
It is currently undergoing the "big rewrite" (which includes some proper classification work rather than shooting in the dark), however, but it's still in daily use on several sites. Hopefully I can learn a few lessons from Diffbot!
I should also point out BoilerPlate - http://code.google.com/p/boilerpipe/ - an interesting Java based content extraction project that's being worked on an by an actual PhD student rather than a dilettante like me ;-) Again, Diffbot's stuff goes a lot further than this but there are lessons to be learned nonetheless.
Last but not least, a paper by the aforementioned PhD student called Boilerplate Detection using Shallow Text Features is available at http://www.l3s.de/~kohlschuetter/boilerplate/
I suspect that there's going to be a lot more work in these areas in the medium term because of the growth of the "e-discovery" market and because the dreams of a consistently marked up "semantic Web" have been washing down the pan for a while now.
I use boilerpipe a lot and highly recommend it. We've discussed it before on MetaOptimize: http://metaoptimize.com/qa/questions/3440/text-extraction-fr...
I ran a few qualitative tests on Diffbot's "Article API" and the results also look good. I haven't gotten a chance to run a detailed or quantitative comparison.
What I'd really love to see is a combination of the RSS API and the article API to produce full article RSS feeds for any site.
The combination of the two APIs is a great idea.
You guys should really open up a tagging API. As a developer working on a social site, I'd love to be able to auto-tag content that users upload.
1) User submits post
2) I create a page for their post
3) I call your API to analyze the page
4) I update the page with the auto-tags
5) I redirect the user to the post
This is kind of slow. I could do it with AJAX calls too, but it's still an awkward flow. A better flow would be:
1) User submits post.
2) I pass the text of the post to your service.
3) You analyze the text and send me back the tags.
4) I add the tags into the post and create the URL.
5) I redirect the user to the URL.
This is a much more natural flow.
There are other solutions in this space, like Zemanta, but they generally suck. If I enter a term like "social network" into Zemanta it will tell me Google Buzz is a tag... which is ridiculous.
Enjoy! :)
As someone who's right in the middle of Mining the Social Web, this almost seems too coincidental to be true.
btw, our blog has no rss feed, but you can just use our RSS API :-) http://www.diffbot.com/api/rss/http:/www.diffbot.com/blog
If technical considerations were the only considerations, we would find a way to get at the content directly instead of using this Rube Goldberg mechanism. But of course there are also economic considerations. Content owners don't want to give you unadulterated content for free; their business model requires that ads be served along with it.
Will an arms race develop between scrapers and publishers, similar to the arms race between spammers and spam filters? Will publishers start randomizing their HTML generation, or otherwise making it difficult to separate content from peripheral material?
Just testing another article on HN[1] that the tags are pretty far off. I expected iPad 2 and photos/pixels but i got 4G and manufacturing instead[2]. So I am really interested in how the system came up with the right and wrong tags (which I guess sound more important than find the body of the article, as people are making that easier for facebook and others through Open Graph/RDFa/hNews etc.)
[1]: http://daringfireball.net/2011/03/bending_over_backwards
[2]: tags received: Recyclable materials, Battery, 4G, Apple Inc., Rechargeable battery, Walter Mossberg, Technology, Computing, Manufacturing, Technology_Internet
Is this something you can do fast enough that the delays wouldn't be user-perceptible? Or maybe something that would be feasible to do in batch mode for a bunch of articles?
(In any case, it's amazing. Congratulations on making this.)
In any case, if the feature extraction is taking too much time, what is sometimes done is to dynamically select which features to extract for a test example based both on the expected predictive value (e.g. via mutual information or some other feature selection method) as well as the time it takes to actually compute the feature. This can be measured by, say, average computation time per feature on the training set. This can speed things up a fair bit if the feature extraction takes too long, since you only bother computing the features you really need, and are biased towards the ones that are quick to compute. This may not translate to your particular application, though, if I remember correctly, I've seen it used a while back for image spam classification.
My guess is that they need to render the page so they can determine the visual layout. So regardless of which visual features they use, the rendering step cannot necessarily be avoided.
For example, some of us might be interested in the CSS parsing and feature extraction, but in an alternate machine learning technique.
If you are going to assume tech savvy users, then you might as well expose low-level functionality and see if people like it.
I'd like to have some kind of boilerplate removal that works well for forum content (e.g. phpBB and related), and boilerpipe (the library that I tried) gives relatively mixed results.
Does anyone know an existing solution for this?
Get it working there and you will have a lot more consumers.
Thanks