Diffbot launches a web page classifier API: analyzes a day of Twitter
betakit.com
betakit.com
The page tagging technology looks good though.
Classifying individual parts of pages (as Diffbot seems to be doing) is difficult, but I suspect google could take screenshots of pages reported as spam or whatever as one class and compare those to screenshots of pages w/high pr to get a pretty interesting classifier they could use as an extra datapoint. Could be an interesting experiment anyway, using data they've got lying around.
How does caching works? Is there any focus on security? Multiple geolocations?
I liked the TOS :) ---- Diffbot.com is made available for personal, non-commercial, and commercial purposes. Services are provided as-is, and we do not make any guarantees on the quality or performance.
Keep it up!
Beyond that, it's useful for companies that perform analytics or use the information to do better search or auto-categorization.