On edit: if only there were some way to do some kind of really simple syndication of your data.
On edit: if only there were some way to do some kind of really simple syndication of your data.
Of course, I know that some version of this can and does occur with classic web scraping too, but that is an arms race that a search engine can win
Cloaked links and cloaked ads still happen on direct requests, too -- a search engine's crawlers come in a widely known IP range (or if they start using unknown or new IPs, they become known soon enough) so even spoofing the user agent of the bot isn't a reliable workaround.
I'd say the arms race is still escalating, though I've been out of that game for a little while I'm still rather sure of that.
People started serving up pages that required Javascript to show content, so they had cope with that. I'm sure it's dramatically more expensive for them as well.
1. You can really only subscribe to particular URLs. So this would require millions of subscriptions. It would only make sense if most of your pages are changing every few days.
2. You need to also subscribe to feeds to fine new content.
There's no law that says "you can't send more than n packets per hour".
But I would agree it's a very outstanding and real problem that is YC-worthy - sharing structured webpage data with trusted partners in a generic and efficient way. I've heard about various AI companies that perform such data scraping and structuring with AI, forget the name - this is many notches in sophistication above a Selenium-headless type driver. If only html were made into a model-view-controller neatly and users were let to bring their own views & controllers.