Why Gnip Will Displace Google - Push Model
mattishness.blogspot.com
mattishness.blogspot.com
This is where search has to be fixed. When I type window handle, I want a HWND. When a builder types 'window handle', his first result should be a local store where he can buy window handles or descriptions of window handles.
That's where search has to be fixed, not in crawling.
I used to be able to craft a search query that would be able to cut out the fuzz from the search and give me the specific results I wanted, but now Google is working very hard to bring that same fuzz back in.
Sorry, until computers can read my mind, trying to make your program guess at what I meant to say means you've already lost.
I think that Guy L. Steele said it best, although he was talking about DWIM (Do What I Mean) in Lisp implementations:
So to this DWIM
Let's say farewell;
The crocks therein
Prove it can't win
And ring its knell:
(Google 'A Time for DWIM' to find the rest of it, although it's not quite as relevant to my point as this stanza. And look at the rest of GLS' poetry/songs too -- very amusing)They do try to get exemplars of clusters, though.
Google needs to spider because the websites themselves can't be trusted to give Google accurate listings.
RSS is another animal, though, because users opt-in by subscribing. Apples and oranges, I think.
Pinging/Pushing is a feature and if it can generate better results Google is best positioned to implement it by combining that data with their existing index and infrastructure.
Push or pull, Google is absorbing URLs to crawl and index at a rate which no start up could match. Eventually, someone may dethrone Google, but it certainly won't be a tiny startup.
Also, there is a significant portion of the web that could never be educated enough to ping some server. This approach is doomed to failure without a Crawler. That's not to say that there might not be interesting applications of their technology is news analysis or aggregation.
Basically, your main sitemap file would consist of a sitemapindex which then links to at least 2 more separate files. The first being your recently updated list that gets flushed when a search engine hits your site, and the rest containing a full index of the content on your site.
Finally, if you're going to turn off your spider, /everyone/ better be pushing to you. So this seems more likely to be next-next-next generation.
I believe Gene Kan (RIP) also had this in mind with his Gnutella work.
I also wanted to build a system based around Gnutella, searching text documents across a whole P2P network, but nobody in the Gnutella community wanted to adopt my ideas.
Push-to-google might be useful for providing more up-to-date search results, though.
Wow this is such an idiotic post.
Err. No. You have a RSS client which regularly pull. It's exactly the same thing. Either the guy has no idea what he is talking about, or he is really bad at making examples.
In a Gnip world, every website would have a feed – whenever content changes – the index gets pinged.
Right. So it was probably just a bad examaple, even though I have my doubts. Still doesn't google more or less have this feature trough it's webmaster tools?
the webmasters need to know how to turn the switch to publish a feed delta, or even write this themselves
with google or ysearch, you just publish and wait for the crawler to find you
the zero-effort solution always wins