Larry Page on Real Time Google: We Have To Do It
readwriteweb.com
readwriteweb.com
Google real-time search had picked up the answer from both a TV forum and Yahoo! Answers in about 5 minutes and it took Twitter about 55 minutes before anyone had an answer in the search results. I felt a bit let down by my first usage of "real time" search from Twitter.
I actually used Yahoo! this week for the first time in a decade because Google just doesn't return good results anymore.
http://www.google.bg/search?q=site:http://news.ycombinator.c...
(for better results try "800 days ago" with quotes, HN strips them for some reason)
site:news.ycombinator.com alex3917 "rule of thumb"
Only half of those posts show up, and to find the rest I need to use searchyc.com
http://news.ycombinator.com/item?id=576727
Now google today for "80-19-1"... it's not even on the second page (and it was the 3rd result to me too previously).
I guess OP's point is that Google is "happy" with fresh content, puts it on the front page, then the evil algorithms ;) push it farther.
On the other hand, this is not an either or proposition. I am ok with this as long as Google keeps the Tweet search results separate but equal (somewhat like they keep the blog search results separate via blog search but equal in that blog posts with good pageranks do appear in search results. Although I can imagine few if any individual tweets having a very high page ranks).
2) Search is won on the margins. Yahoo and Google do equally well on most queries, but users decide which engine is better based on how it performs over all of the types of searches they have to do. So when you take the 80/100 searches you do that are not real time and use unique keywords and you know what you're looking for Y! and G come out the same. It's on those other 20/100 that Google wins users.
I think indexing everything in seconds could definitely be a competitive advantage. I haven't tried Yahoo in years and back then it did much worse than Google. If it has improved this much, it makes sense that the competition is at the margins.
On the other hand, TechCrunch/Twitter et al's idea of real time search seems to be limited to indexing Twitter and Facebook updates as they happen. The arguments they present amount to "Someone tweeted about a plane crash from the crashed plane". I don't think indexing such tweets is going to be Google's edge in search. OTOH I can think of using such information to generate Google Alerts being very useful for some people.
Still, Twitter doesn't do "realtime search", it does "realtime twitter search", so what Googlers ough to do will be more complex.
But with the net at large, blogs etc, this becomes difficult. Incoming links etc are hard to determine in real time (primarily because they haven't occurred yet).
An experiment: here is a search that matches an exact phrase in this comment. http://www.google.com/search?q=%22By+this+measure+HN+is+near...
At the moment it returns nothing. within a minute or two this comment till be the first result.
http://www.google.com/search?q=confusingly+called+copy-regio...
Bizaro world for sure but interesting to ponder.
It was called AwayGrabber (www.awaygrabber.com).
I wrote a overly complex crawler in C to grab away messages from IM networks as fast as rate limiting would allow. Then created a web frontend for viewing all of the status messages from your friends.
It was cool since in most clients at the time you needed to click on a friend and select "get info" for each status you wanted to read. Feeds for status make much more sense. However, I got tired of trying to reverse engineer the changes in various closed protocols (oscar, etc). So I did more than ponder this when I was in college, I tried it.