Our new search index: Caffeine
googleblog.blogspot.com
googleblog.blogspot.com
new index: solution: caffeine, Google's new index will analyze the web in small portions and update its search index on a continuous basis, globally. As Google find new pages, or new information on existing pages, Google can add these straight to the index. improvement: user can find fresher information than ever before—no matter when or where it was published.
recent changes in Google search result page suggested that the searched results appears to be categorized further to the different structured data. This could be inline with the rolling out Caffeine.
Is there a way to remove domains form google results as a preference? I know I can always add "-site:codeweblog.com" But can I save this preference for future search queries?
See my blog post at: http://www.joewein.net/blog/2010/07/01/codeweblog-com-a-pile...
The right hand side of the diagram....
So you were right about the "pithy one liner".
This is the kind of one-liner that perhaps Jon Stewart would make...and while there is nothing wrong with The Daily Show, I thought that we, the HN community, are collectively better than that.
Sorry for the rant, but I'm just sick of wading through pithy one-liners on HN.
Can we discuss its poor choice of color? Its lack of visual appeal? The obvious selection of a Home Depot paint swatch to represent Google's old index, and what that means for their database technology? Not really. Not in response to "Wow is that ever a terrible infographic".
I'd be hesitant to flag everything that you don't initially understand. Cut HN some slack, we're intelligent people and we enjoy humor as much as the next man. I certainly appreciate when someone makes me chuckle in the comments, and I'd loathe HN if that went away by your hand.
Given something like http://news.ycombinator.com/item?id=1403672 is it HN that has to change, or the person in the chair? I realize your time may be valuable, but you don't have to read everything and get offended by its presence -- including my comment.
Re my general complaint about one-liners, it could well be that I'm just not getting any of them....
Again, my apologies for jumping the gun. I shouldn't have been so forthright in saying your comment should be flagged.
Is the joke "hey this diagram looks like that diagram"?
Regardless of what pg says, there's many of us that feel that HN has slowly decreased in quality over the last two years (that's putting it diplomatically). Perhaps the time has come to fork HN.
(cf. http://stopdesign.com/archive/2009/03/20/goodbye-google.html)
Ordered and color-coded, as opposed to flying around.
Caffeine is not new and has been talked in SEO community from a quite sometime. I wonder why they make it public now.
To silence iPhone 4 news? No no,just coincidence...
You can fit around 0.5 PB into one rack nowadays. 200 racks then sounds a bit less impressive than 100 Petabytes.
However, that ofcourse doesn't account for redundancy, nor for doing anything useful with such a pile of data. Both of which impose some interesting challenges at that scale.
It is to laugh. I don't think that's been true since around the time they stopped building server shelves out of lego or cardboard.
They've been keeping the full text of the web in RAM (for the snippets), with indexes, several times over. With independent live siblings in multiple datacenters.
Let's suppose for a second that they can get 8GB sticks of RAM for 1/3 the retail price of around $600, and for the sake of the back of the envelope calculation, let's round it up to ten GB. So let's call it $20/GB of RAM. Four gigabyte sticks might seem cheaper until you consider they'd have to double the number of machines to hold them. Now they mentioned that the size of the index is 100,000,000 GB, which means that they would have spent $2 billion on this RAM alone, not to mention all the other components that are required to house it.
Especially considering how much their capital expenditure has fallen lately (http://www.datacenterknowledge.com/archives/2010/01/22/googl...), it doesn't seem very likely that they could afford to spend even $2 billion on this. And the assumptions I made about the price of RAM are pretty generous, considering they'd have had to have been acquiring this for some time, and therefore that the RAM would have been more expensive previously.
So the live data consulted for almost all web queries might be be much much less than 100 PiB of index data, and thus fit far more economically in RAM.
You know how Google has always had that little blurb about how long it took to process your query on the results page? Back when they were for sure running everything out of RAM, it was usually something comically small like 0.00025 seconds. I just checked and it's now more like 0.25 seconds.
Perhaps they've done testing and found that absurdly fast results don't matter (or no longer matter) as much as they thought?
I am under the impression that most of Google's index is still served out of ram. Certainly I never saw any announcements to the contrary. The other poster also pointed out that sub-millisecond latencies are pretty unlikely. Considering all the diverse sets of data they need to pull from, it would be similarly difficult to believe that tens or hundreds of disk seeks could be carried out in a timely manner for each query to yield a 200 ms total calculation time. If I had to guess they have probably just started taking factors into account that they did not before. For example, perhaps they are measuring total internal latency rather than just the latency of a particular sub-system. Or maybe there are components that go to disk while the posting list lookups are done in RAM.
Or you could use PCIe SSD drives. You can get a TB of PCIe for $3000, and a 5 PCIe x16 mother board at Fry's costs around $300. You could get 100PB of PCIe SSd storage for around $300 M.
You could use RAM for very fast cache, and SSD storage for faster than HDD indexing and retrieval.
I merely tried to put the figure into perspective.
http://www.google.com/search?q=java+6+%s&ie=utf-8&oe...
This should also work in Chrome.
The Nexus One is a phone, not a device with a large storage capacity. If Google made one of those, I'm sure it would be referenced.
edit: the Nexus One box also includes a 4GB micro SD card, but that'd be comparing flash memory with harddrives, and is hardly an intuitive explanation.
I wonder how much is junk?
+5 for an intuitive analogy. Priceless!