Urs Hölzle on Google's first data center
plus.google.com
plus.google.com
You'll see a second line for bandwidth, that was a special deal for crawl bandwidth. Larry had convinced the sales person that they should give it to us for "cheap" because it's all incoming traffic, which didn't require any extra bandwidth for them because Exodus traffic was primarily outbound.
This shows that there's more to being a successful business than just good code; you need to know how to run a business, too. How many people working on web search ranking would ask for a special pricing exception from their bandwidth provider and get it? The answer is, apparently, one.
Anyone who knows how peering and bandwidth pricing works.
Basically, bandwidth prices amongst ISPs is a game of chicken and relative strength and perceived benefits; it's one of those areas where being good at talking and having sufficient traffic can make the difference between you paying consumer rates and the other party paying you to get you to host hardware with them in some circumstances.
(E.g. consider if you run a CDN - for smaller ISPs that pay high rates for bandwidth, it may save them a lot of money to get a pop in their datacentres to reduce their bandwidth usage; and in some circumstances it may even pay for them to pay to get a big bandwidth hog to serve traffic destined for certain other ISPs on their network, to "tip the scales" in peering negotiations)
http://www.dodgycoder.net/2013/02/googles-fiber-leeching-cap...
http://www.statisticbrain.com/google-searches/ https://investor.google.com/earnings/2013/Q4_google_earnings...
2 trillion searches at 0.07 cents per thousand is 2 * 10^12 / 10^3 * 0.07 = $140,000,000.
On the other hand, your post mistakenly reads 0.07 cents per thousand, though you did the calculation correctly.
AFAIK, Google doesn't really publish a precise number. They often say things like "Over a billion searches a day". But then they keep that published estimate the same for years during which growth is obviously happening. At some point, they'll publish a step function upgrade to the estimate.
Of the top of my head 2 trillion searches sounds reasonable but I think there is a lot of fudge factor in that guess.
Even in the collocation facilities google was developing this skill for example by installing their infamous corkboard servers with the drives attached by Velcro.
A lot do have a delusional expectation that building on AWS to handle that expected 10x overnight growth that they dream of justifies it, but that "just" justifies good caching and the ability to spin up frontends on AWS or similar if neeeded. Which ironically makes you likely to spend less on your dedicated hosting, as if your setup is ready to spin up cloud servers when the load goes too high, you can afford to get much closer to the wire before you add more dedicated hardware for your base load.
For most people, renting a dedicated server (there's a huge number of providers that charge month to month, with no commitment) will come out far cheaper for anything they use more than ~8 hours a day on average.
To me, if a startup puts everything on AWS, it's a sign they have poor cost control.
I noticed at silcon mikabout that prity much all the start ups where using aws - I suspect that 90% of them would struggle if they had to set up their own colo.
AWS is an order of magnitude more expensive than owning your own hardware, and Google was leveraging economies of scale to operate much cheaper than people that owned their own hardware (e.g. fault tolerant software vs. "enterprise grade" hardware)
People didn't realize it at the time, but "search" wasn't Google's only core competency. It was also "warehouse scale computing".
Of course Inktomi under Eric Brewer was another search company that realized that search is a perfect application for clusters of commodity x86 hardware, or networks of workstations (which AFAIK came from academic projects like http://now.cs.berkeley.edu/ )
(Don't outsource core functions: http://www.joelonsoftware.com/articles/fog0000000007.html )
http://www.flickr.com/photos/nationalmuseumofamericanhistory...
In terms of the hardware, I guess it was a few years, whenever Google started building its own servers, building its own data centers, etc.
There are a lot of challenges as you scale, but it's safe to say that it wouldn't have made sense at any point to outsource/use AWS. New ground was being broken.
It strikes me that both AWS and various PaaS providers (especially Heroku) encourage strong separatioin between the application and its storage. Applications are supposed to treat all persistent storage as network-attached, while the VMs or containers running the appliication are ephemeral. I guess you're saying this is untenable for some kinds of application.