Bill Gates Flabbergasted By Gmail
fool.com
fool.com
'He began firing questions. "How many messages are there?" he demanded. "Seriously, I'm trying to understand whether it's the number of messages or the size of messages." '
I don't interpret this to be gates questioning the necessity for more than 1GB of e-mail, just trying to get to the bottom of how this guy managed to use that much in a few months.
Often I can't think of a reason I'd possibly ever use a message again, and lo and behold there's some strange reason I want to refer back to it 2 years later (e.g., figuring out my start date for an old job on a new application ... weird little cases like that).
It may be a language issue or a query-building issue, but I feel no difficulty searching through my gmail mailboxes. I often adopt an iterative approach where I build up my query by watching what gets returned.
Outlook 2003 maybe, but that was eight years ago. Microsoft bought lookout and Outlook indexed search is pretty much instant.
That was the clever thing about Gmail's original invite system. Google could offer an amount of storage that Microsoft could never match, and the invite system ensured they could control the initial volume of users.
The author seems to think that Google's innovation was to simply ignore the storage problem and then it went away.
What they did was work immensely hard at the storage problem, until their innovations in data compression and retrieval made it possible to store larger amounts of data. http://highscalability.com/google-architecture
Yeah, such an old paradigm... Just try managing backups for an ever increasing mail store...
Yes, today storage is cheap (even middle-tier storage arrays where the cost per-GB is about 20x of consumer-level disks can be considered cheap), but keeping large volumes of data safe from disaster is not "cheap". Just ask Google how much time it took them to restore those Gmail mailboxes from tape after they got lost in production a few months back. And that's google we're talking about...
http://www.wolframalpha.com/input/?i=%281+gigabyte+%2F+3kilo...
the answer is 1826 mails per day. I assume what happened here is that they didn't have good stats of the mail sizes in the conversation and Bill tried to estimate something like this and as a result, he thought it was ridiculous.
consider, those are newspaper people. they send HTML emails, PDFs, scans, photos, all that stuff. wouldn't surprise me if they collected a few megs a day. busy people.
And Gates explicitly asked how big each email was, on average.
Current storage offered in a gmail account: 7,602MB
Current maximum size of an email for gmail: 35MB
So you could technically fill a gmail account with 217 emails.
However, looking at my mail backups, I've received approximately 17,000 emails so far this year. The average size of an email was 6.6KB. Altogether, they're using up just over 100MB of disk space.
My live mail has only 15 of those 17,000 emails remaining in it, using up 300KB of disk.
"With Gmail, you can send and receive messages up to 25 megabytes (MB) in size."
mike@alfa:~$ telnet gmail-smtp-in.l.google.com 25
Trying 209.85.143.27...
Connected to gmail-smtp-in.l.google.com.
Escape character is '^]'.
220 mx.google.com ESMTP z4si16632304weq.140
EHLO mail.cardwellit.com
250-mx.google.com at your service, [178.79.145.246]
250-SIZE 35882577
250-8BITMIME
250-STARTTLS
250 ENHANCEDSTATUSCODES
Their MX states that it accepts messages up to 35882577 bytes in size, which is just under 35MB."We'll accept up to 25MB!" And then when you send something 26 or 27MB because you don't notice that it's so close to the limit, the system forgives you.
where something is shared among a sufficiently large set of participants, there must be a number k between 50 and 100 such that "k% is taken by (100 − k)% of the participants
A lot of real world things are close enough to continuous for all practical purposes, though I'll concede that we aren't actually sending and receiving email infinitesimals.
It's not a rule. It's the idea that things are not always as balanced/equal/proportioned as one might hope or expect.
For example, there's the claim that, in the USA, 25 percent of households own 87 percent of all U.S. wealth. It's an example of the Pareto Principle in the form of 87/25. 100 only plays a role insofar as nether value can exceed it (since the numbers refer to percentages).
The Gini coefficient is another scalar that can be used to rank distributions by inequality.
But they don't, and that's a key point. The numbers are, in fact (in this example) 25% and 87%. Different numbers would describe a different situation. They are not percentages of the same thing; they are percentages of two different things.
I will explain this more carefully so that you can understand what the key points actually are. Please take the time to read and understand the explanation below.
You are correct that they are percentages of two different things.
However, if 87% of the wealth, whatever that is, belongs to the richest 25% of the population, then it's entirely possible for 88% of the wealth to simultaneously belong to the richest 28% of the population. That would just mean that 1% of the wealth belongs to the 3% of the population between the 72nd and 75th percentile, which is an entirely plausible state of affairs.
Consider, for any number X from 0 to 100, you can find a number Y that makes the statement "The richest X% of the population owns Y% of the wealth" true, without changing the distribution. Y is continuous and increases monotonically with X; and when X=0, Y=0; and when X=100, Y=100. Under those conditions, there is guaranteed to be exactly one point in [0, 100] where X = 100-Y.
If you want to compare two different distributions, it's helpful to oversimplify them to scalars, since otherwise you have vectors in an infinite-dimensional Hilbert space, which are tricky to compare. If you know that in the US, the richest 25% of the population controls 87% of the wealth, while in Argentina, the richest 10% of the population controls 70% of the wealth (it doesn't), you don't know which country is more unequal. It could be that the richest 10% of the population in the US controls 87% of the wealth, or 34.8% of the wealth, or anything in between. Furthermore, it could simultaneously be the case that, in the US, the richest 10% of the population controls 80% of the wealth (making the US seem more unequal), while in Argentina, the richest 25% of the population controls 90% of the wealth (making Argentina seem more unequal).
There are lots of possible choices of scalar. The smallest percentage of the population that controls 50% of the wealth is one reasonable candidate. The percentage of the wealth controlled by the richest 50% of the population is another. The Gini coefficient is a third. And that unique point of intersection where the richest X% of the population controls (100-X)% of the wealth is a fourth.
Does that clarify matters?
However, if 87% of the wealth, whatever that is, belongs to the richest 25% of the population, then it's entirely possible for 88% of the wealth to simultaneously belong to the richest 28% of the population.
Sure. And you can look through the data to find assorted pairs like that, if you have all the data. If you don't then you're guessing.
Consider, for any number X from 0 to 100, you can find a number Y that makes the statement "The richest X% of the population owns Y% of the wealth" true, without changing the distribution. Y is continuous and increases monotonically with X; and when X=0, Y=0; and when X=100, Y=100. Under those conditions, there is guaranteed to be exactly one point in [0, 100] where X = 100-Y.
Sure, and I understand the value of normalizing data in order compare like things. What I do not see is that all expressions of the Pareto Principle must be given in some normalized form that asumes complete knowldge of the distribution. That knowledge may not be available. That doesn't mean one cannot observe and convey an instance of the Pareto Principle.
Basically, my points are that a) examples of the Pareto Principle do not have sum to 100. For example, if I have a team of five hackers, and one writes 90% of the code, then 20% of my team does 90% of the work. It's a 90/20 thing, and that's a valid example of the Pareto Principle as given.
Could this be adjusted to some X/Y such that X+Y == 100? Here's where I may be missing your point; If all know is that one person, hacker #1, is writing 90% of the code, how can I know the actual distributions from 20% to 40% to 60%, etc.? Suppose hacker #2 is writing the other 10% of the code; then I have a 100/40 situation. If hackers 2 and 3 are each writing 5% of the code then that's 95/40, 100/60 (and 100/80 as well).
This is why I say shifting 90/20 to some X+Y==100 formulation is expressing something different (albeit possibly a true one). So, point b) is that some sets of numbers are both more accurate and more germane to making a particular point. E.g. 90/20 is more striking, and accurate, than perhaps 74/26, which may reflect a truth about the distribution but fails to convey anything salient (though it may be handy for comparison with some other data).
In other words, there's an important difference between making an arbitrarily true statement about a distribution, and pointing out a uniquely interesting aspect of that distribution.
My apologies if I'm still being dense, or missing your point entirely, and I do appreciate your explanation. I get the sense we've been talking past each other.
Around 95% of my (360) emails consume 6.5% of the space (360 messages take).
I believe it took spreadsheets for 'personal' computers to really start taking off in the smaller business world; it wasn't until - IIRC - Windows 95 that computers took off in households. I remember that you were something special prior to '95 or so if you owned a computer.
Okay, I'm extremely suspicious of anyone who claims to process one gigabyte of email per month. Barring large attachments, that is a tremendous amount of sheer data. If you're getting that much email, you probably need culling tools more than you need space.
Regarding the two gigabyte limit being shocking, it was at the time. The only reason Google could pull it off is because 1) very, very few users would get anywhere near that and 2) they were already scaling data at a ridiculous rate. A traditionally desktop-oriented software company like Microsoft would not be able to offer anything similar without tremendous investment of hardware and development.