More On Gmail’s Delivery Delays
googleenterprise.blogspot.com
googleenterprise.blogspot.com
As far as their telling us, the failure didn't cascade causing their queues to fill up and other horrible things to happen. As far as we know, no mail got bounced, no mail got dropped on the floor and lost permanently, etc.
A 2.6 second delay isn't even worth mentioning, how could one even tell the difference between a 2.6 second delay due to a problem on google's side, or routine problems in the email network outside of google that could easily lead to a second delay or whatever.
Did people really get upset about this? Was it worse than it seems because of details they are ommitting? Are they preemptively advertising how well they dealt with this barely-a-problem-noticeable-to-users problem, because the message is really "see, look how damn stable gmail is, our worst problem ever is barely noticeable"?
Or, what?
Maybe those 1.5% of messages were all to the same users, so some people had no delay, and some people had ALL their mail delayed a couple hours? That would make it more noticeable.
But in general, occasional (like once a year or less) couple-hour mail queue delays should probably be expected from any mail provider, no?
I've hosted a dozen or so mailboxes at Rackspace Mail (formerly Mailtrust) for several years now. I have experienced zero delayed delivery or service failures of any kind in all those years.
For $2/mo/mailbox ($10/mo minimum total), you get a 100% uptime SLA, 25GB per mailbox, IMAP-push/POP3, good spam filtering and 24/7/365 support. That's 100% uptime or you get paid, and a real person will pick up the phone if you have a problem at 3AM on Christmas, for $2.
Having your worst damage in case of critical failure being only 2.5 hours of delay, on a very large scale architecture ? That's genuinely great work. 2.5 hours is not even a failure in email terms.
So if this counted as downtime (and that's unlikely) and your email was down for 3 hours during your business hours, and if the mail SLA is the same as managed hosting's network SLA you'd get back 60 cents per user.
Anyway, Google Apps also has a SLA which would credit 10% of the monthly for less than 99.9% -- but again, that's uptime so "some users not getting email on time" is not going to count.
Which is still really fast, especially when you know how mail work and remember the "old days". I guess people got used to email being an instantaneous form of communication, which says a lot about how great it is working most of the time, given that instantaneous is not in email's job description.
Mmm, dialing in to QuantumLink at 300 baud... ahh, the good^Wold days.
Probably this. Gmail was just about unusable for me yesterday.
Not that I'm complaining, for exactly your reasons.
9am it was ~10min delay.
10am about 21min
noon about 1h
3pm emails with ~3h delay arrived.
Then it got better. We werent' able to help our users and they of course blamed that we responded slowly. Actually a bounce message would've been preferable since then they could've used a phone or stopped by.
You don't really expect your company @example.org to take 3h. Either bounce withing ~3-4min or you expect it to be delivered by that time.
After this event are you considering a different architecture?
I guess the people who got upset are those whose email got delayed hours (that tiny percentage of users which amounts to some dozens).
"A 2.6sec delay" is just the median (as per the OP), that leads nowhere if you do not know the mean or other statistics. Which was the maximum delay?
Just trying to clarify that the fact that you are not upset does not mean the 50% people above 2.6 secs should not.
We use Google Apps for Business and were experiencing delays of several hours, and I did receive some bounces when sending to others within my organization. Strangely, most of this happened when using Outlook with Google Apps Sync - not when using the web interface.
Still, we've been very happy with the service - things like this happen every once in a while. We made do until it was straightened out.
The problem does not appear to be with individual user accounts, but from Google's ability to receive emails from certain parts of the internet universe.
The result was lost time and lost money on my end. I imagine this had very little impact on the experience of individual users -- but for businesses that depend on getting emails to their customers, this had a much bigger impact.
My experience was this:
1. This problem was not a 2.6 second delay. It was a multi-hour delay.
2. The problem was not with 1.5% of Gmail users, but with Google's ability to receive messages from other parts of the internet. (A major pipe was down, according to the OP.)
3. No Gmail user had all their incoming mail delayed for a couple hours -- Certain senders to Gmail had all of their mail delayed for several hours.
Technically this is not a problem. SMTP does not guarantee instantaneous delivery, just that if a SMTP server accepts the message it will be delivered.
In reality, this 'delay' could be perfectly normal and is just fine according to spec.
(Also: it's probably not feasible to do for everyone at your companies, but offlineimap[0] is a great tool for ensuring that you always have a local copy of your mailbox.)
Once the issue is resolved and the queue starts flowing again it can take a very long time for these messages to be sent successfully, especially when you need to scan each one for viruses and account for disk latency that doesn't usually occur on these servers.
If 97% of users are inactive, and this affected 1.5% of users, then 50% of active users were affected.