Gmail having issues
google.com
google.com
> Dear ProtonMail user,
Starting at around 4:30PM New York (10:30PM Zurich), Gmail suffered a global outage.
A catastrophic failure at Gmail is causing emails sent to Gmail to permanently fail and bounce back. The error message from Gmail is the following:
550-5.1.1 The email account that you tried to reach does not exist.
This is a global issue, and it impacts all email providers trying to send email to Gmail, not just ProtonMail.
Because Gmail is sending a permanent failure, our mail servers will not automatically retry sending these messages (this is standard practice at all email services for handling permanent failures).
We are closely monitoring the situation. At this time, little can be done until Google fixes the problem. We recommend attempting to resend the messages to Gmail users when Google has fixed the problem. You can find the latest status from Google's status page:
https://www.google.com/appsstatus#hl=en&v=issue&sid=1&iid=a8...
Best Regards, The ProtonMail Team
Or returning one of the 4xx status codes which indicate less-permanent failure state like:
- 451 Requested action aborted: local error in processing
Which is kinda like a HTTP internal server error as it can mean anything.
Another option would’ve been to accept everything with a very lightweight smtp ingest service, journal it all, and play it back to the full frontend after their code fix was pushed out.
Not an SRE so ¯\_(ツ)_/¯ just some thoughts from my time in a similar role and similar pain points (but thankfully not at this scale)
This is not a cheap shot, but a message to inform users that it's an issue with Google that Protonmail can do nothing about.
> Running a datacenter is no easy task.
Sure, but then there are very view companies which have more experience with running data-centers and (normally) providing reliable email service.
So any outage for more then just a short time is very unusual. I'm really interested what went wrong.
Many of them auto-unsubscribe after a bounce.
The underlying issue (wherever this occurs) seems to be lack of nuance regarding error codes when people try to implement robust systems. Different codes imply different things and shouldn't all just fall back into generic buckets.
Gmail screwed up here, returning a 550 error, it's not anyone else's job to try to second guess that or retry in contradiction of the accepted standard.
Re: the RFC, note it says "should not", not "must not". That seems to suggest they acknowledge repeating might actually make sense in some cases. And honestly the practicalities of this situation and the risk-reward tradeoff seriously tilts toward repeating the request later regardless of what the RFC says. The world isn't going to end.
For any small provider, getting on the shitlist is catastrophic as unlike the big providers, getting off of it will be hard / impossible.
It's a difficult one though, because as you rightfully state, covering up for Google is not the best course of action for the system as a whole, yet it's likely a good course of action for those users who didn't get their emails.
[0]: 4. SHOULD NOT This phrase, or the phrase "NOT RECOMMENDED" mean that there may exist valid reasons in particular circumstances when the particular behavior is acceptable or even useful, but the full implications should be understood and the case carefully weighed before implementing any behavior described with this label.
That is exactly the thought process that leads to non-standard mess that we see numerous examples of.
If you believe the standard is not robust enough to handle problems like this, first work towards a fix to the standard and then implement the solution. Not the other way round.
I didn't suggest people should apply this thought process in arbitrary cases. I said it should be applied in this case. You can take any thought process that gives a good outcome in one situation and obtain a bad outcome by applying it to the wrong situation. That's not an indictment of the thought process. It's just an indictment of the person failing to correctly judge its applicability.
That said, by all means, do try and go fix the standard; I wasn't trying to imply you shouldn't do that.
Exactly, that is why it is important to follow standards. Most engineering decisions are not clear-cut and are born out of tradeoffs. That is why we agree on standards that define those tradeoffs instead of every one of us having our own take on situations.
> Nobody cares if their mailman's knocks follows an RFC or not
If there is a Mailman RFC which says: "If someone opens the door and says `Mike does not live here' then DO NOT attempt delivering the same package"
THEN I expect the mailman to not bother me again, EVEN IF it was actually my mistake that I forgot my roommate Mike actually does live at this address.
These incorrect responses could be caused by mistakes which the remote server admins could reasonably avoid, like software bugs. I understand not having much sympathy for that case, especially from an organization with no shortage of resources. But they could also be caused by, for example, hackers or governments exerting control over the remote server temporarily.
A standard which explicitly refuses to acknowledge these possibilities is not what I would describe as “robust.” An obvious better alternative would be to set some standards around what constitutes a polite retry policy.
I saw someone on Reddit say his SES was suspended for sending tons of bounced emails in a short period of time - it's taken very seriously by ESPs.
E: also user rtx a few comments below
If you want revenge for modal popups, your best bet is to create a bunch of throwaway email accounts, subscribe to the mailing list from them, and start reporting the individual messages as spam when they arrive. Flag them as junk at the mailbox provider (Gmail, Outlook, etc.) and use the links in the List-Unsubscribe headers to flag them at the ESP's end, too.
I have good experience with them fixing issues related to their spam-related flagging for messages that are coming from our self-hosted email server, but never got any specific reply.
Like HTTP, SMTP is also designed to be stateless so, in the first place, the remote server shouldn't return a permanent error in temporary failure scenarios.
The default error should be 450: "Requested action not taken – The user’s mailbox is unavailable”, not "the user has deleted everything and left".
These standards worked well before big players came and told "My responses tell what I chose them to say, and these meaning doesn't always overlap with the established standards". The only exception is spam and we now have standards for helping to reduce it.
Google's mailserver could genuinely believe that the user doesn't exist, if the user service doesn't fail completely but cannot access part of the data and thus doesn't find a user record. In this case the returned "user doesn't exist" error is intended behavior of the mail server and the post you replied to still stands. If you sent to that email successfully earlier, it's much more likely that the server is responding erroneously than that the email actually got deleted.
Mailing lists believing what an email provider tells them and acting in an overly cautious way is a separate issue.
This can't work; you can say that gmail's system should have a component that recognizes the difference between various failures, but that new component can itself fail. You can't solve the problem of "what if something fails" by saying "just add a new component that won't fail".
Note that this is rather different from physical, mechanical systems which can fail in all kinds of exciting and unpredictable ways due to physical wear and tear, things getting jammed in places, component failure, etc.
That is true in a perfect world. In the current world, there are all sorts of ways that code implemented one day does not run the same the next day. Say the code is in an interpreted language and an unrelated sysop updates the language runtime in a way that changes the behavior. Again, in a perfect world that doesn't happen, but that is not always the world we live in. I have great sympathy with people who treat software systems AS IF they were "physical, mechanical systems which can fail in all kinds of exciting and unpredictable ways".
That's true, but human behavior is also fundamentally deterministic, and those two observations are about equally useful.
> Note that this is rather different from physical, mechanical systems which can fail in all kinds of exciting and unpredictable ways due to physical wear and tear, things getting jammed in places, component failure, etc.
No it isn't. Those are deterministic too.
If the protocol is stateful, why the state should be kept by the "sender" and not by the "receiver"? Being stateless removes this ambiguity in my opinion.
Also we should remember how bad is for spam reputation sending emails to a non-existent address and thus I would not blame it on the mailing list for being "overly cautious".
Hard-failing good addresses is a much worse bad than soft-failing bad addresses. In the latter case, remote sender tries again later and eventually gets a hard bounce. In the former, good addresses are permanently dropped from numerous services, and sent mail is lost rather than retried.
Critical failures should soft bounce until positively determined otherwise.
If the a mail server can't tell whether a user/email is valid, it should either return a temporary failure or accept and queue.
Unless of course you're too big to fail, then you just do whatever you want.
Actually, I don't think so.
> Google's mailserver could genuinely believe that the user doesn't exist, if the user service doesn't fail completely but cannot access part of the data and thus doesn't find a user record.
As a system administrator and/or provider you have to think about worst case scenarios and provide sensible defaults. Your mail gateway should have some heartbeat checks to subsystems it depend on (AuthZ, AuthN, Storage, etc.) and it should switch to fail-safe mode if something happens. Auth is unreliable? Switch to soft-fail on everyone regardless of e-mail validity. Can hard fail others later, when Auth is sane.
Storage is unreliable? Queue until buffer fills, then switch to error 421 (The service is unavailable due to a connection problem: it may refer to an exceeded limit of simultaneous connections, or a more general temporary problem) or return a similar error.
SMTP allows a lot of transient error communication. Postfix, etc. has a lot of hooks to handle this stuff. Just do it. Being Google doesn't allow you to manage your services irresponsibly. If we can think it, they should be able to do it too.
Google SMTP servers should have returned a soft bounce here (not hard bounce), so then retry can work.
That's the standards-compliant way. Also I'd argue that spec'ing your code to handle cases where Google fails that badly is (was?) a poor allocation of LoCs.
1. If the user's mail service penalizes you equally regardless of whether the recipient's addressed existed 1 day vs. never existed, that itself is absolutely inexcusable nonsensical behavior that needs to be fixed. You shouldn't do that, just as you shouldn't shoot the mailman (or even arm yourself...) merely because he knocked a second time.
2. Notwithstanding the previous point, I don't buy this as valid justification anyway. The proposal isn't that you should blast 100 emails toward the mailbox every time you get a bounce due to an address not existing. The idea was to just exercise some intelligence in the matter. Like maybe just retry a couple times, spaced out by a day or two. The bounce rate increase due to such an event is very negligible here—people don't suddenly delete their accounts en masse. When that happens, it's clearly due to an outage, not because half the users at that domain suddenly decided to delete their accounts. (Which is something you can also easily detect across the domain as another useful signal to drastically lower the bounce rate across the entire domain, btw, if you're absolutely paranoid about your immaculate delivery rate dropping by an epsilon. But it shouldn't be necessary given how negligible the impact should be.)
So I don't buy this excuse one bit.
What you're proposing is to explicitly ignore the specification (which says that you should _not_ retry when you receive a 550) and try to implement a custom smart retry logic that handles temporary error cases, but also does not get you blocked.
> So I don't buy this excuse one bit.
I'm all for building resilient services, but "try to detect when a server incorrectly returns 550" is not something I would prioritize at all. I'll happily manually clean up after this occurrence than to have this complicated retry logic. It's not an "excuse", it's a very sensible trade-off.
That means there are two sides to the interpretation of what SHOULD NOT means. And in this case, senders have, due to experience, interpreted what Google does when someone SHOULD NOTs:
- The sender SHOULD NOT send us the same sequence again when we reply 550, if they do they MUST go on our shitlist.
Obviously it's not so binary and it takes retrying to several different recipients, but people have very good reason to interpret this SHOULD NOT as MUST NOT.
Email service providers are HIGHLY incentivized to act 100% in accordance with the wishes of the system where the mailbox exists because it’s highly likely that acting in any way that’s considered abusive could get your emails landing in a spam folder.
Mail boxes cease to exist thousands of times a day at places I’ve worked previously. Employees leave all the time and people shutdown mailboxes, this is Google’s fuckup, nobody else’s.
So if you are not getting any notifications from GitLab, even though your email is correct, I suggest contacting them and asking if you have been blocked due to an error.
[1]'Check email service status before sending emails' - https://needgap.com/problems/178-check-email-service-status-...
I fear that this will lead to many lost mails. In my experience, users are often confused by the technical "Mail delivery failed" mails and tend to ignore them or write them off as spam.
https://pbs.twimg.com/media/EpUE20UXYAEa_Uv?format=jpg&name=...
That's a peak of 90% of Gmail inboxes bouncing – and this has been going on for almost 24 hours.
What a joke. And this after we're leaving AWS Workmail because of bounced emails.
No luck with signing up so far.
About your query
I gather that you are concerned about your Ads Disapproval for your Google Ads Account.
Observation
I understand that this is taking a bit longer as we are working with a limited staff due to Global pandemic and there is another team who reviews the account so there can be a slight delay in the decision I apologize for the inconvenience caused as I understand this is not the answer which you are looking for but be rest assured I will get back to you on coming Friday 12/18/2020 end of business day.
For any further assistance, I am just an email away.
Sincerely,
I logged into our sendgrid and mailgun accounts and manually purged all the failed gmail records.
Customers generally cannot change this on their end as far as I can imagine -- this is on the ESP end and is a protection built in because you are sending from their IP / Server and they don't take kindly to that.
The reason why this is so nasty is not because Gmail went down, but because they returned a 5XX permanent failure and not a 4XX temporary failure for these bounces. Literally every email provider will respond to a permanent bounce by suppressing all further emails to that email address (it's permanent, after all!), so the fallout from this will be huge.
The action for rectifying isn't too difficult, but the implications are still pretty big...
This year i decided to do "something" about it, so every mailing list mail received in my inbox that i don't want/care for gets an unsubscribe. It has already reduced my daily mails by a somewhat large amount. It's hard to say exactly how much, but i estimate around 10 emails less every day.
Most of the unsubscribed lists are from companies where i've purchased something andthe seller took the liberty of subscribing me to their mailing list. Those are mostly pre-GDPR that i've just never gotten around to dealing with.
The execption is of course obvious spam mails, to which unbsubscribing will probably do more harm than good.
Rant: As I side note I usually try and buy direct when shopping online rather than through Amazon (for all but the most trivial purchases) and this is the 2nd largest drawback (behind filling in CC and shipping info) - because I bought one item from you, once in my life does not mean send me a daily email, and then when unsubscribing pretend like I signed up for them! For me it’s one of the easiest ways to destroy brand loyalty/reputation.
Plenty of critical communications get caught in this storm...
A lot of clean up is going to be needed as a result of this.
To add some more details, when using a 3rd party email delivery service, those services will either black-list or just outright remove email addresses when they get a hard bounce "email address no longer exists" message back.
Some providers make re-adding an address after a hard bounce a non-trivial task, since after all, the authority on that email address just said it doesn't exist.
This is going to be really ugly.
Kind of happy I had to do something else and I didn't burn hours investigating.
Most systems operate more immediately in isolation on individual addresses than that right now, because such analysis is generally not needed (until today, of course ;-)).
The problem here is Gmail has been throwing out "NoSuchUser" errors which are an instant unsub in most systems because Gmail takes repeated delivery to non-existing addresses into account for deliverability purposes.
I'm extremely paranoid about email hygiene, tiny bounce rates and high delivery rates, so we aggressively unsubscribe troublesome addresses (often to the point of getting reader complaints about it) for many reasons beyond that, however.
I think you mean "reputation purposes"?
If so, wow, that sucks. Their opaque rules have conditioned their counterparties to punish Google as hard as possible for a screwup.
Good for karma, bad for everyone though.
That better describes what I was trying to say, yes. Reputation then affecting deliverability.
Over 80% of our subscribers use Gmail so to say I'm paranoid about maintaining a good record with them is an understatement ;-) Gmail is a huge weak link for us.
"Gmail going down" would not have caused this problem. Even if all their SMTP servers went offline.
(And if no such thing is detected deleted quarantined mail addresses.)
That simple fix buys them 24-72 hours to solve this properly.
Yeah, it burdens servers sending mail to them because now they have to hold on to all mail (including mail that really is permanently undeliverable) for another day or so, but that's still better than what's happening right now.
His solution would result in exponential retry failures baked into most services, which would buy them a few hours, and result in no lost emails, and no suppression list additions.
That is a better scenario, than 5xx.
There is no way that putting in a hardcored hack like that would have been faster. Making the change is, of course, fast.
But then you need to review it (and this is a super risky change, so the review can't be rubber stamped). Build a production build and run all your qualification tests. (Hope you found all the tests that depend on permanent errors being signalled properly). And then roll it out globally, which again is a risky operation, but with the additional problem that rolling restarts simply can't be done faster than a certain speed since you can only restart so many processes at once while still continuing to serve traffic.
The kind of thing you describe simply can't be done by changing the SMTP server, in 2.5 hours. The best you could get is if there was some kind of abuse or security related articulation point in the system, with fast pushes as required by the problem domain but still with the sufficient power to either prevent the requests from reaching the SMTP server at all, or intercept and change the response.
As a trivial example, something like blocking the SMTP port with a firewall rule could have been viable. Though it has the cost of degrading performance for everyone rather than just the affected requests.
My mail server logs show about 20 failures in all of the last week until yesterday 20:43 CET, then 350 failures between 20:43-00:21, then nothing after that. So fair enough, from the client side rather than the status page it looks like 3.5 hours rather than 2.5.
But still, given that resolution time, the suggested solution of changing the SMTP server is absolutely ludicrous.
That means I can't just resend the the emails blindly, because I'm too scared to trigger some sort of automatic suspension...
(I don't do this regularly, so I'm not familiar with all features... additional mail verification could help probably ....)
[1] - https://en.wikipedia.org/wiki/List_of_SMTP_server_return_cod...
I am astonished that either (a) this switch has not been flipped yet or (b) this switch does not exist.
Somebody is incompetent here.
https://www.google.com/appsstatus#hl=en&v=issue&sid=1&iid=a8...
On the other hand, their status dashboard reported similar issues yesterday and here we are again: https://www.google.com/appsstatus#hl=en&v=status
I'd absolutely hate to be hit by this at this time. Thankfully I've made an time investment to run my own mail server years ago. A handful of times it broke down, it either went offline or started returning 4xx codes due to misconfigured or broken milter after an update. Neither meant lost messages from normal senders that use queuing MTAs.
Is it? Is dealing with IP reputation, getting your emails accepted by major providers, and being on the hook for fixing everything yourself very easy? I haven't tried, so I don't have personal experience, but I've heard enough horror stories to think that it's not a good use of my time.
TLDR: Before you spin up a mail server, check if your IP address is on any of the blacklists [0]-[1] as well as Proof Point's list [2]. If it is, then try and get a different IP address.
I spun up a hosted server on Digital Ocean and received an IP address. I checked several black lists from a few email testing/troubleshooting sites [0] and [1] and all was groovy; my IP address wasn't on any list.
I got a bunch of 521 bounces when I tried emailing a neighbor who had an att.net address.
So, I checked the troubleshooting websites, and my IP address was listed as clean.
My logs said I should forward the error to abuse_rbl@abuse-att.net, so I did.
Those emails were never delivered, because abuse-att.net had its own blacklist. I was getting 553 errors. In the logs, the message from their server told me to check https://ipcheck.proofpoint.com.
Proof point runs their own blacklist that some enterprises use (e.g. att and apple [3]). I checked their list, and lo and behold, my IP address from Digital Ocean was blocked [2]. Digital Ocean wasn't able to remove the IP address from their blocklist and suggested I spin up a new droplet with a different IP address.
I didn't want to do that, so I sent Proof Point an email that went unanswered; the email asked them to remove my IP address. I forgot about the issue for five or six months (this is a personal server), and ran into the issue again a few months ago. So I sent Proof Point an email again, this time with different wording emphasizing that "my clients" were having delivery issues. Within a day, they removed my IP address from their block list.
So, my main suggestion is to check if your IP address is on any of the blacklists as well as Proof Point's list before you start on your server. If it is, then try and get a different IP address.
Does anyone have more "enterprise" lists, like Proof Point, to check?
[0]: https://www.mail-tester.com/
[1]: https://mxtoolbox.com/blacklists.aspx
[2]: https://ipcheck.proofpoint.com
[3]: https://www.reddit.com/r/email/comments/6toxzr/ip_blocked_by...
Regarding getting a bad IP rating, normally that's due to having an insecure config, like acting as an open relay, or not having DKIM enabled. There are lots of tutorials online about this, if you know Linux it really is easy.
Receiving side is where there is a great range of options, and many things to try and have fun with. You can have anything from a single catchall mailbox with no filtering, no GUI, and a simple IMAP or POP3 access for MUA, to a multi-account, multi-domain setup with server side filtering, database driven mailbox and alias management, proper TLS, web MUA access, etc. It can also be built up gradually, starting from very simple setup to something more complicated so that you never lose account of how things work.
The triggering event may be an email bounce. I get a lot of github notifications sent to my email, and the failure of just one/a few may trigger the reverification.
When this happens, you can spin up a temporary server and have a mechanism in place to redirect email so you don't go down when your provider does.
Losing incoming email is pretty much the worst case scenario when it come to configuration errors. It about as bad as not having backups, in that both cases results in unrecoverable loss of data.
> When this happens, you can spin up a temporary server and have a mechanism in place to redirect email so you don't go down when your provider does.
Use a commercial provider, but fall back to your own server when it goes down without changing your email address.
Sure there will be some internal turmoil going on right now, but isn't there some non-confidential info to share? Can't imagine this will hurt the image of google neither in the short nor long run, quite the opposite.
It should not be a problem that gmail is "down". Unless this would be happening for more than a few days, noone would lose e-mail. It's a problem that it's not returning a temporary error code, but permanent one.
I think a lot of time and effort is spent categorizing errors from external systems into transient or permanent, and it's always kind of a one-off thing because some of them depend on the specifics of the calling application. It definitely takes some iteration to get it perfect, and it's very possible to make mistakes.
You don't have milliseconds. You can take quite some time to handle the client. 10s of seconds for sure. For example default timeout for postfix smtp client when waiting for HELO is 5minutes.
https://status.cloud.google.com/incident/zall/20013#20013003
Well I guess the thing is left unanswered for now is why the quota management reduced the capacity for Google's IMS in the first place.
Maybe we will know someday :)
* We have a lot of automation/tools to prevent incidents when mitigation is straightforward (e.g. roll back a bad flag, quarantine unusual traffic patterns), which means that when something does go wrong it's often a new failure mode that needs custom, specialized mitigation. (e.g. what if you're in a situation where rolling back could make the problem worse? we might be Google, but we don't have magic wands)
* Debugging new failure modes is a coin flip: maybe your existing tools are sufficient to understand what's happening, but if they're not, getting that visibility can in itself be difficult. And just like everyone else, this can become a trial and error process: we find a plausible root cause, design and execute a mitigation based on that understanding, and then get more information that makes very clear that our hypothesis was incomplete (in the worst case, blatantly wrong).
As Douglas Adams says, "The major difference between a thing that might go wrong and a thing that cannot possibly go wrong is that when a thing that cannot possibly go wrong goes wrong it usually turns out to be impossible to get at or repair."
Here comes the poison pills!
However, I'd speculate that in this instance, when you get that .0001% problem, less hands on deck makes work from home aspects less easier. Akin to remotely fixing somebodies PC over standing behind them.
With that premise I'd speculate in this instance that whilst not the root cause, may of been a small ripple that led to that root cause and/or lead to a slower resolution than what would normally get.
Those speculations aside, it will only highlight what that some tooling needs to adjust for remote workers as does design and set-ups more. Water cooler talk is not just for gossip and a counter would be more regular on-line group socialising at a work level so that not only the companies but the workers can fully adapt and embrace the work medium; But so the kinks and areas that need polishing can be polished and made better for all.
Lastly, I'd speculate that I'm totally wrong and yet what I said may well anecdote with some out there and resonate with others.
When you operate at Google's scale then everything that can go wrong, will go wrong. Google does an amazing job providing high-availability services to billions of users, but doing so is a constant learning process; they are constantly blazing new trails for which there are no established best practices, and so there will always be unforeseen issues.
Yes, apps are highly distributed. Yes, roll-outs are staggered and controlled.
But some things are necessarily global. Things like your Google account are global (what went down the other day). Of course you can (and Google does) design such a system such that it's distributed and tolerant of any given piece failing. But it's still one system. And so, if something goes wrong in a new and exciting way... It might just happen to hit the service globally.
When things go down, it's because something weird happened. You don't hear about all the times the regular process prevented downtime... because things don't go down.
Sometimes it's a script responsible of deployment that will propagate an issue to the whole system. Sometimes it's the routing that will go wrong (for example when AWS routed all production traffic to the test cluster instead of production cluster).
https://support.google.com/mail/thread/6187016
Maybe time to switch to a more reliable provider.
Did you try pulling them down using the API tester?: https://developers.google.com/gmail/api/reference/rest/v1/us...
Some of the internal formatting that Gmail uses has changed over the years, so more likely than not the API that parses the stored message for display in the Gmail UI is just throwing some kind of error.
Either way my point is that this is a pretty serious bug and they haven't even acknowledged it! Not a good look.
How much do you hate it as an engineer when sales people make tech promises to customers without asking you? For comms people, engineers leaking info publicly feels the same way.
1) Harmless to share 2) Will never be shared by PR teams
I don't see anything wrong with asking people to share what they can.
But the next best thing you can do is simply just use your own domain. That way, you can at least decide to migrate your email elsewhere. Don't use the free domains you get from things like gmail or other providers, because then you have to _change_ your email address, and not just your MX records.
Services also tend to be better if lock-in effects on customers are low. ... and using your own domain for email does reduce lock-in significantly.
unrelated: the link in your profile does not seem to work.
But there is a case for making legislation forcing email providers to allow moving emails to other providers (how it should be done technically is another question). This is already in effect for telephone numbers many places.
As such, migrating these email addresses was easy enough.
My older @gmail and @googlemail ones though, not so easy. I've been moving each account I have used with these addresses one-by-one, but you never catch them all and even when you do, some services simply will not let you change your email address.
I recall being so excited when Gmail first launched and was one of the first people to get a Beta invite. I regret ever signing up for them now, given the headache it has been to get off it.
The main thing here is to avoid single point of failure as in both domains (politically induced problems) and infrastructure (technically induced problems). If people would use more than just a handful domains/providers for mail then single failures would not have that big of an impact.
I have been hosting my mail on my own domain for the past 3 years and have not been impacted by this incident. Currently I am happy for protonmail that I use to host my mails at. But I know that I can easily move on, if service drops, and even selfhost.
and the weight of google as well!
an individual mail getting onto an blacklist is most often than not a dead sentence for the address, the domain and sometimes even the ip.
but if google is at fault and email get into a permanently removed bucket, like in this event, it's in the interest of the other to play nice and accommodate for the fault.
I think people severely underestimate how hard it actually is to consistently deliver email in 2020 between dkim, spf and domain keys while tiptoeing around everyone else ip/email antispam services.
I also run my own and its been easier than expected, especially with maddy: https://github.com/foxcpp/maddy
I'd still suggest to put "secrets" in a free global trusted email provider with 2fa
I know people here dont think its a scam but it is
Spam filters have been content driven for a long time now. IP addresses and domain names are ephemeral and so are 'blacklists'. With the amount of spam being send, we would have blacklisted the entire internet by now.
If a spam filter gives false positives, it hurts the receiver just as much as the sender.
The real problem with self-hosting is that the majority of self-hosted e-mail servers are terribly configured. Getting the SMTP server running is one thing, but getting DKIM, DMARC, SPF, TLS and MTA-STS running properly is often overlooked. What was the last time you checked the validity of the TLS certificate of your SMTP server?
Get your server and domain setup properly. Sign your email with DKIM, setup an SPF and DMARC policy and perform DMARC monitoring to spot problems. Setup TLS and an MTA-STS policy service for your incoming email. Throw in SMTP-TLS-reporting for good measure. E-mail servers are not set-and-forget if you want to do it right. And this is not the fault of large corps, it's the spammers who got us in this situation.
It's really easy to blame large services or blaming your email deliverability problems on being on the same IP block as a spammer, but really it's almost always a misconfiguration on your side.
Disclaimer: I'm the founder of Mailhardener (https://www.mailhardener.com), we do e-mail hardening and solve deliverability issues.
https://www.hetzner.com/de/webhosting
Just because you can (theoretically) run your own infrastructure does not mean you should. Trust the professionals. You don't do your own surgeries, do you?
(not affiliated to Hetzner, just was the first offer I thought of.)
It's still putting all your eggs in one basket in a sense, but being a paid service there's a sense of privacy, security, and permanence that Gmail and the other free providers don't offer. I do own my own domain as well, and I have mail accounts tied to it that I use for certain services and communications, mostly medical and local businesses, but I'm still at the mercy of my hosting provider for that domain. With that said, my provider (Tiger Technologies) has been astoundingly awesome and has never let me down in 12+ years of service.
As a learning experience, sure, but most people are not prepared for what running a 24/7 mail service requires of them.
First of all, a static, non-residential IP is likely needed. The big players will flat out refuse receiption if your IP is registered as residential, so that rules out hosting it from your home despite having gigabit internet.
You also need SPF, DMARC and DKIM working, or major players will also flat out refuse reception.
On top of that, you need to implement the infrastructure to actually host a server 24/7, including patching and backups, as well as monitoring it for unauthorized access.
Despite all of the above, you may still find yourself on a spam/block list, and removing yourself from these can also turn into a large task.
Part of the irony of Gmail having outages is that Google and other "large players" have fought long and hard for a decade to make it harder to host your own mail server. It has been done in the name of fighting spam, but i doubt any of them minded it making it harder to run your own.
So yeah, build your own mail server as a learning experience. Then move the domain to someone dedicated to running it.
I purchased a lifetime subscription (limited promo offer) with mxroute.com. 10GB mail storage, unlimited domains and accounts (limited by space only), as well as a Nextcloud instance for all your users. Service and uptime has been nothing but exceptional. Customer support is actually reachable. The only downside is that the spam filter (SpamAssassin IIRC) is not as highly trained as the GMail one, so more spam comes through.
If you want to directly send mail that's true. But if you send mail through a smarthost, like your isp's smtp server, you can easily receive mail on a dynamic, residential ip.
> implement the infrastructure to actually host a server 24/7,
email is really tolerant of downtime. You can be down for hours without losing mail. The sending servers will retry for a while.
I’m aware senders will retry for days if your server doesn’t reply, but it still requires monitoring and is not just a “fire and forget” solution.
Also, if your server starts bouncing emails, chances are you’ll be missing mails. Again, needs monitoring.
Downtime hasn't been a major issue - senders will retry sending email, usually multiple times over several days. I've been able to have downtime of 24-48 hours without losing any messages.
A SPF record is just another easy to create DNS entry. If you know how to manage DNS, setting up SPF is a matter of minutes. DKIM is just slightly more complicated, with an extra key generation step. Sites like mxtoolbox.com can help you validate records.
The biggest problem I think I have with my own server is security. I do patch the machine regularly, but of course I don't have the same kind of security that Google or another big player would. On the other hand, I suspect I might have a smaller attack surface and better security than plenty of small websites.
At the risk of stating the obvious, note that 'lifetime' refers to the lifetime of the company, not the customer. Which underscores the risk of buying lifetime subscriptions.
And as much as I like the idea of avoiding recurring costs (I have a 'lifetime' Plex pass), it seems to me that these can't be sustainable for the company on the long term.
It’s really no different than Google, where a single bad comment somewhere in their vast eco system can end up getting your account banned.
In my case I try to stay as far away from Google as I can with my everyday services. I’m also well aware that chances are extremely high that any email I send will make its way to Googles servers.
The “easy” solution would be to self host, and I do that to some extent, but as I work with system administration I really don’t want/need another day job. I’m actively looking for relatively secure, privacy aware and affordable cloud solutions for everyday use. I wrote affordable because nothing is free.
"One reason is that free software gets the whole community involved in working together to fix problems. Users not only report bugs, they even fix bugs and send in fixes. Users work together, conversing by email, to get to the bottom of a problem and make the software work trouble-free."
And Service as a Software Substitute (SaaSS) takes away your freedom: (https://www.gnu.org/philosophy/who-does-that-server-really-s...)
"The basic point is, you can have control over a program someone else wrote (if it's free), but you can never have control over a service someone else runs, so never use a service where in principle a program would do.
With free software, we, the users, take back control of our computing. Proprietary software still exists, but we can exclude it from our lives and many of us have done so. However, we are now offered another tempting way to cede control over our computing: Service as a Software Substitute (SaaSS). For our freedom's sake, we have to reject that too.
With SaaSS, the server operator can change the software in use on the server. He ought to be able to do this, since it's his computer; but the result is the same as using a proprietary application program with a universal back door: someone has the power to silently impose changes in how the user's computing gets done.
Thus, SaaSS is equivalent to running proprietary software with spyware and a universal back door. It gives the server operator unjust power over the user, and that power is something we must resist."
> never use a service where in principle a program would do
Email is definitely a service.
Don't agree with this bit. You are ceding control of your data for sure, but this isn't quite the same as running a proprietary piece of software on your machine which you have no idea what it's doing.
Ceding control of data is also worrisome I agree, but giving control of your data to a custodian you trust in exchange for said custodian promising its careful curation is a trade-off that most people would find acceptable. Some won't, and I respect that, but the advocacy above might be counterproductive to most people who take it and try hosting their own email servers.
I run my own mail server, and you can screw up badly with free software. And you probably will more than Google or the big players, especially if it is not your job. Free software does not make your admin-sys screw-ups or hardware failures something you can "get the whole community involved in working together to fix problems".
(Fortunately, I haven't screwed so badly that my server started responding "no such user". Just some regular downtime that, as far as I know, has not make me miss any mail)
Using someone else's computer and services can be problematic for a lot of reasons, but uptime / reliability is generally not the issue.
-----------------------------
> One reason is that free software gets the whole community involved in working together to fix problems. Users not only report bugs
This works only for popular open source software, and still doesn't apply to infamous Unix mail servers or likes of GNOME. The 'community' is often more interested in adding features than fixing bugs.
> so never use a service where in principle a program would do.
Comfort vs Freedom tradeoff. Sometimes data privacy / freedom isn't just that critical to justify costs and difficulties of self hosting.
> Service as a Software Substitute (SaaSS). For our freedom's sake, we have to reject that too.
For what it is worth, SaaS has only benefited freedom of people working in big corporates. It has loosened the grip of enterprise software directly sold to C-suites by wine-and-dine sales methods.
And know what, software writers have to make good amount of money too. Just giving away free desktop / server software doesn't work out for most developers. I will happily accept if an open core software has a value added SaaS. And "live cheap and write freedom respecting software for some semi-arbitrary definition of freedom", is just disrespectful to talented software developers.
> someone has the power to silently impose changes in how the user's computing gets done.
Again a convinience and security etc.. thing. Most use cases, you don't care.
Stallman sees everything as black or white. (Linus Torvalds has also written on this)
I'd suggest Stallman once watch 2017 Tamil movie "Vikram Vedha" :-)
When I followed up today, I got the “this email account does not exist” error from Gmail and proceeded to dispute on PayPal.
I only found out through hacker news that this was a Gmail bug.
Google, I expect a follow up retraction of the incorrect error messages. It’s one thing to give a temporary error. It’s another to say “this email account does not exist”.
They should be 2 separate things. 2 separate css files with @media select. 2 separate button sizes and styles.
I know what happened - designers got lazy.
All of the above is general trends of how these decisions are made. There will always be counter examples or situations those aren’t good ideas (or that someone has made a mistake applying a lesson to the wrong situation).
Saying an entire group of people is lazy or dumb is not a particularly insightful way of looking at anything that helps your understanding of the situation or learning what kind of results different incentives yield.
Maybe, but when you're talking about the armies of designers at the multi-billion-dollar tech company Google not going through the effort to maintain two stylesheets, I think "laziness" is an accurate description.
Sure, but they already do all the work twice.
I get a completely different site on my phone than on a desktop/laptop.
In fact, they maintain far more than two designs - in addition to two native apps, there's a mobile website, the desktop website, and the basic HTML version. On top of that, they have multiple display density options for desktop (which admittedly is mostly just adjusting padding), redesigned the desktop site a few months ago, and had Inbox for a few years. On top of that, you can (some of these without a reload) change whether there are separate inboxes for various labels, add/remove a reading pane, and split threads into individual emails.
I don't think Google is lacking in potential to maintain a website.
Desktop (plain HTML): https://mail.google.com/mail/u/0/h/
The mobile vs desktop versions you posted are likely the same codebase with minimal (if any) differences. My understanding is that generally such things are accomplished transparently with flex layouts that automatically adjust to screen size
I doubt much work is being done on it, but presumably they at least make sure it works; I mentioned it because a few posts up (edit: you) mention testing (rather than initial design) as the reason why having multiple designs is so difficult.
> The mobile vs desktop versions you posted are likely the same codebase with minimal (if any) differences....
It's entirely possible that they're derived from a similar codebase at some point, but what reaches the browser is significantly different - it's barely responsive, based on user-agent, and appears to be significantly different obfuscated blobs of HTML, CSS, and JS.
For example, the feature set required to support the HTML page could be frozen and the APIs backing them stable with no need to change. So testing isn’t really necessary. Alternatively, there’s just API changes being made to remove dependencies on deprecated code and so the testing coverage comes from the testing that happens of that API surface through other means. Finally, it could be that the HTML page is even fully staffed to support emerging markets. That’s a different budget potentially than the budget for the “rich” UI.
Again, my point isn’t to argue over the specific business pressures and practices Google has for their email UI. This requires a level of knowledge I don’t think either of us possess. All I’m trying to do is illustrate that there could be all kinds of pressures why the system is the way it is, but dismissing it as “laziness” or “stupidness” on the part of the designer is itself a lazy and stupid conclusion to make without concrete evidence. I generally assume that’s not the case and look for the incentives/pressures those people are under until there’s overwhelming evidence those people are actually stupid/incompetent (and even then, the question becomes what structures, incentives, pressures were in place to put those people in positions they shouldn’t occupy).
It used to be that you could check the width and height of the viewport and say something like “320px wide? Must be a touch interface, deploy the big buttons”. Then tablets got big and it was like “1024px? Could be a laptop, but it’s probably an iPad, which has a touch interface, deploy the big buttons”. Then laptops got touch screens, then the Surface Studio came in and was like “HAHAHAHA”.
Now the game is “1920x1080? Could be a big tablet with touch, or a 1080p monitor without touch, or a non-maximized window on a Surface Studio with touch, or maybe it’s a monitor without touch hooked up to a laptop with a touchscreen and our window could get moved between them at any time...”
Nowadays, there’s no single reliable way to tell if a page is going to have to support touch until it gets a touch event, by which point you’ve already rendered the UI and it’s too late.
[0] https://developer.mozilla.org/en-US/docs/Web/CSS/@media/poin...
Not easily so on mobile - the need to zoom multiple times and scroll in 2 directions to read simple information is a PITA.
Money
Good thing gmail wasn't completely stupid and didn't try to forward that one, getting in infinite loop...
Other people’s automation scares me. I’m sure mine scares other people as well.
If anyone from AWS SES is reading - please do not deliver bounce receipts for gmail for the time being - it makes everyones situation much worse.
Edit: oops thought it was a text post
An interesting problem - I have a gmail address, and also one on my personal domain (which uses g-suite for email). The personal domain's backup is a gmail address, and the gmail's backup is my personal domain. When I set that up ages ago, I genuinely never thought about what happens if gmail itself implodes.
I guess I'm setting up a... icloud??? backup email for a bunch of stuff shortly.
> We will provide an update by 12/15/20, 10:30 PM
> We will provide an update by 12/15/20, 5:30 PM
I agree that a uniform time zone would be clearer.
The images on different websites I've visited don't load (e.g. twitter, bbc etc). When they do they load, they load verrry slowly
Use dnscrypt.
Plus, how would he send tweets if no internet? I think he likes being able to use Twitter more than he likes being President.
For whatever reason, gp seems to be under the impression that it's Trump who would use this to his own advantage, which is not consistent with the prevailing conspiracy theories.
The "internet" deserves these outages to make people - and CEOs, CIOs, etc - realize that in-house ~~engineering~~ sysadmins used to exist for good reasons.
This blew away about 10% of our newsletter and marketing subscribers. I can't imagine the time you're having if you send to millions of Google accounts.
Pardon me for being conspiratorial, but I have to say that the timing of these particular issues is of concern in light of the massive attacks the US government and others have faced this week. It seems adversaries would have plenty of fun if they got a new toy that could mess with the user/resource permission mapping at Google would want to use it to go after inboxes to do email confirmations. Even those this is 99% likely to be a DevOps chore that created an SRE nightmare, I'll allow myself believing 1% this was SecOps locking down huge swaths of accounts to mitigate a mass email verification attack.
Still, I have good enough reasons to think that some sort of Pandora’s box of backdoors was opened this fall and its fallout is yet to be felt.
Thankfully, we have backups, but we will have to move them to the inbox or elsewhere to have them processed.
Edit: As an update, we usually have at least 100 emails come in every hour, and I am seeing none since 4:02 pm EST
Not a good week for Google
For business critical operation, leave such answers unanswered can be pretty risky.
I know for example companies that use email for handling customer orders at their stores. I can just imagine the loss if the system was down for hours a few days before Christmas, or worse, actually lost the data.
Yes.
under 99.9% is a 3 day credit, and they're currently at 99.4%
I use it for privacy (am a fan) but I feel pretty smug knowing I’m getting better reliability too. At least this month :)
No affiliation with the company.
Is your data backed up anywhere in case your Helm box burns in a fire?
Very surprising it’s not localised.
That's a considerable difference.
Intl.DateTimeFormat().resolvedOptions().timeZone
> "Australia/Melbourne"
Which is UTC +11. new Date().getTimezoneOffset()
> -710
Which is roughly UTC +11.8.If you're going to shoebox me into a particular timezone - tell me what it is, and let me change it.
Edit: Apparently it's browser settings, which means you need scripting enabled for this page.
> 550-5.1.1 The email account that you tried to reach does not exist. Please try double-checking the recipient's email address for typos or unnecessary spaces. Learn more at https://support.google.com/mail/?p=NoSuchUser x62si100799otb.139 - gsmtp
[1] https://www.reddit.com/r/Stadia/comments/kdr2ps/its_not_just...
This outage has now convinced me I need a new email provider.
IMAP seems to be working, though.
Could be inconsistent though.
All good now apparently
By the end of the thread I wondered why it was not affected, until I remembered my small business email is actually on Exchange 365 :-D
So if you want to be able to receive emails from smaller mail servers - don't switch to Microsoft.