Google App Engine Silently Stopped Sending Email 5 Weeks Ago
code.google.com
code.google.com
The issue reported here is linked to App Engine and Gmail tightening up their spam filters. The root cause was an increase in organizations sharding out their spam systems to utilize App Engine’s free tier in such a way that is (a) in direct violation of our ToS and (b) making all of our lives suck a bit more (raise your hand if you want spam). It’s unfortunate that while App Engine is trying to provide a free tier that enables developers to easily use our platform, others see it as an opportunity for exploitation. Even more unfortunate is that it has a negative effect on legitimate users. It’s a fine balance that has been highlighted by several users within this thread.
Spam filtering is not a perfect science, and we’re constantly tweaking things -- with our customers in mind. This issue should be limited to new applications where the trust signal might be a bit lower. Thus existing apps / customers shouldn’t be experiencing issues (which was also highlighted by a few within this thread). If this isn’t the case email me: cramsdale@google.com. For those asking, “hey, why am I being penalized for being a new customer?” See my previous comment about spam filtering not being a perfect science. Then email me.
We’re here and we want to help.
-- Chris (Lead PM for App Engine)
The best part for me when using it for customers has been the bounce/click through rate tracking. When someone asks me, "how do I know they got it?" It's incredibly nice to point them to the dashboard and show a less than 1% bounce rate (people putting in bad email is almost always the reason) with a log of every single email sent.
Most my clients get this service for free because their volume us low enough. They have quite a generous free tier.
Anyways, it's up to Mandrill to choose who they want to serve. I really liked the service though. Best of Luck.
edit: Not to say Mailgun isn't fancy, I have no idea if it is or not, I can say that it works in minutes though.
> This issue should be limited to new applications where the trust signal might be a bit lower. Thus existing apps / customers shouldn’t be experiencing issues.
1. My app is not new. It's been running without issue for 2 or 3 years.
2. My app is not on the free tier.
3. On average, my app sends out under 8 emails a day, and it's done the same for over 2 years.
How your algorithm considers my app a spam risk sure beats me.
I'm missing 11 days worth of quote requests from customers. Are these recoverable? (is there a hidden outgoing-spam bin?)
My app started sending mail again today after I changed the src of an image in it from https://example.appspot.com/images/logo.png to http://www.example.com/images/logo.png
How well do you think email clients are going to like emails with images embedded in them without HTTPS?
I sent him an email 14 hours ago and haven't received a response. I'll wait a day and if I still don't hear from him, will contact you.
The numbers Mandrill released about that business suggest that they had a large number of low volume senders, who may now be looking for a new home.
Or are you afraid that "rogue" applications will use it to produce messages that are SPAM but yet not trigger the SPAM filter ?
As an email service provider it's like that and more, since (1) once an abuser uses your service, they've gotten the benefit immediately and keep it even if their account is discovered as fraudulent later, e.g. stolen CC number & chargeback. (2) Abusive users can directly harm good users such as by harming the deliverability of the overall platform. It's not just bad debt, it's bad experience too. (3) Unlike Candy Japan where fraudsters mostly just wanted to check CC numbers and not actually buy product, email abusers really want to send emails (4) It can be hard to tell good and bad senders apart because some companies with an internet presence aren't email savvy and might make mistakes or might get hacked.
Spam filters are always tough because if you give someone transparency into which actions of theirs that you consider abuse, then they will quickly detect and route around your attempt to block them. (See Candy Japan article) It's pretty easy for a human to guess what might be the sign of their fraud and run a few experiments to see what gets flagged e.g. By comparison a machine learning system might be hard to outsmart, but then it's also challenging to explain and troubleshoot false positives. Hence what's effective is often a combination of machine-learned filters and heuristics along with manual overrides by human judgment.
All other things equal, new users are a lot more likely to engage in fraud than existing ones, and so tend to be under more suspicion. Aside from B2B fraud where companies take out lines of credit and then go bankrupt intentionally, it's uncommon for existing established customers to turn fraudulent - they're already vetted. (Consider: who is more likely to be fraudulent. The first time subscriber to Candy Japan, or a subscriber who has been using it for 12 months and is about to buy their 13th month?) It's not a great experience as a new user to be under suspicion, but if it's temporary and easily overridden by a human it can be a decent trade-off - the need to reach out acts a deterrent to spammers but does not deter legitimate users as much (speaking generally).
I've noticed the spam filter on Gmail for Google Apps has gotten significantly less accurate recently, resulting in far more false-positives than usual. Any ideas if this is a known issue? I can only presume it's more accurate overall, but it was definitely a noticeable change for our organisation.
Also, App Engine is integrated with both Sendgrid and Mailgun, and we strongly recommend using these if you're planning on sending larger quantities of mail:
https://cloud.google.com/appengine/docs/python/mail/sendgrid
Whether it's your outgoing spam filter that's overzealous, or if a datacenter is on fire is irrelevant. Your infrastructure stopped doing what your docs said it would.
Paying customers like myself have apps running on your infrastructure. Your infrastructure stopped delivering the exact same emails that it had been delivering for years. It did this without notification and without triggering any errors.
Approximately 35 quote requests that should have been delivered have disappeared into thin air. You've cost me money, and I'm disappointed to find out that you'd been notified of this issue weeks before it began affecting me.
At the VERY least you could have easily set up a gmail account and sent them from smtp. Choose the right tool for the job.
I partially agree with this - dedicated email services have a higher deliverability rate than a random web server, especially a cloud-based server that might be using an external IP that was previously used to send spam. However, I can understand the parent poster being annoyed that their emails were being delivered one day, and not the next, without any changes on their end.
> At the VERY least you could have easily set up a gmail account and sent them from smtp. Choose the right tool for the job.
This is absolutely the wrong tool for the job. You should not be using a Gmail account to send out automated emails that you care about.
Agreed, but better chance of delivery than directly from the server.
I've noticed that they've added more support people since the early days but it's still a pain as they seem to have an incentive to answer the tickets as fast as possible without doing any real investigation.
The linked issue is a good example of this behavior.
If you're running production services, you should probably sign up for a paid support package, which have guaranteed response times down to 15 minutes: https://cloud.google.com/support/
Disclaimer: I work in Google Cloud Support.
More often than not, what most people want is simply an acknowledgement. They're not seeking a bug bounty. And definitely not the brush-off you just gave the GP: when someone's house is on fire, you don't demand prepayment for the water.
You absolutely do not have to pay us to report a bug, period. The same day this bug was filed, one of my colleagues (pay...@google.com) jumped on the report and started asking questions to determine what was happening.
It seems to me that the GP's complaint is that this process was too slow: too much back-and-forth, too little dedicated attention to get the problem figured out immediately. That's a frustration which is easy to understand.
The point about support contracts was most likely intended to emphasize that if your livelihood depends on a service, you should have an agreement in place that guarantees you can wake up an engineer on the weekend.
the "free tier" of municipal services is actually just a form of socialism!! eek!
I feel little bit bad for them sometimes. They don't seem to have the time to look into problems. It must be quite a boring job for those assigned on the public issues.
Maybe their higher-ups aren't aware of that.
Form my perspective: sometimes when I report a bug it's more so that other users don't waste their time troubleshooting it than to have it fixed immediately.
Maybe sometimes I'm not 100% clear but it is a bit discouraging to have to send follow-up messages on clear-cut issues like these: https://code.google.com/p/googleappengine/issues/detail?id=1... https://code.google.com/p/googleappengine/issues/detail?id=1... (the comment #10 of the former issue was written a bit too quickly don't you think?)
Then there are issues that "work as intended" but that seem like bad product design: https://code.google.com/p/googleappengine/issues/detail?id=1...
Do they really reach a product manager?
Finally there are issues like these ones that are not clear cut but that I can't investigate myself because it requires time and resources (they eventually cost me a few bucks running instances for testing purposes). https://code.google.com/p/googleappengine/issues/detail?id=1...
You guys have a good product. I think that the promise of not having to manage infrastructure hit a cord with many people. However you do have many bugs to fix to make the platform more stable and (non intentionally) discouraging people from reporting issues will make things improve at a slower rate.
Good luck!
No. (Or they don't care.)
$10-20/month VPS providers often respond to, and actually take action, on tickets within 5-15 minutes at no extra cost. Of course there is no guarantee, and a ticket can certainly wind up unresolved for 24+ hours. In my experience, if the ticket you file clearly explains your issue and represents something they can realistically help you with, the level of service provided for such a low server cost can be exceptional.
I'm not sure what you mean by this. Every AWS service I can think of has "free tier" for lower use and then charges for higher use, like AppEngine. For example, AWS's email service (SES) is free up to 60,000 emails per month.
To me, the whole point of Appengine is to abstract out system administration. GCE (like AWS) abstracts out the systems, but not the administration - you suddenly become responsible for all the scaling, fault-tolerance, and all the subsystems that Appengine manages for you.
I've tried Containers with Google App Engine Flexible Environment (in beta now) and liked that it's much more customizable than standard GAE. Basically, you just need any Docker container listening on port 8080. But deploy time took 10-20 minutes, so I'm still preferring Standard. Hopefully they'll fix that before GA release.
My request is for a tool which would provide this functionality in the form of code which exposes an App Engine like interface to the developer, except it's running in Container Engine. Given Container Engine has the required code to restart processes that fail and oodles of other automated goodies, this give the fault tolerance desired.
> We wouldn't use it.
Given the "tool" doesn't appear to exist IRL, that statement is illogical. You can't not use something that doesn't exist. And, if there really did exist a migration tool that provided additional operational abstraction inside the framework you were migrating to, you'd run the risk of being in conflict if you didn't consider using it.
http://penguindreams.org/blog/how-google-and-microsoft-made-...
I have seen more phishing/ransomware e-mails come through mine (lots of malicious Javascript made to look like Excel files), so I can understand how great the risk is. The fact still remains though. E-mail is broken and is less reliable than an envelop dropped in a letterbox.
I don't know about Microsoft, but I get a couple of DMARC reports per day from Google's email infrastructure. I also get reports from aol.com, yahoo.com, linkedin.com, comcast.com, fastmail.com, etc.
It was sending an HTML email with the appspot domain used for the logo: https://example.appspot.com/images/logo.png
Changing that to http://www.example.com/images/logo.png fixed the problem, but now I have to host the image elsewhere on a HTTPS enabled domain.
A separate question is what has google done with all the messages that users of this app sent? Are they gone for good, or stuck in a queue somewhere?
Of course the support organisation here (Google) should be well equipped to set up this kind of testing themselves and work with a customer to root cause the issue. In my experience though these issues are rarely taken seriously until there is no-one else left for support to blame.
This thread is a great example of why support is so hard.
If my customers don't get a reply within a few hours, then all hell breaks loose. Even if, there request comes in at 3am / 0300 on a Sunday morning my time.
To be fair, when I point this out to them they do apologise, didn't realise time difference, are in a different part of the world where Sunday is not a day off etc. I then ask what they expect from some billion dollar companies support wise and they say they are happy with a reply within a week. Asking them why a one man shop doing quite a bit less than billions of dollars is expected to provide so much more.
Anyway, rant over. Support isn't hard. You hire enough competent people as your customers require. The end. Happy customers
Email is hard.
Running a mail server with reliable outbound delivery, and inbound spam filtering, is anything but simple.
How is this any harder than anything else in computing?
At work, we handle pretty large traffic for mail, with self-hosted mail servers as well - those are exim. No delivery issues either, although these are only the real, important, not even remotely spam-like, core business messages. Spam-like newsletters are done via sendgrid there.
The GAE docs say "do x and we'll send your mail". If we do x and mail isn't sent, then that is a problem.
I agree with you that some measure of notification that there's spam filtering would be good, but it's Google we're talking about. There's spam filtering.
Nonsense! I had no idea this was a feature of GAE. I pay for the use of GAE, and so long as I send less than my quota of emails, it's not Google's business what the content of those emails are.
There's no reason that there should be silent spam filtering on outgoing mail sent by my app. If they suspected that my app was sending spam, they should have notified me and disabled my app. That would have alerted me to the problem.
I find it absolutely abhorrent that Goole would just send private communications sent to me by prospective clients into the trash where they can never be retrieved or restored, and to do so without telling me is beyond belief.
The point: G is accepting cash from one group of people to send a ton of messages (mostly unwanted by addressees) and earning ad dollars from another group to stop unwanted messages.
30 years ago this would be textbook conflict of interest but the 'veil of automation' means we're blind to a lot of unsavory business practices.
That's been tested for and eliminated. My app logs that an email has created and is going to be sent. The send_email function is called to send the email. Then my app logs that the email has been sent.
The emails are sent to addresses hosted at both gmail and outlook. They are not received. They do not appear in spam bins.
GAE has a bounce api that will log if an email is bounced back. In my case, no reports are logged by my bounce handler.
The content of my emails has not changed for 2 years. For 2 years both gmail and outlook received my emails in their inboxes (not spam bins).
Given all this, I'm pretty sure there is an outgoing spam filter.
You are very wrong about this. Senders of email need to make sure they aren't being used for spam or they get blacklisted.
But then they should have notified me: "We suspect you're using the service in violation of our TOS, so you are suspended" or something like that.
I'm paying for a service, they continued to take my money while pretending to provide that service.
Nothing whatsoever changed in my code. I haven't touched it in a year, and all of sudden email stopped on the 1st of April.
There's nothing in the changelogs to indicate there was a change that should cause this behaviour. It's absolutely an App Engine problem.
https://cloud.google.com/appengine/docs/python/release-notes
https://cloud.google.com/nodejs https://cloud.google.com/ruby https://cloud.google.com/python .... etc
If you have any questions at all about the service, feel free to ask me anything.
Also, why the long delay? Is it because of the switch from C to Go in 1.5?
l.Infof("version=%s", runtime.Version())
...and it gives me:
version=go1.6 (appengine-1.9.36)
That's not true. We continue to invest heavily in App Engine as a platform.
"A commitment to Moore's Law: https://cloud.google.com/pricing/philosophy/ "
-- Chris
They're not the same product, and in this case they serve different points in the space with different underlying requirements. App Engine Standard's cost structure on the backend just hasn't changed nearly as much as compute engine's lately where it's a much more direct: oh look Haswell released, more cores per host at similar dollars yields less $$/core.
Disclosure: I work on Cloud (Compute Engine mostly) and care a lot about our prices.
What does this even mean.
> Just because some company moves a lot of bits doesn't mean they know what they are doing.
You have to have some measure of competency to be able to operate at Snapchat scale. You really do.
> Fuck just run a rack of open relays and bad NTP/DNS servers.
You really have no idea what you're talking about.
>What does this even mean.
I find their app to be very poorly designed, I wouldn't be surprised if their back end is also poorly designed.
>You have to have some measure of competency to be able to operate at Snapchat scale. You really do.
Meh, they are big but they aren't that big. There are about 3000 bigger sites out there and not that many are built on GAE so that's a strange way to measure the competency of GAE or snap chat for that matter.
You can run at scale and do it poorly, Enterprises do it all the time. Size and scale aren't really a good indication of quality.
If you want to make a case for appengine, do that, but don't tell me that 'these guys vetted it, and they are kinda big, so it must be good' There's a lot more big sites that don't use app engine. Size doesn't equal competency.
Or not.