Dodging S3 Downtime with Nginx and HAProxy
blog.sentry.io
blog.sentry.io
# matches /s3/*
location ~* /s3/(.+)$ {
set $s3_host 's3-us-west-2.amazonaws.com';
set $s3_bucket 'somebucketname'
proxy_http_version 1.1;
proxy_ssl_verify on;
proxy_ssl_session_reuse on;
proxy_set_header Connection '';
proxy_set_header Host $s3_host;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header Authorization '';
proxy_hide_header x-amz-id-2;
proxy_hide_header x-amz-request-id;
proxy_buffering on;
proxy_intercept_errors on;
resolver 8.8.4.4 8.8.8.8;
resolver_timeout 10s;
proxy_pass https://$s3_host/$s3_bucket/$1;
}
Adding NGINX caching on-top of this is pretty trivial.Also, heads up, in the directive proxy_cache_path, they should consider enabling "use_temp_path". This directive instructs NGINX to write them to the same directories where they will be cached. We recommend that you set this parameter to off to avoid unnecessary copying of data between file systems. use_temp_path was introduced in NGINX version 1.7.10 and NGINX Plus R6.
use_temp_path=off
Also, they should enable "proxy_cache_revalidate". This saves on bandwidth, because the server sends the full item only if it has been modified since the time recorded in the Last-Modified header. proxy_cache_revalidate on;This is vulnerable to path expansion attacks. If someone passes a URL such as your site/s3/..EVIL_BUCKET/EVIL.js all of a sudden your site is serving someone else's content. Bad idea. Use virtual host style buckets instead, i.e S3_bucket.S3host/content.
One sort of weird case is if I have an image key (sha-based) and want to store thumbnail sizes: 'bae6ff187e4c491e5de9cfa3b039ce7da8255798' makes sense as a base key, but really I want bae6ff187e4c491e5de9cfa3b039ce7da8255798/400x400 for thumbnails rather than storing individual thumbnail shas, hah.
One argument for self hosting the proxy is that I don't care if s3 is working when my server is down anyway.
Remove the C in CAP and you can go far.
What is the name for this phenomenon where folks think they can out-available a thing that has multiple engineers singularly dedicated to nothing more than its availability /and/ operation? It it just hubris? Surely there must be a more clinical name.
Additionally, not keeping all your eggs in Amazon's basket means you're not SOL when they have a datacenter hosting all your content go down. It also means that if and when a service that better fits your needs comes along, you are more readily able to migrate without problems.
Finally, site reliability is not something that takes a team the size of Amazon's. A lot of the things that are required for availability on AWS -- redundant systems providing services, standbys, etc -- are things that sysadmins were doing before AWS was extant. AWS' biggest gift to reliability is that its instances are less stable than most dedicated servers; you're taught from day one not to rely on a single server, so you build it right the first time.
So no, it's not hubris. It's calculating price/performance, it's applying things you're probably already doing to a new problem, and figuring out what the best solution really is, which rarely involves just throwing money at Amazon.
We gracefully fall back to S3 directly if our cache server is down without a hiccup. So there is no operational overhead of this additional cog. If the server has a failure, we'd go back to slightly degraded performance by talking across the country until we brought it back online.
If that's an earnestly literal statement from you, then it means you simply haven't encountered the failure modes that these kinds of set up are inclined towards.
I've worked at several BigCo's, seen them all implement this pattern, and seen every single one of them have fleet-wide outages due to these innocent "local proxies".
Remember FB's 2-3 hour outage 2 years ago?
https://www.facebook.com/notes/facebook-engineering/more-det...
It was /exactly/ this kind of "local proxy for higher availability/caching over the downstream thing" that caused the outage.
And for what it's worth, I've definitely fucked this up in the past and caused downtime as a result of a setup like this. The pros still outweigh the cons in practice.
Who deploys the thing?
A human? God knows they can screw it up.
Automated deployment?
Well, that's how you get a simultaneous failure and total outage.
Automated incremental deployment?
Ok, slower road to total outage.
Automated incremental that will halt itself or rollback based on reliability metrics?
Ok, getting there.
Wait, was the local proxy load tested?
Was it load tested when one of your data centers is down and everything is doing 30% more work?
And on and on and on. It's all operational overhead, it's all ways to fail.
Can you tell I used to work in monitoring? Maybe I just have PTSD now. :P
Correct, but it's an existing process. So you're right, we could ship a blatantly bad config.
> Who deploys the thing?
We do, humans, yes. We can definitely screw up a config.
> Automated deployment?
We tend to do blue/green deploys on critical pieces of infrastructure just to sanity check it. We might even pull a node out of production, test on a staging server, etc.
> Wait, was the local proxy load tested?
Yes. The load we need for this case is not even close to significant.
> Was it load tested when one of your data centers is down and everything is doing 30% more work?
Yes, it's literally just a proxy to S3 doing no additional work. For our traffic, the load is not a concern. Especially since it's running on every machine, it's distributed pretty well. A single box cannot overload our haproxy process compared to the CPU needed to run the Python application itself.
> Can you tell I used to work in monitoring? Maybe I just have PTSD now. :P
DataDog, it's pretty dope. It gives us lots of super good insight into all of these things and is what alerted us because haproxy reported S3 down in the first place. It'd also tell is the moment a process like this crashes, etc.
It all comes down to risk vs. cost vs. gain.
Also worth noting, that this isn't really a single point of failure as a system wide thing. It'd only be a single point of failure on that single node. So if haproxy decided to explode, only that one machine would have a problem momentarily, while the process got started back up with our process manager.
The worst case scenario is a human error where we ship a bad config and break everything.
Haproxy sends to caching nginx if available, else directly to s3.
With AWS you don't have any control over "higher risk" times. If you have a massive launch coming up, or you are nearing peak usage for the year, or your clients need you to be stable for the next few months, you can't put updates on hold, you don't know to get a few more people on standby, you can't choose to not make changes to your system, because it's not your system.
With an in house solution you can choose to lock it down for a month, or do the risky upgrades/changes at your lowest traffic time, or even give your customers a heads up if needed. Hell even just being able to mak e sure that your best sysadmin isn't out getting hammered when you go to make changes could go a long way.
There is some merit to that idea, but I personally feel the track record of many of these services is so near perfect that the chances of unexpected downtime is still smaller than most could realistically manage.
Do you host your own DNS?
Your analysis assumes all other factors are constant. Change causes downtime, cloud services have a high rate of change (constantly pushing new configs, new code), many other servers don't.
Most importantly, while S3 is relatively stable it's a black box. If you really care about high availability, you want to bring all of the points of failure under your direct control. On the other hand if availability is just a nice-to-have, relying on S3 is probably a better use of time.
This is what exactly I'm talking about. People can't accept that trying to control it is not in any way guaranteed to make it more highly available. Bringing points of failure under your control makes no innate guarantee of improving anything. It can easily make it worse.
The only thing it can do is let you blame yourself of blaming S3.
You can only be hardened for what you have anticipated or experienced before.
There are tens thousands of wall clock hours of operational experience w/ S3. Availability is one of the top concerns of all AWS.
Thinking you can be more available is just fooling yourself. Believing you are more available will entail a willful ignorance or distortion of metrics.
Do your engineers carry pagers and have a <15 minute engagement time? Do your engineers sit at home when they are on-call because they know they can't simply let a page slide because they were in the middle of dinner? Or is your company more lenient than Amazon when it comes to operations?
Do your engineers spend a quarter fixing some failure mode of your infrastructure, or are they too busy working on features?
Is your team's performance measured by availability of your service? Or your actual core business?
I don't think I could run an S3 service at the scale AWS does with higher reliability, but I do believe I can run a pool of redundant HAProxies with higher reliability, and in fact we already run pools of HAProxies for other reasons so we have quite a bit of operational experience with it.
If my company had three engineers I certainly would not go this route, but if you are big enough to have a dedicated ops team that already has experience with this sort of thing, you can architect something that is more reliable than just relying on S3 alone.
I think the fallacy here is that you're not comparing apples to apples: I would be the last to argue that I could run a globally distributed S3 competitor better than Amazon. But I can (and have) run a massively simpler service with better overall uptime because it increases our options during upstream black swan events.
I'm not saying it's impossible. But I am saying it's dangerous to omit from this conversation the idea that introducing the very --point of option-- can cause worse reliability than just using the downstream thing in the first place.
We also have less than 15 minute incident engagement times, and don't let important pages slide through dinner. It's totally standard ops stuff: if one of the servers is down, we'll replace it when we get around to it. If they're all down, pages are going off.
If you were looking to protect yourself against S3 going down and had bigger drives, I assume you could use cron to sync the entire bucket and use `try_files` to prefer local and fall back to S3 if the file was missing?
I know of a datacenter in Denver CO with 14 YEARS of uptime, for instance.
Whereas - if you just rely on S3 directly, if they have problems, there's not much you can do unless you also have all assets locally on your servers.
S3 is an object storage system, they're adding a proxy. It's pretty easy to make a proxy that has better uptime than S3 because it's far, far less complex.
@mattrobenoit -- neat idea and the 70% savings in bandwidth is awesome. The side effect of helping mitigate the S3 issue for you was a sweet little bonus!
Talk about great timing!
One solution that I haven't seen much of is to just use a service that gives you multi-region without any extra work, such as https://cloud.google.com/storage/docs/storage-classes#multi-...
Also from the public description this sounds like one big system, ie the second "region" may not be a public Google Cloud region. Just as much chance of an outage,
As with most services, you pay as you go and only for what you use. The prices are aligned with typical AWS storage, request volume, and data transfer costs.
Obviously this method has cost associated with it, so you should probably only do this if you need complete data availability.
And yeah, we have 0 resiliance for write data here. Again, fortunately, we can afford this tradeoff since the amount of uploads is significantly lower and much less critical for us.
Well, when the primary goes down, your write operations would get busted.
If the upstream source allows it, just write to both buckets at once. E.g. with Logstash this is trivial.
With S3 replication, you still have a primary/replica setup in which only one of them can accept writes, but you can accept reads from both. So we'd gain HA between multiple regions, but we wouldn't solve our original goals: speed. The round trip to S3 was too slow for us.
Another idea might be to use Varnish for the caching layer, but I haven't compared Varnish to NginX in many years so the gap has probably been closed now?
Good work. I've stuffed this one in the back pocket for future use.
And in our case, performance between reading some bytes from disk vs memory isn't significant. A disk seek is still many many orders of magnitude faster than a round trip to Amazon.
With that said, Varnish does offer the ability to use mmaped files, but the performance is really appalling out of the box, and just not worth it. Varnish is way better if you want strictly in memory cache.
Another benefit of nginx is the cache won't be dumped if the process restarts, unlike Varnish.
Of course, if you have complete control of the client you can just change the hostname. But if you have references scattered throughout html files, you'll likely want a reverse proxy in front of S3.
Is HAProxy there to get more insight into things, like the slack notification? Or does it serve another purpose?
Is the next step locally cache uploads so that works while S3 is down?
Back in the day, every production web site had something like this "HAProxy" in place.
Now, S3 goes down, and everybody has their thumbs up their asses because that wasn't covered in the Rails bootcamp they went to.
It's not uncommon for applications to have MRU access patterns and be able to keep partially functioning during partial data availability. For these applications, a cache will lower costs and mitigate S3 outages. It would have been nice for the article to give the criteria up front.
That paragraph says your connection to S3 is slow. Your solution doesn't fix general S3 slowness. However, it does make S3 slowness less of a problem for your application.
That's why it would have been useful to describe how your application accesses S3, so others can quickly determine if their applications would also benefit from a similar solution.
For example: "90% of our S3 reads are of blobs that were read in the last 30 days."
Many applications might upload data to S3, then quickly download it for processing. For those applications, this solution won't work.
> But I never set out to mitigate S3 failures like this.
We know. It's just that some people might read the title and think this is a more general solution than it is. OP was just warning people against that.