S3 was down
status.aws.amazon.com
status.aws.amazon.com
https://support.cloudflare.com/hc/en-us/articles/200172256-H...
This downtime made me realize we weren't caching any html on cloudflare. I just turned it on and all our static sites are doing fine now (and our bills are smaller!).
If you're fancy you can even programmatically purge the cache when you do CI deploys using the cloudflare API.
EDIT: Oops. Misread cloudflare/CloudFront.
For inner pages, browsers do consider /your-page-url?v=1505418887 to be a completely different page from /your-page-url?v=1505418888.
However, just expiring cache and/or properly setting HTTP cache control headers at most CDN's is a cleaner and more correct option.
I was just pointing out that some people adjust their menu links (etc) to include some sort of dynamic variable (ie /?ts=xxx). I don't recommend this sort of scheme, though, except in unusual circumstances: IMO, it's fragile, leaky, and inefficient.
Thanks for replying, but unless you're willing to explain to the rest of us why so, why bother? I can't make heads or tails of this pissing contest thread.
"Site Delivery" is purged globally within 7 minutes.
"Fast purge" enabled products (site delivery) is purged globally in under 20 seconds.
All assets for websites are site delivery. Only media services family of products are not site delivery. Those purge within a couple of hours. If you are someone who uses media services products, you know how to force Akamai to go back to origin globally instantaneously and how to trigger invalidation. In media services you should never serve stale.
Where it doesn't help is if Cloudflare itself is having issues (all it does is move the weak link one layer up in this case). But that is easy enough to disable (assuming the cloudflare API or dashboard is up).
However, I do agree that it should be possible to turn on without pagerules for these sorts of scenarios.
However, the case of a static website is not really hitting me as something can't afford a few seconds downtime...
> If you're fancy you can even programmatically purge the cache when you do CI deploys using the cloudflare API.
You need to be really careful with this.
I'm not sure how you've got things set up but isn't this going to lead to issues unless you have a cache busting or revalidation strategy? You can purge Cloudflare's cache but that's not going to purge the cache on browsers that already visited your site and cached a page already. You might get cases where an old cached HTML page is asking for page resources that don't exist anymore or pages break because the user is seeing the old HTML with the new CSS/JavaScript.
Also, if you can log in to your site, it can potentially cache one user's logged in page content and share it with others if your Cache-Control settings don't include "private".
You could set things up to make sure the browser and then the Cloudflare cache always asks your server first if a page has been updated recently (revalidation) but your server has to be configured right for this (e.g. etag or modified date usage).
Cloudflare is actually meant to have a feature that keeps a static version of your site to show in the event of outages:
Ref: https://support.cloudflare.com/hc/en-us/articles/202238800-W...
Edit: It's not about bragging. It's not about the ROI. I want to (1) experiment & learn, and most importantly (2) show what is possible with very simple technical architectures. HN is the ideal place to "show and tell" these sorts of projects.
edit: punctuation
But yeah, you are right it's a very, very simple site.
That's almost certainly lower than combining reputable VPS providers geo-redundantly, and likely a much higher cost for the "convenience".
And it's not just for geeky pride. You learn the most when things break. Far too many funded startups run poorly architected apps in a single AWS EZ, unlike GP.
GP is "grandparent", the poster who started the line of discussion. (mrb)
My browser won't automatically try 2.2.2.2 or 3.3.3.3, or would it?
Any HA solution I've seen that attempts to reliably achieve this five nines capability relies on network-level things like virtual IP's and what not. And I don't consider it a five nines solution if only some customers can access it. How the browser behaves in this case could be critical depending on how your visitors use the site. I would not consider a site "up" if it's only available to some people and not others.
Well, that depends on the SLA/SLO, which is really what "nines" is speaking to. Intuitively I agree, but it can, realistically, not be the case and be "valid". Doesn't make it right. Just is.
Last time I researched this, behavior was quite different across the board, and it's something one should test extensively when designing HA for HTTP. In some situations, the same browser on another platform will defer to the system resolver versus its own, for example, which will potentially change behavior #1 even for the same browser. Mobile is starting to perform weird tricks with TCP, too, so you really have to dig into this one to do it right. Then throw in HTTP/2 and you've magically created yourself about a decade of justifiable work ;)
On the other hand, I'm feeling a strong case of "Not this shit again". Wondering if US-East-1 is more trouble than it's worth, as these outages seem to happen mostly there.
For example, it's likely that the actual master data stores for IAM are solely within us-east-1, but the key data is cached and services run in each region.
Similarly, Cloudfront is theoretically a global service, but only ACM certificates set up within us-east-1 can be used with Cloudfront.
Not true at all. It'll be among the first for newer features, in as much as the US regions usually are among the first, but it's certainly not the first place code is deployed to.
All services in AWS start out software deployments in smaller regions where the blast radius of a software bug is likely to be smaller.
Looks like November 17 to June 14.
[1] https://aws.amazon.com/about-aws/whats-new/2016/11/amazon-sq...
[2] https://aws.amazon.com/about-aws/whats-new/2017/06/amazon-sq...
Oregon sees almost the same speed of adoption of new features with way better reliability.
Happy to answer any questions about the tool, either here or over email.
(Disclaimer: I work for Gremlin)
[1]: https://twitter.com/AlecSanger/status/908402829349572608
It's also hosted on S3 but still up and running. (The main service is independent of S3 anyway after setup)
Spun up this few hours ago, will be gone soon but if anyone wants to check it out without waiting for the metrics to trickel in:
http://pwd10-0-33-3-3000.host3.labs.play-with-docker.com/das...
[0]: http://principlesofchaos.org/ [1]: https://github.com/Netflix/SimianArmy/wiki
Please contribute if you know alternatives.
I've seen a handful of test libraries but none of them seem to make realistic up-to-date error injection a priority.
they have multi-az support if you use it
I think you really meant s3 objects are redundant in each region, which actually is spanned across multiple DCs.
Active/active cross-region is possible, but far more complex.
https://www.openstack.org/software/releases/ocata/components...
If you're serving a static site out of S3, put Cloudflare (ugh) in front of it and enable the feature to serve your cached content when the origin is down [1]. When S3 goes sideways, as everyone is learning, it goes down hard and you're probably not going to be able to make changes to objects, read object metadata, etc.
[1] https://support.cloudflare.com/hc/en-us/articles/200172256-H...
Our Airflow tasks that use S3, they are all down.
Basically, we are down :)
Also, we weren't like all down, we just saw lots of time out issues when reading/writing to S3.
Very interesting re redshift => google analytics though, I've never heard of it done in that direction.
Do you think airflow is suitable for one-off/cron task management as well?
As for the ssl, I messed up my link. We are in the process of moving the site, so some of the redirects are not working perfect at the moment >.>
Internal validation error
11:58 AM PDT We are investigating increased error rates for Amazon S3 requests in the US-EAST-1 Region.
12:28 PM PDT We are investigating increased error rates for Git Push, Git Pull and API calls in the US-EAST-1 Region.
hopefully this raises awareness on how important planning for failure is before you make a design choice to introduce a dependency.
Console says: Failed to load resource: the server responded with a status of 503 (Slow Down)
EB interface is affected too
[1] I think they even have a graph somewhere whose axes are in UTC, but whose tooltips are in local browser time, but I can't recall for sure.
error: S3ServiceException:Please reduce your request rate.,Status 503,Error SlowDown,Rid