AWS can do so many things, reporting critical outage updates in UTC is not one of those things.
Life is too short to remember what each timezone name means and converting to it, UTC offsets are much easier on the mental calculator.
GCP's various products have gotten a lot better at this lately, but just a few months ago I could click around between various dashboards and explorers, some showing the time in UTC, some in your browser's tz, and some in your profile's tz (if I recall correctly). Some of them were showing the tz, and for some you had to guess. Sometimes you had multiple tzs on the same page. Sometimes the date picker for a control was in one tz and the widget it was controlling in another (leading to quite a lot of confusion).
The worst offence IMO was not showing the tz at all. Especially given the overall lack of consistency.
Showing both at the same time is peak design for me personally. UTC compares for relative sequencing, local time for "was that before or after I ate lunch".
(You probably still meant IANA the Internet org not IATA the aviation one.)
Come to find out that this was some sort of entrenched, company-wide standard that was deliberately imposed. I made a lot of noise about this and appealed to some rather highly-placed directors, because I felt like it was wildly inaccurate and deceiving people; if you schedule a meeting in EDT but you say it's in EST, and we have employees all around the world, who's going to know? You're inviting off-by-one errors. Especially with me who lives permanently in MST.
3 years on, I've been unable to change this fundamentally; while a few people acknowledge DST, 90% of the company still adheres to this crazy false standard.
I encourage everyone at my company to do the same. Easy way to eliminate errors while typing 1 less key stroke!
Also, your clock can get confused driving North from PHX to Zion National Park.
In summer you start in Mountain Standard Time, drive into the Navajo Nation which does observe Mountain Daylight Time, containing through the Hopi Reservation, which is Mountain Standard Time. Then you end up back in Navajo Nation with Mountain Daylight Time. You keep on driving towards Page which is in Mountain Standard Time. However, when you cross the state-border of AZ/UT you're back in Mountain Daylight Time.
My clock threw a segmentation fault.
I literally had cases when I was woken up in the middle of the night for production issue because some people are too sloppy about this kind of thing.
Thankfully, I learnt a long time ago to use ISO 8601 and UTC for dates and times. I still revert to PST/PDT if my audience is primarily left coast based.
Heh. After the first few instance of confusion, we switched to saying Bangalore time and Dublin time.
Or go from rabbits being OK to some 5 figure fine if you're caught with one :)
Many people also get the timezone names completely wrong. I've had multiple scheduling email exchanges where someone says X pm EST not realizing that at the time it's currently EDT and that EST ≠ EDT.
And yet, for some reason, the two-letter abbreviations (e.g., ET) that are technically correct year-round, never seem to have caught on in the wild.
I've given up on the abbreviations and just say "Eastern" now to avoid confusion.
Names don't carry any information intrinsically, they are only a reference to the actual information, and the offset information is pretty short, so why not just provide the information directly?
"X pm GMT-3" only requires the reader to know their own timezone offset, unlike "X pm Brasilia time" (which is inaccurately known as São Paulo time outside Brazil) or "X pm BRT", which requires the reader to both know what that timezone means, and their own (or, more likely, requires them to look the conversion up).
(And if the difference between GMT and UTC is significant, I hope it didn't take my comment to convince you about using offsets :> )
- the keynote will start when this post is 5 hours old
- the rocket launch is scheduled to when this comment is 30 hours old
date: illegal option -- I
usage: date [-jnRu] [-d dst] [-r seconds] [-t west] [-v[+|-]val[ymwdHMS]] ...
[-f fmt date | [[[mm]dd]HH]MM[[cc]yy][.ss]] [+format]time encoded as a float of trecenti-seconds since year -8435 of the Georgian Calendar
why?
'caus it hurt to even just think about implementing that anywhere
UTC is human readable even if it is not calculated correctly. yes, i'm saying that if you can read epoch seconds, you're not human. 1970-01-01 00:00:00 is always a give away that something is a foot
Never mind dealing with India, Australia, etc etc.
OK to use local time in your statement, just say what that time is.
ftfy
AWS is powerful and very popular, but for the console, "it functions" must be the only condition the UI has to satisfy. Should every page use a unique table and sorting widget and UI language? Yes, please!
I'm assuming this helps them move fast, not having to coordinate with anybody or wait for a UI designer to tell them how it should look. But it's striking when compared to GCP.
Thank you for reminding me about one of my biggest mildest annoyances from working at AWS.
There are also people who tell me GMT (because they think that term means "the time in London") when they meant BST (because in summer, London doesn't operate on GMT).
I have found this site very helpful for linking people to when they are confused or using them incorrectly:
The time zone comparison feature is nice as well:
https://time.is/compare/0800AM_14_June_2023_in_Cincinnati/Lo...
If all systems talk this we'd save tens of thousands of man hours. Just do the conversion for us mortals, or other necessities. Tech side of incidents is definitely "system", I'd argue more often than not consumers of AWS are also tech side with systems in UTCs so health dashboards should also be a UTC first system. Doubt this could get prioritized tho
If you login, you can specify what timezone for timestamps and for the text to be parsed into your timezone preference.
https://health.aws.amazon.com/health/status#settings
As I'm logged in, it persists across browser sessions.
Imagine you have two browser opening the same page, one showing UTC, another showing your local time.
There are no indicator showing which time zone is used. You have to mentally correlate "this browser windows has logged in..." with the time shown on screen.
I don't know what the architecture of IAM looks like, but somehow it's never suffered a global outage.
AWS is really, really good at regional isolation.
Authentication possibly, but the control plane has gone down preventing changes.
Not being able to update your existing resources is still an outage from a DevOps perspective.
It might be an API level outage vs an end-user level outage from your customer's perspective, but if the functionality is down, it's an outage.
Our data plane was fine (for example, ec2 instances and s3 buckets in other regions were fine).
Increased Error Rates and Latencies
Jun 13 12:08 PM PDT We are investigating increased error rates and latencies in the US-EAST-1 Region.
They list Lambda as the only affected serviceThis seems like an odd limitation. Do you know the technical reason?
And why would ordered and paid for shopping be so inaccessible? It's not going to be locked away in a vending machine type thing right, you just need to show up with ID/log-in and collect it. Or so you should.
It's ridiculous that everything stops, but I do understand.
The pickup is even more ridiculous. Like you say, it's already through the system. They would just need to mark it as delivered later.
Are you an Amazon Flex driver? Or do you even need to check-in to pickup your order as a customer?
Contents:
> Has anyone ever actually had customers accept an outage because AWS was down; or is this just cloud evangelicalism copium?
> [ ] Yeah, outages free pass
> [ ] No, they say to use AZ's
Using 3 AZs in us-east-1 won't save you.
I guess a demanding customer would have said 'you should have implemented disaster recovery so you could failover to us-east-2' but that's easier said than done. The more regional AWS services you adopt, the bigger the impact is. How does one recover from a regional outage if their pipeline is in that region?
Another alternative studied was to use a thirdparty ci/cd service, outside of our network. It was discarded bc you never know where that would actually run
Yep, I considered that switching to GitHub Actions would _theoretically_ eliminate the need for disaster recovery for CI/CD (since the handling of disasters is out of your hands) but in practice their SLA is far worse than just running CodePipeline in a single region.
then you get to eat popcorn when stuff explodes.
* single server event. $
* multi server event. $$
* single az event. $$$
* multi az event. $$$$
* global provider event. $$$$$
* cross provider event. $$$$$$
* alien invasion. $$$$$$$$$$$$$$- memo from Enterprise Sales Dept.
Short of alien invasion level are strategic military resistance levels to global/regional wars with differing levels of weapons and devastation.
AWS really does have an easier time than old school datacenter providers. I guess the complexity is higher but it's shocking that they can charge so much yet we hold them to a lower standard.
I worked for one for some time and whenever we had issues, some people would call and ask if we were going bankrupt. It gave me a feeling they also have way smaller customers that might not understand the underlying stack.
edit: turns out AWS is the one with geo distribution, not Azure
If your customers are tech, they're too busy running around with their hair on fire too.
Whether customers "accept" it or not just comes down to what's in your SLA, if you have one in the first place, and if they are on a contract tier that it applies to. [Many servies provide no SLA for hobby / low tiers, beta features, etc.]
Firebase Auth, for instance, offers no SLA at all [1].
I would be curious to see statistics across a range of SLAs for what % include a force majeure or similar clause which excludes responsibility for upstream outages. I would expect this to be more common with more technical products / more technical customers.
edit: for those that would downvote: HN _just_ yesterday: https://news.ycombinator.com/item?id=36295352 https://news.ycombinator.com/item?id=36295305
Do you really think other (smaller) orgs can do a better job at hosting a datacenter than Amazon / Google / Microsoft / Cloudflare? They have some of the brightest minds in the industry working there, and they can price things at a much better price than anything you can build yourself.
Yes, I get it. All the computer processing power in a handful of actor's hands is probably not the most fantastic thing. However with the price of some cloud vendors compared to the DIY approach, it's hard for organizations to ignore.
If you really want to combat this, make the cost of running your own data center less. Reduce risk. Reduce the amount of money it costs for hiring good people or MSP's. Reduce the cost of acquiring and installing hardware.
Organizations pay attention to dollars so if you want the trend to shift, come up with a less costly alternative to the current cloud offerings.
Amazon et al even contract with them.
But cloud isn’t selling rented rackspace; it’s selling APIs for billions of things. Much different.
There are middle grounds.
But let's be honest: 99% of companies have never done the napkin math, because nobody ever got fired for choosing IBM^W AWS.
We joked about this in my company: we had a variable-load thing that we used autoscaling in the cloud for, but it had a baseline load that purchasing a real machine might have made a lot of sense for. The napkin math probably checked out. We never suggested it more than jokingly, though, because even when we suggested it jokingly, we got shut down: "You don't understand the cost of that." No, actually, we jokingly did enough math that we do understand, better than the people criticizing us did. We never did it.
Whenever the "own it" argument comes up, eveybody is real quick to hop on the "but maintenance cost" train. But as I perceive it, those who believe in the cloud budget exactly $0 for maintenance of managed cloud resources. As someone who's only done cloud, that number is unadulterated bullshit: the number of hours I've had to spend chasing cloud vendors to do the job that we're paying for is just silently flying under the budget radar. In the minds of the finance books, I'm 100% SWE, but in reality, I'm 75% SWE, and 25% support ticket monkey.
At least with a real machine, it'd be interesting, and I'd have some agency to actually solve the problem. As a support ticket monkey, I'm utterly powerless. I'm tired of having to beg.
That's not to say I'd move everything off cloud; I actually think the vast majority of what we do is well-suited for cloud, mostly because upper management can't make up their mind about product direction enough to be able to say "yes, we can purchase this and we'll use it." But those nuggets of stability do happen from time to time.
> disaster planning
"Disaster planning" is something every org wants, because they're trying to tick the box with the regulator. But the requirements that get passed down border on absurd: "what if a meteor hit AWS and they were never able to recover from it?" … we're literally never going to plan for that, because the $ needed for that level of eng. work is not going to happen. A sane scenario would be "can we handle an AZ outage?" (or, let's start there, and maybe, maybe if we can get that down pat, then we can graduate to regional outages.)
> cyberinsurance
… you don't get out of this via being in the cloud, if you need it. (I wish we did, because ours pushes some utter inane requirements.) I can mismanage a machine in a DC just as easily as I can mismanage a VM in the cloud.
> Organizations pay attention to dollars
No they don't. This oft-repeated mantra is nonsense. Finance dept. get an invoice that has a total; even were they to have access to the finer billing information, they're not technical, and cannot understand it. I've yet to be at a company that's dedicated sufficient resources towards infra eng such that we could do the legwork necessary to present a sane organizational view of what cloud infra dollars go to what high-level objectives or teams. The resource tagging isn't there, and even if it were, some things cannot be tagged, and you still have to aggregate bills from a dozen different vendors, and then figure out what weights to apply to shared resources across OUs. I'm on employer #4? and have yet to see anyone scratch the surface of that.
Which is why you see articles about cloud $ waste all the time.
What happens far more often in my life is someone from management descending with "why are we spending $X on Y?", where $X is usually an order of magnitude wrong, or Y is … something we're not even doing anymore? And then you have to go round the mulberry bush of "how did you arrive at that figure?" "okay so here's what those numbers mean" "here you're adding $/mo and $/yr and you can't do that"
> Do you really think other (smaller) orgs can do a better job at hosting a datacenter than Amazon / Google / Microsoft / Cloudflare?
Than Microsoft? Absolutely yes. The others, probably not.
> Yes, I get it. All the computer processing power in a handful of actor's hands is probably not the most fantastic thing.
The long-term end state of not investing money into R&D is that it is centralized into those who do, and you become beholden to them. You get what you pay for, here. It's not good, and I think there's discussion to be had around that, but my real problem is the cognitive dissonance that follows. If you want to centralize on one of the cloud duopoly, then you also need to acknowledge that your own eng cannot be held responsible for the cloud's reliability: they have no control over it.
Could you imagine a 1,000 person org with a 20 person IT department rolling their own DDoS solution?
But yes you are also right, cyber insurance is still required, and even AWS touts an expectation of a “shared responsibility” model.
I’m still skeptical that cloud hosted offerings are a bad thing. For a long time there were only Ford, Chrysler and Chevy in America, then foreign imports became popular, then a few years ago Tesla became a contender.
I still think new entrants can come into the cloud space, particularly in Europe, but they need to do their due diligence and understand their competitors offerings very well.
everyone knows, nobody seems to care.
Another comment of mine in this thread asks the question if you can excuse downtime of your service due to AWS outages.
Consensus seems to be: yes
which is a pretty huge deal, well worth the insane cost increase of AWS by itself. No other hosting provider would grant you such an excuse.
I would weep for the centralised future of the internet, but its already here, so theres no point.
This is a good reminder to avoid cloud-centric products, but they are getting harder and harder to avoid.
I have always stayed away from that region because it seems significantly less reliable than other regions.
* Largest (DDoS'd most, most complex, scaling issues etc)
* Oldest (More time for weird idiosyncrasies to take hold)
* Where most testing happens
* Where new products are deployed first
IAM, Cloudfront ACM certs, etc
* The only place where the IAM dashboard can be accessed from. I need to access it NOW. I can't.
It is also a massively complex beast in itself spanning dozens of datacenters with massive amounts of fiber between them. Much more fragile than having everything in a single building and as you scale up the number of components you increase the rate of failure.
https://statusgator.com/blog/is-north-virginia-aws-region-th...
Deployments start off very conservative, maybe 1-2 small regions on the first day of deployments. As you gain confidence, the pipeline deploys to more regions/bigger regions.
A pipeline that deploys to 22 regions over one week might go from 2 small regions on monday, 4 small/medium regions on tuesday, 8 medium/large regions on wednesday, 8 regions on thursday.
us-east-1 is usually going to be deployed to on the wednesday/thursday in this example, but that isn't always the case because sometimes deployments are accelerated for feature launches (especially around re:invent), or retried because of a failure.
There are best practice guides within Amazon that very closely detail how you should deploy, although it is up to the teams to follow them, which they usually do an okay job of.
@dijit is right: https://news.ycombinator.com/item?id=36315736
Typically, a team would group their regions into batches and deploy their change to one batch at a time. Usually they follow a geometric progression, so the first batch has one region, the second batch has two regions, the third batch has four regions, and so on. This batching was performed for the sake of time; nobody wants to wait a month for a single change to finish rolling out.
One reason not to deploy to us-east-1 in the first batch is so you don't blow up your biggest region. The fewer customers you break, the better.
One reason not to deploy to us-east-1 in the last batch is that there are a lot of batches. If a problem is uncovered after deploying the last batch, then someone has to initiate rollbacks for every single region.
Some teams tried to compromise and put us-east-1 in one of the earlier batches.
https://ca-central-1.console.aws.amazon.com/console/home
This assumes you don't actually need anything from us-east-1, though :)
We don't use big cloud were I work, so maybe I'm missing something. Does East-1 offer something other don't?
Plus, as others have noted, there are critical AWS services in their control plan that only run in us-east-1 behind the scenes. So you're kind of out of luck.
Then you realize that they just switched you back to us-east-1 for some reason and a wave of familiar relief washes over you.
For example, do you want your Cloudfront CDN to have a custom (secure) domain?
Then you have to host your ACM cert in us-east-1.
But I still prefer EU region =).
For example, in Poland, your typical restaurant or shop needs to generate tax receipt (as well as properly calculate the tax), and uses either a separate receipt POS device, or POS with appropriate receipt printer (the devices are certified and for example do simultaneous two prints - one for client one for seller - or use digitally-signed storage for seller copy).
If the POS isn't designed properly to operate in case of network failure... welp, can't take cash either, at least not legally.
When the POS system goes down, restaurants take down credit card numbers, and then charge them later when the POS comes back up.
Even management console is down, and their suggested region specific workaround does not work, at least for us-east-1. I can see some processes via api but I don't have code prepared for monitoring every service from my local.
i wonder if it will work first try? the true test of devops culture.
<rant>
It's also quite annoying sometimes that some things _need_ to be in us-east-1, and if e.g. you are using Terraform and specify a different default region, AWS will happily let you create useless resources in regions that aren't us-east-1 that then mysteriously break stuff because they aren't in this one blessed region. AWS Certificate Manager (ACM) certificates are like this, I believe.
</rant>
Sigh.
i find it stupid that clients get to know about regions at all. They should only notice latency hits and batch job queuing latency if something bad happens underneath, but no services should go down.
pretty cool that stuff was stored in a backlog and eventually processed!
I wonder what % of the internet went down because of the us-east-1 today.
Loss of DNS causes inter-service api calls to fail, then IAM and all other services fail. Anything not built to handle those situations with backoff causes a 'stampeding herd' of failure/retry and exacerbate the outage
Review the AWS statements about outages here - https://aws.amazon.com/premiumsupport/technology/pes/
And what a surprise it's US-EAST-1 again...
* DNS
* Misconfigured switch
It's always one of those two.EDIT: thanks to those of you who have signed up in the past few minutes! Let us know if you have any feedback.
Though a lot of practical thermal-related causes of electronics failure seem to operate on timescales of years to decades, like electromigration https://en.m.wikipedia.org/wiki/Electromigration or even just cooling fan bearing failure. And I don’t think it would be a huge stretch to point to electromigration as a case of diffusion, a natural entropy increasing process, re-randomizing the arrangement of atoms within a transistor (and therefore making it fail eventually).