"7:42 AM PST We are investigating Internet connectivity issues to the US-WEST-2 Region."
https://status.aws.amazon.com/
Edit: They added US-WEST-1:
"7:52 AM PST We are investigating Internet connectivity issues to the US-WEST-1 Region."
Edit: Found root case, maybe?
"8:01 AM PST We have identified the root cause of the Internet connectivity to the US-WEST-1 Region and have taken steps to restore connectivity. We have seen some improvement to Internet connectivity in the last few minutes but continue to work towards full recovery."
"8:01 AM PST We have identified the root cause of the Internet connectivity to the US-WEST-2 Region and have taken steps to restore connectivity. We have seen some improvement to Internet connectivity in the last few minutes but continue to work towards full recovery."
us-west-1:
7:52 AM PST We are investigating Internet connectivity issues to the US-WEST-1 Region.
8:01 AM PST We have identified the root cause of the Internet connectivity to the US-WEST-1 Region and have taken steps to restore connectivity. We have seen some improvement to Internet connectivity in the last few minutes but continue to work towards full recovery.
8:10 AM PST We have resolved the issue affecting Internet connectivity to the US-WEST-1 Region. Connectivity within the region was not affected by this event. The issue has been resolved and the service is operating normally.
us-west-2:
7:43 AM PST We are investigating Internet connectivity issues to the US-WEST-2 Region.
8:01 AM PST We have identified the root cause of the Internet connectivity to the US-WEST-2 Region and have taken steps to restore connectivity. We have seen some improvement to Internet connectivity in the last few minutes but continue to work towards full recovery.
8:14 AM PST We have resolved the issue affecting Internet connectivity to the US-WEST-2 Region. Connectivity within the region was not affected by this event. The issue has been resolved and the service is operating normally.
I assume there are different teams responsible for each, is the west-2 team just more on top of things?
2. They don't really have much "legacy" stuff to deal with since they likely turn over racks quickly across their whole fleet and software deployments should be standardized, so any US-east-1 flakiness has to do with the fact that its where amazon houses their control planes often.
I agree in principle, but clearly something is hobbling them because of (probably) legacy stuff
So yeah, having done that I now understand that it was probably a joke but it really puts into perspective just how ridiculous things can get with a few 9's.
[1] https://earthsky.org/clusters-nebulae-galaxies/triangulum-ga...
[2] https://www.verywellhealth.com/why-do-we-blink-our-eyes-3879...
[3] https://en.m.wikipedia.org/wiki/Flash_(photography) (wikipedia's won't let me deep link on my phone, it's in electronic flash section under types)
It’s a clever way of getting reasonably accurate data very quickly and easily, though it does have it’s flaws — the data is pretty noisy and users often attribute outages to the wrong service (e.g. blaming their ISP or Microsoft or something when YouTube is down, or vice versa).
I'm surprised with FAANG hosted their stuff in their competitors cloud services without providing a fallback cloud service if the primary service is down. Sure it cost money but it would be effective this way than putting all eggs in one basket.
Like what, the networking is broken but if you could send packets, the services would still work so they are green?
Is this a thing?
Eat your own dog food shows confidence, but monitoring it is a different dimension, you need use anything but your own dog food there.
They could have a related outage, or even a coincidentally timed one
Maybe add a CDN. This shit isn't rocket science and being able to accurately monitor your own systems from off infrastructure is the one time you should really be separate.
Why not? The big tech companies use each other all the time.
For example, set up a new firewall on macOS and you can see how many times Apple pulls data from Amazon or Azure or other competitors' APIs and services.
The postulation was that Apple and Amazon weren't competitors. Not that they're not competitors in a specific niche.
Apple uses their competitor's services because they can't build their own cloud and host their own shit. The big boys don't use competitors for services they are capable of building themselves.
It seems to have all the look and feel of AWS, and somehow has more up to date info than the official AWS status page?
Pretty funny actually.
1. Remove all the thousand green services that no one cares about when looking at AWS status
2. Upgrade all yellows to reds because Amazon refuses to list anything as "down" no matter how bad the outage is.
3. Insert a snarky legend
Will large players flee because of excessive instability? Or will smaller players go from single-AZ to more expensive multi-AZ?
My guess is that no-one will leave and lots of single-AZ tenants who should be multi-AZ will use this as the impetus to do it.
Honestly, having events like this is probably good for the overall resilience of distributed systems. It's like an immune system, you don't usually fail in the same way repeatedly.
Only during this beta period, AWS will start charging for this feature soon enough.
Side note: terraform is pretty good for causing various kinds of chaos, deliberately or otherwise.
There is no possibility that outages are good for AWS. Nor is there more money to be made from "publicity" of the outages.
>Or will smaller players go from single-AZ to more expensive multi-AZ?
I... don't think you know what S3 is. Or maybe what FTP is.
(Also S3, EC2, RDS, etc. were named long before GCP had competing services)
All of these make sense.
If you're gonna complain about names, at least pick the really sucky ones, like Athena, Snowball, etc.
Do you know how many non-technical CEOs/boards/bosses have told their tech people that they need to go multi-region/cloud because that's what the one-paragraph blog and/or tweet told them to do in response to last weeks event?
Yes! When you have a service interruption pay 2x more! With a region down I am sure other regions wont have any interruptions either! /s
1: https://auth0.com/blog/auth0-architecture-running-in-multipl... 2: https://twitter.com/auth0/status/1471159935597793290
In the next 5 calendar years the bottom line will still grow.
However, the brand damage means they permananently lose market share. Which impacts their growth ceiling.
Gov Cloud Status Page: https://status.aws.amazon.com/govcloud
Cloudflare having some significant issues as well on certain domains.
The nice part about something like that is that properly wrapped you can change your durable storage as needed, and can easily even selectively pick "cheaper but less trusted" options for less critical data. It also allows you to leverage AWS features to ride closer to the wire. E.g. to take another example than storage, I've used this to cut the cost of managed hosting by being to spill over onto EC2 instances in the past, allowing you to run at much higher utilisation rate than what you can safely on managed / colo / on-prem servers alone - as a result, ironically the ability to spill over onto EC2 makes EC2 far less competitive in terms of cost to actually run stuff on most of the time.
Seemed to work for Dropbox.
I suspect DownDetector itself suffered some outages during this period, which it shows as outages of every service it monitors.
AWS seems to be working for me, but I’ve worked with clients in the US and spectrum internet tended to drop connections to us sporadically, which looks like an outage to our clients but is something we obviously can’t control.
I guess it could be an ISP thing but I guess we're all assuming 80/20.
We've taken a massive turn away from a "decentralized" internet.
just like Cavendish bananas are grown in multiple places...
(This is two similarly spec'd boxes on us-east-2 and us-west-2). Looking at GeoIP of connecting clients, the only pattern I can see is the region itself.
The response is that this actually works well enough, so the investment required has not pushed anyone to do it (with that meaning building the core infrastructure to make that easy).
What?
Every time Github went down multiple people post on HN saying "every since they were bought by Microsoft, ...". As annoying as those Rust evangelists on every single memory corruption bug.
First of all, how dare you!
Second, shoulda used rust ¯\_(ツ)_/¯
Plz don't disparage Rust evangelism!
Rust is awesome. yes it is complex, frequently annoying, easy to learn difficult to master. I'm speaking from a 30 year dev career.
a few months ago I intended to do a quick investigation into RUST to validate my "i really don't need to learn this" specifically for an embedded project. Within a few hours I found I had become a zealot. Rust has too many "omg, i should tell everybody about this" behaviors that I can't even find my favorite aspect yet.
It's equivalent to a lost soul finding Christianity and accepting the lords blessing and forgiveness! The weight that is lifted of being forgiven to your sins resulting == no more guilt, it's all forgiven! immediately reduction of cognitive dissonance. in this example with rust, it's pointer tracking and memory management, but it's basically the same thing. Rust is for the pious developer.
Those people who are still using C++ for fresh starts are the same folks who love to do things the hard & wrong way, or at least those who don't know any better, infidels, unwashed heathen.
Join us. join rUSt.
Until those Rust evangelists managed to rewrite the world with Rust (and I promise you there still will be a lot of security bugs), we still have to fix our shit in a low-cost way and their evangelism does not help at all and is pure annoyance.
No, I understand it. I started out in vuln research and have been in defense for a decade. It's probably fair to even say I'm an expert on it.
I'm going to keep advocating for rust as one the highest ROIs for improving security.
Slack wasn't sending messages and Pagerduty was throwing 500's.
This cloud-for-everything-even-local-devices thing is both hilarious and sad.
I wonder if anyone had trouble doing their dishes or laundry today, because I'm sure someone thought dish washers and washing machines needed cloud.
also, building access systems should be hosted in the building they reside in for security reasons anyways.
Depending on the cloud is certainly a very stupid decision. keeping everything inside the building is better, but still not ideal.
Cloud-based badges make sense if you have locations with small staffs and no HR people or managers. Like if you're controlling access to a microwave tower on the top of a mountain.
But badges-in-the-cloud for an office building full of people who are being supervised by supposedly trusted managers, and all of whom has been vetted for security and by HR, is just being cheap.
Like the 1980's AT&T commercials used to say: "You get what you pay for."
I'm not convinced that's true, or at least certainly not an order of magnitude. Wouldn't a badge system hosted on-prem also need a user management system (database), a hosted management interface, have a dependency on the LAN, and need most of the same hardware? Such a system would also need to be running on a local server(s), which introduces points of failure around power continuity/surges, physical security, ongoing maintenance, etc.
In addition, you're forgetting the thousands of points of failure between the building and the cloud provider. Everything from routers being DDOSed by script kiddies to ransomware gangs attacking infrastructure to Phil McCracken slicing a fiber line with his new post hole digger.
Adding complexity and moving parts never reduces points of failure. It can reduce daily operating worries as long as everything works, but it can't reduce points of failure. It also means than someday when it breaks, the root causes will be more opaque.
Many logical people have decided to abstract away their soul-crushing anxieties and legal gray area during outages to incredibly stable and well-staffed cloud infrastructure providers.
If you and your team are better at taking care of hardware than an entire building full of highly paid engineering specialists, then that's cool for you, but also, no you're not.
That's not to say you're not capable of running on-prem hardware that is stable.
I'm just saying that the high-handed swiping away of everyone else who's made an incredibly safe and logical decision to host their stuff in the cloud makes me question your general vibe.
If you plan to replicate all of AWS I'd agree with you. But if all you need is a handful of servers, you could end up with better uptime doing it in-house just because you don't have all the moving parts that make AWS tick, reducing the chance for something to go wrong.
My bare-metal servers stayed up during both of the recent outages, not because I'm some kind of genius that's better than the AWS engineers but just because it's a dead simple stack that has zero moving parts and my project doesn't require anything more complex.
The trade offs aren't quite that simple. Those specialists are necessary because they're building and maintaining infrastructure that's extremely complex since it has a crazy scale and has to be all things to all people. When you're running in-house, your infrastructure is simpler because it's custom tailored to your specific requirements and scale.
There are tradeoffs that make cloud vs local make sense in different contexts and there's no one right answer.
This appears to be cross-provider.
Edit: We have IPv6 back.
EDIT: Cognito auth seems down for us too
EDIT2: our ALBs are timing out as well
EDIT3: us-west-1 looks like working now!
The thing that really gets me is the reports from the last major outage a few days ago about how pervasive lying inside the company is. This really doesn't work well for engineering and we're possibly seeing the results of that. We should certainly expect to see that becoming visible the more time goes on without a major cultural shift. Which given that the guy who ran AWS now runs all of Amazon.com....
Is this enough of a push for organizations to actually move over their infrastructure to other providers?
The other cloud providers have had their own outages.
Organizations can more easily swallow an AWS failure when they aren't the only ones hit. They move elsewhere, those outages look more unique
Folks may think multi cloud is a good idea... But you're just as likely to suffer from the extra points of failure as you are to benefit
Discourse is reporting trouble, too. https://twitter.com/DiscourseStatus/status/14711403698992906...
> AWS Internet Connectivity (Oregon): 7:42 AM PST We are investigating Internet connectivity issues to the US-WEST-2 Region.
Source: https://status.aws.amazon.com
Their CDN, CloudFront, always works reliable for me. Couldn't they put the status page on CloudFront?
Edit: Hitting refresh a bunch finally got it open.
[1] https://aws.amazon.com/blogs/networking-and-content-delivery...
My sites run on Cloudflare and Vercel, and I can't even log in to those right now.
I'm curious — what does Hacker News run on? It seems impervious to any kind of downtime...
On a dirty, disgusting dedicated server.
I'm adding "reliable" into that mix. Too bad they're too expensive and hard to setup for side projects, but HN is probably one of the most stable site I frequently visit, and I don't even think about it.
Hard to setup is relative. It all depends on what you're doing and how much reliability you need. For a side project or a dev server you can just start with Debian, stick to packaged software (most language runtimes and services such as Postgres or Redis are available) as much as possible and call it a day. You can even enable auto-updates on such a stable distro.
The knowledge you'll gain by dealing with bare-metal is also going to be useful in the cloud even in container environments.
I mean they're not particularly. Unless the use is extremely minimal, it'll always be cheaper to buy a small server even for a side project.
I use cloud VMs for projects that can live on $5/mo VMs because at that usage rate I'll never break even to buy a machine.
But as soon as your AWS bill is even like $50/mo, worth to start looking at alternatives.
Pobody's nerfect.
It amazes me how many projects exist that don't even have multi-region capability, let alone no single point of failure
All to go from, idk, 99.9% uptime to 99.95% (throwing out these numbers)? The thing is when AWS goes down so much of the internet goes down that companies don't really get called out individually.
As a customer, it's unclear what the right approach is. Invest more with your vendor who caused the problem in the first place, or trust that they'll improve uptime?
I have some services which can cope with a 98.5% downtime, as long as they are available the specific 1.5% of the time we need them to run, as such "the cloud" is useless for that service
Really only 3?
Let's consider RockCo and CloudCo. They both provide a B2B SAAS that is mostly used interactively during the working day, and mostly used via API calls for the rest of the working week. Demand is very much lower on weekends. Both RockCo and CloudCo were founded with a team of six people: a CEO who does sales, a CTO who can do lots of technology things, three general software developers, and one person who manages cloud services (for CloudCo) or wrangles systems and hosting (for RockCo).
In the first year, CloudCo spends less on computing than RockCo does, because CloudCo can buy spot instances of VMs in a few minutes and then stop paying for them when the job is done. RockCo needs a month to signficantly change capacity, but once they've bought it, it is relatively cheap to maintain.
In the second year, they are both growing. CloudCo buys more average capacity, but is still seeing lots of dynamic changes. RockCo keeps growing capacity.
In the third year, they're still growing. CloudCo is noticing that their bills are really high, but all of their infrastructure is oriented to dynamic allocation. They start finding places where it makes sense to keep more VMs around all the time, which cuts the costs a little. RockCo can't absorb a dynamic swing, but their bills are now significantly lower every month than CloudCo's bills, and the machines that they bought two years ago are still quite competitive. A four year replacement cycle is deemed reasonable, with capacity still growing. And bandwidth for RockCo is much cheaper than the same bandwidth for CloudCo.
Who's going to win?
Well, you can't tell. If they both got unexpectedly sudden growth surges, RockCo might not have been able to keep up. If they both got unexpected lulls, CloudCo might have been able to reduce spending temporarily. RockCo spent more up front but much less over the long term. CloudCo could have avoided hiring their cloud administrator for several months at the beginning. RockCo's systems and network engineer is not cheap. And so on, and so forth.
Mostly /s; I wish the aws engineers the best of luck through this.
e.g. our services running on AWS are fine right now, but new sessions dependent on Auth0 are not.
[07:42 AM PST] We are investigating Internet connectivity issues to the US-WEST-2 Region.
/edit: https://phd.aws.amazon.com/phd/home#/dashboard/open-issues
[0] https://downdetector.com/status/crunchyroll/
[1] https://downdetector.com/status/aws-amazon-web-services/
A lot of websites use a cache in front of databases (or template rendering engines, or many other systems). That cache might evict entries based on time - after 5 minutes, the entry is considered invalid.
But that means that if you have no traffic for 10 minutes, the cache completely empties. Then when traffic returns, it all skips the cache and actually triggers a real hit to the backend - which is now overwhelmed with traffic. The cache protects the backend in normal behavior, but now it's not doing its job, so the backend has many more requests than usual.
In the worst case, those requests are enqueued in a big serial sequence... but the ones at the back of the queue may time out. The client may do something like say "it's taken me 5 seconds and I still don't have a response - I'll abort and retry!" and now you have even _more_ traffic to deal with.
So cold caches and retries can conspire to keep a service down for a long time even after the root cause is fixed.
IIUC the parent comment, it's describing a policy that evicts entries even (a) and (b) are false. Is that common in the web-hosting / CDN world? Or is age considered a proxy for stale?
A lot of web systems work this way - DNS records for example use a "TTL" which means "time to live." If the TTL is 60, then you throw it out of the cache after 60 seconds even if you have room in the cache, and you have no reason to believe it's invalid. This lets independent entities (like a DNS authority) make a change and get it rolled out everywhere.
I think the reason this is common is that proving cache invalidity is so hard, especially with the typical "dumb" cache appliances that are widely used. They just do stuff like cache the response bytes for a particular URL; they might not even understand HTTP beyond interpreting the request's headers, and certainly don't really understand the response.
All sorts of issues still unresolved for years, including the ridiculously annoying "Finishes playing season English sub, autoplays first season of German dub, which then gets stuck". Still no profiles (nerfing their super-premium offering). Auto-resume points are unreliable, the Android app is hot garbage at dealing with network disruption...
I can only imagine their back-end is mostly Visual Basic running on a single AWS-powered VM.
Seems like a really bad idea.
* EC2 instances
* AWS Workspaces
* FSx for Windows
* AWS Directory Service
* S3 Buckets
"No, no, of course not"
"Should I check?"
"No, don't waste time checking, get back to your TPS reports"
Seems to be down in a major way. Lots of various AWS services are down. However, so many things depend on AWS that it could just be EC2 is down and it is causing a rippling affect.
One issue is that outbound requests from our servers us-west-2 timeout. Other than that, it seems that we are running ok so far.
Website outages in the past hour 86,967
Lowest 16,208
Average 16,208
Highest 16,209
HTTP Error 500 internal server error
Some users are clueless, but the clueless users average out over time and the spikes make it clear when there are actual issues.
edit: some streams back up, chat still buggy as of 09:55 local time
edit2: appears to be back ~10:00 local time
There is zero excuse for this shit. Be professional. Acknowledge reality. It is logically impossible to run your own status page. Trying to do so just wastes everyone else on the internet's time when you have an outage.
See discussions from the last outage about the VP signoff needed to admit, I mean announce, an outage.
And that, ladies and gentlemen, is how I passed my system architect interview!
This morning we saw some weird behavior in us-west-2, our traffic just _vanished_. I thought: there is no way this is us.
Went to https://status.aws.amazon.com/
Top of the board showed “Internet Connectivity Issues (Oregon)”
And that was that. The board worked exactly as it should - it immediately explained my missing traffic and kept me up-to-date with the status of the outage on their side.
Most of these people just happen to be employed at AWS or azure right now.
I could not even try to discover which berries are edible without killing myself.
However, I can teach advanced maths to a largish group of students without much trouble.
Cluster berries, from raspberries to pineapples, are never poisonous. Avoid berries that resemble blueberries or currants unless you're able to identify the plant: we grew up with blueberries and know the leaves, but we avoid anything currant-like because we'd have no idea if they're actually, say, chokeberries. Avoid anything that looks like baneberries.
Here in Maine, we forage for raspberries, blackberries, wild strawberries, and (mostly low bush) blueberries, but don't risk others.
I know nothing but that still seems too generalized that it doesn’t have an exception somewhere in the world.
Historically here in Maine, the core diet seems to have been seafood, freshwater fish, maize, Capreolinae, game birds, eggs, honey, roots, and greens. While only a tiny fraction of fish/seafood remain, deer are over-populated and make a fine sustainable food source, the limitation mostly being the contemporary appetite for venison.
1. https://www.smithsonianmag.com/smart-news/indigenous-peoples...
Yes, that is a lot of time to go hungry and testing just one item. And then you still don't know what actually gives you nutrition vs just not killing you (for example leaves that you can break down such as leaf lettuce, vs eating grass).
The prevalence of meat comes from a society of abundance.
Example source:
https://onlinelibrary.wiley.com/doi/10.1002/ajpa.24247
A contradictory source says meat was less prevalent but still at 40-50%:
https://asu.pure.elsevier.com/en/publications/the-diet-body-...
Both of these estimates are way higher than what we know people eat today, where meat and dairy are 18% of worldwide calorie consumption (27% in the US).
So I think the abundance we see today is actually due to the availability of non-animal dietary sources.
Source - used to teach these skills before it was cool to be a "survivalist" on TV and social media.
These items aren’t in my earthquake bag (I have enough energy bars to last until the National Guard shows up). Instead these are for a Carrington Event type of solar storm, civil war or some sort of other long-term disaster.
Eat a small amount and see if you get sick? Science in the wild...
i wonder if there's anything like that for adults
I doubt even a very skilled engineer would know how his own machine works all the way down. What I think happens mostly is the skilled dev can use his experience to know where to investigate and where to look for solutions.
The question is organisational. Might it be that certain orgs have gotten so convoluted that they cannot do this investigation on an org level? Essentially, letting the right people look in the right places, unhindered by politics, legitimate security concerns, and practicality?
You'd think there'd be a limit to scale at some point. A bit of redundancy makes sense. There's probably a lot of people with multicloud setups patting themselves on the back at the moment.
For instance, I manage a team that does "full stack" development, where full stack means I regularly interact with mechanical and manufacturing, operations, electrical engineers, battery and radio people, embedded developers, mobile, and most aspects of backend engineering. We had an issue where one of our chip suppliers changed their FW, didn't tell us, and we literally were taking apart units to get to the bottom of why units off the line weren't working properly. We go pretty deep. Still, at some point we throw our hands in the air and say "Hardware is hard, it's in the name."
So you are disagreeing with this statement and saying that, in fact, you are the person who knows how the whole thing works?
I knew tech work produced some large egos, but sheesh.
But even if I didn't mean that, there still is nobody who could do it.
I'm saying no such person, even one who built up a team that they taught, exists for the modern computer.
But I am also incredulous by the parent.
I wouldn’t be able to diagnose traces on a motherboard or a defective (but partially functioning) CPU.
I wouldn’t be able to diagnose irregular voltage conditions or drop offs.
The amount of stuff that I know I couldn’t diagnose is absurdly high, but the amount I don’t even know that I can’t diagnose is higher still.
And my job, like the parents, is to drill down and spend time in specifics.
You know how AWS virtualization works and how to diagnose a problem with the AWS networking stack? Yes, and yes assuming that everything AWS is responsible for is operating within spec. Obviously, I don't have access to their switches, and cannot see anything at layer 1 or 2.
Knowing how it works, and being able to build a new one, are also two very different problems. For example, there are plenty of Computer Science folks who learned how to design chips (layout the circuits, write the microcode, etc) - but you need a whole extra background in EE and Physics to be able to fab said chip...
The problem is that this is really only possible with a toy model of a computer and very simple programs. Modern chips with their branch prediction and caching and threading and advanced vectorized operations and so on are vastly more complex. The 6502[2] was perhaps the last chip that one person could fully grok. Maybe a chip designer at Intel or AMD could understand the whole circuit in detail but no one else has the time - it would literally be a full time job. The same thing is true for operating systems - even if you're Raymond Chen, you can know a lot about Windows, but you can't know everything.
We learn just enough about the other parts of the system to convince ourselves that we understand the principles. We build the basic mental model we need to interact with other systems but all we can really do focus on our own specialized areas and hope that everyone else is doing their job. This works well enough until something like Spectre[3] or Meltdown[4] crops up and that's when we realize that we've been building castles in the sand.
[1]: https://www.nand2tetris.org/
[2]: https://en.wikipedia.org/wiki/MOS_Technology_6502
[3]: https://en.wikipedia.org/wiki/Spectre_(security_vulnerabilit...
[4]: https://en.wikipedia.org/wiki/Meltdown_(security_vulnerabili...
We really depend on three things - knowledge (stored collectively and in various media eg books), materials (tools and manufactured precursor goods, available via active supply chains or existing stores), and most importantly having our basic needs met trivially so that all our time is not sucked up addressing them.
A scenario where someone has a 'wasteland' to pick over for their basic needs, knowledge and materials looks quite different to a return to primitive living where what nature provides is all their is to work with. 'if you want to bake an apple pie, first you must invent the universe' or however it goes...
Then of course there's the question of why someone would have any interest in obtaining computing power were either of those scenarios to occur. Much like the 'how do we warn future civilisations about our nuclear waste' problem perhaps it is acceptable to not bother, they'll figure it out again on their own eventually given enough time.
This stuff is fun to think about at 6am when hay fever is preventing my sleep :)
And there's probably an equal number troubleshooting why it didn't failover the way it should, while their upper management starts questioning what they're paying for.
I've met a few people who can rightfully lay claim, but yeah, an incredibly rare set of skills.
That said, there is a recent revival in building systems from the ground up. While you can't manufacturer your own transistors, it is quite possible to understand everything from simple logic gates to ALUs to older style CPUs and memory buses.
I eventually got a fully-functioning 32-bit cpu with instruction pipelining, two levels of cache, DMA input/output, an asynchronous bus, a custom assembly language with an assembler written in python, and got the Game of Life running on it.
It ran about 2kHz with 8kb of memory or so.
It’s an easy target to romanticize but realistically, any alternative is basically a way of saying: “let’s stop evolving.”
Wanting to evolve differently doesn't mean halting.
It's a call not to devolve.
From my perspective, we require abstractions in order to free our intellectual capacity up for the next layer of complexity.
I thought about my cave men ancestors who during such a storm if they needed water would have to go out and get it, getting themselves soaked.
If I wanted water, the tap in the kitchen would give it to me, in a nice controlled fashion. If I did feel like having water rain down upon me, my shower would do that, again in a controlled fashion, and I could select the water temperature.
If they wanted the cave to be warmer, they had to burn something and deal with the smoke. And they might have to work hard to obtain whatever it is they burn.
If I wanted my apartment warmer, I just had to turn the knob on the thermostat.
They were at the mercy of their environment. My environment is mine to command. I was feeling pretty superior to my cave man ancestors.
Then I realized that I don't know how to build the systems that I was relying on for my supposed superiority, or even how some of them work.
I'm really just a cave man that found a nicer cave.
Commit and push often.
We had solar panels and a generator we used only when absolutely necessary. We were never without power, but we lived with the constant anxiety of optimizing our energy consumption. Some stuff we could only do during the day and at night we only used devices with batteries.
For a couple of weeks we didn't have running water in the cabin because we were rebuilding our water deposit tower. We used buckets for everything.
That was almost a decade ago and I still feel grateful at having unlimited energy or running water on demand.
I also feel guilty at times when doing power hungry stuff like playing video games, knowing electricity production is by far the biggest driver of climate change.
I think everyone ought to do a week in an RV with no connections to utilities. Not to take away from your story, but a similar scenario comes up when we "dry camp" (no water or electrical connections): resources are not unlimited. We have solar panels, big-ass inverter and big-ass battery to go with it. But if we want lights at night, best not run that 1100W microwave for too long, because the panels won't keep up and the battery isn't that big. We have a built-in generator, but unlike most RV owners, we are loathe to use it. It's almost like a game, and if that generator fires up then we've lost.
You want to let the water run while you brush your teeth? Go right ahead, our water tank is plenty big...oh, wait, but the holding tanks aren't. Shut that tap off before there's dirty water coming up through the shower. Speaking of showers, use the outside shower, as the holding tanks won't hold enough for your 30 minute, piping-hot shower.
Point of it all is that it one quickly learns that it all has to come from somewhere, and it has to go somewhere after you've dirtied it. I'd like to think that it has made the both of us more conscious of our usage.
There's nothing like being at sea, 100+ miles from civilization, reliant on the limited capacity systems on your vessel. You manage your food, you manage your water consumption, fuel, electrical usage, you're closely attuned to the weather, the sea state, the charts. There are no other visible people or people-made objects out to the horizon in all directions. If something breaks, you'd better know how it works and be able to fix it, or go without. It feels very freeing, but also provides a "back to basics" accountability.
Standing under a hot water shower with unlimited water in a spacious home shower afterward feels luxurious.
What's probably closer to truth is that many humans were forced to join farming communities. Stronger individuals or tribes probably enslaved others, and then forced them to build and produce.
The patterns of inequity and the march toward hyper-specialization we still see today make sense in that context.
As a tangent, if anyone is interested in that "cavemanness" deep in our DNA, check out the idea of primitive camping. That was my first experience camping, and I expected an idealized tv-ad experience. The trip was not framed as "primitive camping" to me.
I was dealing with intense burnout, stress, ADHD symptoms, immune problems, trouble sleeping... And I was thrown into the desert in the summer with a tent and some beer. It fucking sucked sooo bad. It fucking sucked sooo bad that I forgot every stupid problem I had, because I spent the entire time in survival mode. Setting up camp. Hauling equipment up and down dunes. Staying hydrated in the 100f+ heat. Making food. Making sure my wife and friends were ok. Strategizing how to defend our camp from bugs and psychos.
I really have not had such an existentially-dense experience as that one. And no, I didn't take any mushrooms, as the rest of the group did. I wanted to be lookout. Maybe I come from a long line of hyperaware sentries.
AFAIK humanity is yet to produce a society where the majority of farm laborers are fully free to leave the land they work on (whether via having their papers confiscated, their wages held until the season ends, by having transport provided to a remote farm but the trip back withheld etc). We've seen improvements in the degree of freedom, particularly over the past century and especially the past 50 years, but it's still very low compared to urban dwellers.
1: "Against the Grain", James C Scott
> Then I realized that I don't know how to build the systems that I was relying on for my supposed superiority, or even how some of them work.
I'm sure if you just sat down with a pen and paper you could come up with a DIY solution.
I used to have this joke(?) with my friends: remember Mark Twain's "A Connecticut Yankee in King's Arthur Court"? The titular Yankee basically upends the (faux) medieval society he gets transported to, "inventing" all sorts of technological miracles.
Well, I'm a software developer but don't come from an engineering background (I mean actual engineering, not programming). I don't even understand how electricity or the telephone work (I mean, old fashioned telephones, let alone current mobile networks). If I was transported to 2 or 3 centuries to the past, I wouldn't be able to explain modern technology to other people, let alone actually build it.
I sort of understand how steam machines work, and I could "invent" the printing press. I guess. But anything related to circuitry, electricity, chemistry, engineering of any sort, I wouldn't be able to even begin explaining them to King Arthur.
My introduction to the knights of the round table would go something like this:
"We are questing for the Holy Grail, oh noble stranger from a far away land! How can you help?"
"Depends, which version of Python are you running?"
[I only liked the first four books, but that's enough to cover the original story arc]
Fun fact: printing rates increased from about 120 sheets/hour to over 1 million over the course of the 19th century. Those began with wooden screw presses that differed little from Gutenberg's to cast iron, rotary, steam and later electric powered, and web (continuous paper feed) presses, and from matrix plates (with individual type set in blocks) to offset Linotype (in which the entire print block was cast as a single sheet through multiple stages from the original matrix characters).
Thought just occurs: the falling characters of the iconic Matrix screen somewhat resemble the individual type elements flowing and falling through a Linotype machine. I don't know if that is a deliberate or incidental reference, but it's an interesting one.
https://topatoco.com/products/qw-cheatsheet
And the T-shirt's companion bandana and spin-off book:
https://www.popularmechanics.com/culture/a23286104/how-to-in...
Wing is no problem as long as one can calculate how to make it stiff enough and of a right shape.
If I were a bit more clever, or maybe if I was 50 years older and had played with this kind of stuff growing up, I'd probably try to make a spark-gap transmitter. That seems to be in a sweet spot of not requiring too many super clever bits, and having obvious applications.
Not trying to glorify off-the-grid living or anything, but I think it's interesting to think that in some (very specific) ways, the cavemen were actually superior to us.
>It seems that someone asked the great anthropologist, Margaret Mead, “What is the first sign you look for to tell of an ancient civilization?” The interviewer had in mind a tool or article of clothing. Ms. Mead surprised him by answering, “a healed femur (thigh bone)”. When someone breaks a femur, they can’t survive to hunt, fish or escape enemies unless they have help from someone else. Thus, a healed femur indicates that someone else helped that person, rather than abandoning them and saving only themselves.
You aren't really - most cavemen didn't even understand that fire is possible, and wouldn't be able to consistently operate a lighter if they found one (it'd probably be put on an altar and worshipped instead, as it should). You might not be able to build your entire cave, but your education alone is a _huge_ advantage!
Toilets are amazing and I feel privileged every time I use one. Girlfriend thinks I'm nuts.
https://www.youtube.com/watch?v=ZSRHeXYDLko
See also Foundation by Asimov.
https://www.bing.com/images/search?q=toasting+bread+over+cam...
Sometime in the next few months, I've to troubleshoot and fix the broken 2 yr. old refrigerator. Someone came and fixed it once, now it's out of warranty and fixing it would cost about 50% of its cost. Meanwhile I'm glad I didn't throw away the 10 yr old refrigerator and just moved it to the garage. We just have to keep going to the garage.
I also have to play the accountant for my consulting business pretty soon. This is a task I had outsourced for years and have now started doing myself.
As stuff gets more specialized, I've started noticing that I'm able to do moderately complicated things better than professionals paid at the 50th - 70th percentile. If I want to get a really good job done, my rule of thumb is to be ready to shell out money in the 90th percentile range and look for references.
In case of AWS, I guess the Greasemonkey scripts are getting too complicated ;)?
But somehow there are multiple organizations that "know" how to build another hot bath, and newer and bigger baths are continuously being built all across the Empire.
And occasionally one of them stops working and thousands of citizens are angry, because they feel, being honest citizens of the Empire, they are entitled to enjoy these hot baths. Sometimes their very livelihood depends on the baths running.
Then the bath is fixed, and all is well again.
Maybe we need more baskets to distribute our eggs amongst?
https://www.thousandeyes.com/blog/aws-outage-analysis-decemb...
https://azycqgvwjz.share.thousandeyes.com/view/tests/?roundI...