My $500M Mars rover mistake
chrislewicki.com
chrislewicki.com
As a software engineer, I have a couple stories like this from earlier in my career that still haunt me to this very day.
Here’s a short version of one of them: Like 10 years ago, I was doing consulting work for a client. We worked together for months to build a new version of their web service. On launch day, I was asked to do the deployment. The development and deployment process they had in place was awful and nothing like what we have today—just about every aspect of the process was manual. Anyway, everything was going well. I wrote a few scripts and SQL queries to automate the parts I could. They gave me the production credentials for when I’m ready to deploy. I decided to run what you could call my migration script one last time just to be sure I’m ready. The very moment after I hit the Enter key, I realized I had made a mistake: I had just updated the script with the production credentials just before I made the decision to do another test run. The errors started piling and their service was unresponsive. I was 100% sure I had just wiped their database and I was losing it internally. What saved me was that one of their guys had just a couple hours earlier completed a backup of their database in anticipation of the launch; in the end, they lost a tiny bit of data but most of it was recovered via the backup. Ever since then, “careful” is an extreme understatement when it comes to how I interact with database systems—and production systems in general. Never again.
Great story, thanks for sharing.
We rarely interact directly with production databases as we have an event sourced architecture. When we do, we run a shell script which tunnels through a bastion host to give us direct access to the database in our production environment, and exposes the standard environment variables to configure a Postgres client.
Our test suites drop and recreate our tables, or truncate them, as part of the test run.
One day, a lead developer ran “make test” after he’d been doing some exploratory work in the prod database as part of a bug fix. The test code respected the environment variables and connected to prod instead of docker. Immediately, our tests dropped and recreated the production tables for that database a few dozen times.
if strings.Contains(dbname, "prod") {
panic("Refusing to wipe production database!")
}
Truncate(db)One of the reasons I put interactions between databases behind a cli.
*ideally* devs should not have prod access or their credentials should only have limited access without permissions for destructive actions like drop/truncate etc.
But in reality, there's always that one helpful dba/dev who shares admin credentials for a quick prod fix with someone and then those credentials end up in a wiki somewhere as part of an SOP.
If you need access for a quick prod fix, your key gets added to the machine with that explanation and a week (or lees) lifetime.
There is no code that will protect your db/data. Only replication to a read-only storage will help in such situations.
For instance, test code shouldn't have access to production DB passwords. Maybe that means a slightly less convenient login for the dev to get to production, but it's worth it.
For Django projects, add the below to manage.py:
env_name = os.environ.get("ENVIRONMENT", "ENVIRONMENT_NOT_SET")
if env_name in TEST_PROTECTED_ENVIRONMENTS and "test" in sys.argv:
raise Exception(f"You cannot run tests with ENVIRONMENT={env_name}")It’s useful to distribute the test anyway, especially for non-transactional tests.
If the database initialisation is costly that’s useful even if tests run on empty, as copying a database from a template is much faster than creating one DDL by DDL, for postgres at least.
Of course this bypassed the rewrite process, and there was inadequate separation between QA and prod, so now they were connected to the live DB; and then they ran `rake test`...(cue millions of voices suddenly crying out in terror and then being suddenly silenced). The DB was big enough that this process actually took 30 minutes or so and some data was saved by pulling the plug about half-way through.
And _of course_ for maximum blast radius this was one of the apps that was still talking to the old 'monolith' db instead of a split-out microservice, and _of course_ this happened when we'd been complaining to ops that their backups hadn't run for over a week and _of course_ the binlogs we could use to replay the db on top of a backup only went back a week.
I think it was 4 days before the company came back online; we were big enough that this made the news. It was a _herculean_ effort to recover this; some data was restored by going through audit logs, some by restoring wiped blocks on HDs, and so on.
We never hear about first time launch deploys that wipe ALL data because whoever is so unlucky probably never got to browse hacker news
I can think of a larger blast radius when deleting files on a shared mount point for example but it's not representative to the regular use of sudo.
\set AUTOCOMMIT off
In your .psqlrc and then you can never forget the begin transaction; every statement is already in a transaction, its just the default behaviour to automatically commit the statements for some ungodly reason.Years ago I hired an experienced Oracle developer and put him to work right away on a SQL Server project. Oracle doesn't autocommit by default, and SQL Server does. You don't want to learn this when you type "rollback;". I took responsibility and we had all the data in an audit table and recovered quickly. I wonder if there are still people who call him "Rollback" though.
That's good from the DBA perspective, but relying on that default as a user is risky in itself, when you deal with multiple hosts and not all are set up this way.
…only to reappear a second later. It was just the view refreshing! Talk about awful UI!
I had access to the production database, something I absolutely should not have had but we were a tiny ~15 person company with way more clients than we reasonably should have. Corners were cut.
I write a quick little UPDATE query to update some marketing text on a product and when the query takes more than an instant I knew I had screwed up. Reading my query, I quickly realize I had ran the UPDATE entirely unbounded and changed the description of thousands and thousands of products.
Our database admin with access to the database backups had gone home hours earlier as he worked in a different timezone. It took me many phone calls and well over an hour to get ahold of him and get the descriptions restored.
The quick change on my way out the door ended up taking me multiple hours to resolve. My friend in marketing apologized profusely but it was my mistake, not theirs.
As far as I remember we never heard anything from the client about it, I put that entirely down to it being 5pm on Friday of a holiday weekend.
Oh, and I absolutely refuse to do anything but the most critical stuff against prod on Fridays.
For me personally the much bigger issue would be harming the client, their business or our relationship. Doing a few hours of overtime to fix my mistakes would probably only feel as well deserved punishment...
Luckily the customer sites each had a local db that synced to the central db (so the product could run with patchy connectivity), but the guy spent 3 or 4 days working looooong days rebuilding the master db from a combination of old backups and the client-site data.
I am very worried about doing the wrong thing in the wrong terminal, so for some machines I colour-code my ssh windows, red for prod, yellow for staging and green for dev. e.g. in my ~/.bashrc I have: echo -ne '\e]11;#907800\a' #yellow background
A DBA colleague sitting nearby laughed and had things restored back within a few minutes....
Having worked in enough of these though, I am aware that even they (the "seniors") are seldom entirely responsible for all the issues. It's mostly business constraints that forces cutting of corners and that ends up jeopardizing the business in the long run.
(One of my standard end-of-interview questions is "how easy is it for me to trash the production database?" Having done this previously[1] and had a few near misses, it's not something I want to do again.)
[1] In my defence, I was young and didn't know that /tmp on Solaris was special. Not until someone rebooted the box, anyway.
I’ve had a search but can’t work out why it’s special.
"why didn't they have a hot-spare" They do! Flight spares are complete, flight-rated copies of spacecraft built for exactly this contingency: https://en.wikipedia.org/wiki/Flight_spare After launch the flight spares are used for terrain testing and troubleshooting. (The "mars yard" has flight spares for Curiosity and Perseverance https://www-robotics.jpl.nasa.gov/how-we-do-it/facilities/ma... which were used to test some wheels to destruction after Curiosity started showing some wear https://www.planetary.org/articles/08190630-curiosity-wheel-... )
The blog post lays it on a bit thick with the $500 million number and the "launch only two weeks away" given that the article itself is illustrated with a photo of the Sojourner flight spare. Spirit had the SSTB1 test rover. If he had actually blown out the entire electrical system, they could have launched it instead. Swapping out the entire vehicle right before launch would have been an awful job, but it's not flat out impossible.
I liked that other people pointed out that risk could have been eliminated by using polarized connectors (I hope they started doing this after the incident), but also made me wonder about "back-EMF" caused by solar flares. In other words, maybe all thick wires and ground/power planes should be hardened against current surges simply due to a solar event hitting mars (which may incidentally cover the case of back-powering the driver circuits).
I have been burned by this in some version of Ubuntu and have assumed it was normal behaviour ever since.
This brief moment in time has a name: an ohnosecond.
But you should definitely have bought that man a beer :)
To do that, on the "connect to server" dialog, click "options". On the tab "connection properties" in the "connection" option group, check "use custom color". And I pick the reddest red there is. The bottom of the results window will have that color.
edit: my horrible foul-up was restoring a database to production. The "there is trouble" pagers were all Iridium pagers since they loved climbing mountains (where there was no cell service back then). But then that place didn't use source control, so it was disasters all the way down.
Whether this was a process problem or a human one we don't really get to judge since we do expect more from a FTE.
I'll just say putting myself into his shoes made me tear up as I read the dread and pangs of pain upon realizing what happened - then to have life again after the failure of the ray of hope. That weight, I've never had a project that so many people depended on.
All heroes in my book.
1. The plug was allowed to be connected backwards. Either this should be impossible or this hazard should be identified and more than one human should verify orientation
2. In use tools like multimeters should never be disconnected. At worst you get problems like this at best you annoy whoever was using it
Blaming individuals only gets them fired and weakens the entire organization. You just fired the one person who learned an expensive lesson.
The only time when an individual should be blamed is when they intended harm, at which case the law could kick in
Like say they have one that is setup to test the motor driver circuitry and another one that is setup to test the motor?
Or say the breakout box intentionally has both sides of the connection on it, so that you can get in-between the driver and motor?
So either the ends are literally the same (e.g. Anderson Powerpole), or there is some kind of weird symmetry or inadequate keying. Or maybe the two cables don’t connect directly and instead go through some of kind of interface? The latter is fairly common in networking, e.g. “feed-through” patch panels and keystone jacks and quite a few kinds of fiber optic connectors.
All of these seem like utterly terrible ideas in an application where you would take the thing apart after final assembly and where the person doing the disassembly or reassembly could possibly access the wrong side of the panel.
In that case keying or whatever isn't going to prevent you from connecting to the wrong side, because both sides are present.
I'm suspecting the breakout wasn't literally sitting between the driver and the motor, but rather all internal connections are broken out to the box for testing; and likely the author's mistake was to not mess with the spacecraft to temporarily disconnect the driver.
But I'm not sure if I'd "just" made the right call and done so nonchalantly on a Mars rover to launch in few weeks.
"The incident was featured in season 23, episode 5 of the Canadian documentary series Mayday . . ." [1]
Season 23 - I'm glad I don't fly!
How … how often does that go wrong?!
Control cables also can and do break, but that too is fairly rare.
What is not rare is control mechanisms jamming. Here is an example:
https://www.ntsb.gov/news/press-releases/Pages/NR20230928A.a...
If you are putting people in a situation with absolutely no safeguards, you can’t have them go into it fatigued.
I’m guessing the people working on that team also weren’t getting great sleep by the discussion of high stress and long hours. Recipe for disaster.
The simplest task you can imagine takes incredible proportions (for good reasons).
Disconnect and reconnect that plug? Please inform persons X and Y, person Z must be present, only person W can touch that plug, and do perform a functional test according to the procedure in this document before and after and file these reports etc ...
Cleaning a part? Oh glob. Get ready for 3 months of adventure talking to planetary protection experts and book the cleanest room in the continent.
"It's a waste of time" is very often a fallacy, especially when the risk cannot be easily undone.
I (mostly mentally) complete the phrase "It's a waste of time" with "what's the worst that could happen?", and when I'm actually saying the phrase out loud, stare at whoever said that for 5 full seconds.
Why? Bad error handling in the software (primarily). What is the worst that could happen? An instrument saturate, a variable gets stuck at a value, but keeps being integrated, the spacecraft computes a negative altitude and thinks it'a below ground level (negative altitude) but is in fact in full descent and at 3+ km from the surface. Oopsie !
[1] https://exploration.esa.int/web/mars/-/59176-exomars-2016-sc...
Agree on the blame point, but not on firing point. As a manager, sometimes you need to fire people, that's a necessary part your job. And no, changing the hiring process cannot prevent that.
PS: Also, more rules and better processes are not necessarily a good thing. Sometimes there are just too much red tape and bureaucracy that makes already super-slow NASA even slower. In those first-of-its-kind missions sometimes you need to risk and depend on people, not processes.
At my first real job as a web dev after school, I crashed the production website on my very first day. Tens of thousands of visitors were affected, and all our sales leads stopped.
Thankfully, we were able to bring it back up within a few minutes, but it was still a harrowing ordeal. The entire team (and the CEO in the next room) was watching. It ended up fine and we laughed about it after some minor hazing :)
But by the time I left that job a couple years later, we had turned that fragile, unstable website into something with automatic testing, multiple levels of backups and failover systems across multiple data centers, along with detailed training and on-boarding for new devs. (This was in the early days of AWS, and production websites weren't just a one click deploy yet.)
That one experience led to me learning proper version control, dev environments, redis, sharding and clustering, VMs, Postgres and MySQL replication, wiki, monit, DNS, load balancers, reverse proxies, etc. All because I was so scared of ever crashing the website again.
That small company took a chance on me, a high school dropout with some WordPress experience, and paid me $15/hour to run their production website, lol. But they didn't fire me after I screwed up, and gave me the freedom and trust to learn on the job and improve their systems. I'm forever grateful to them!
The first thing mentioned in the post mortem call was “No one is going to blame the guy who did those trades. It was an honest mistake. What we are going by to do is discuss why a developer can hit the production trading API without any authentication at all”.
I was chatting with him when he noticed the stock the strategy was trading (KLAC) was gradually declining linearly. He looked at the L2 quotes and saw that someone using his brokerage was repeatedly putting out small orders, and then he realized they were his orders.
The fund got a margin call and had to shift some funds between accounts to make margin, and they had to contact regulators and inform them of the bug, and they had to manually trade their way out of the massive short position they traded. However, they ended up making $60,000 that day off of his mistake.
I'm a frequent flyer and I got a feeling that most airline ticket booking pages are broken in some way more than half the time. Maybe not often broken to the point that they're blank, but definitely broken to the point that booking a ticket isn't possible (I prefer blank, so that I don't waste like 30 minutes on not being able to book a ticket).
Also most of the internet seems often broken. Oh hello Nike webshop errors upon payment (on Black Friday) for which helpdesk's solution is: just use the App.
Everything can be a guilt trip if you try hard enough.
Then I met a guy, now a good friend, that made me do my first "pull the plug migration" on his most important website. He lived on this.
I looked at the site going down, horified. He mocked me, then proceeded with the udate. It didn't work. The site stayed offline for hours.
Then it worked again. And nobody cared. It had zero consequences on traffic.
User were pissed off for a few hours, and life goes on.
People who take on high risk projects are underappreciated. But many managers prefer employees who can reliably deliver zero value, than those with positive expected value but non-zero variance.
That development would be an obvious investment that pays for itself. I’m in banking, and terrified of making even a slightly complex deployment without validating it in production first. (Complex here referring to that it might be dependent not just on code changes, but also environment).
I mean look at Twitter, which was famously down all the time back when it first launched due to it popularity and architecture. Did it mean people just stopped using Twitter? Some might, the vast majority and then some didn't.
Downtime isn't catastrophic or company-ending for online services. It may be for things in space or high-frequency trading software bankrupting the company, but that's why they have stricter checks and balances - in theory, in practice they're worse than most people's shitty CRUD webservices that were built with best practices learned from the space/HFT industries.
https://news.ycombinator.com/item?id=13419313
"> A young executive had made some bad decisions that cost the company several million dollars. He was summoned to Watson’s office, fully expecting to be dismissed. As he entered the office, the young executive said, “I suppose after that set of mistakes you will want to fire me.” Watson was said to have replied,
> “Not at all, young man, we have just spent a couple of million dollars educating you.” [1]"
The boss's response makes a lot more sense than the usual fluff, though: "If I fire you now, the next guy to make a mistake won't admit it and we won't find out about it until it's too late."
> That website sold millions of euros worth of tickets every day.
The claim wasn't that a single airline sold a million dollars per day, but that a third party on seller sold a million euros worth of tickets a day.
Is that plausible?
Consider The City in the Sky:
Every day 100,000 flights criss-cross the globe with more than 1 million people in the air at any one time. Dallas Campbell and Dr Hannah Fry explore the world of aviation.
https://www.imdb.com/title/tt5820022/At any instance there are one million people aloft.
At any instance there's at least 50 million dollars worth of ticket sales in play - how much during a 24 hour day would you estimate?
Is it possible for a single third party seller to capture a million euro per day?
However, I don't agree that this is the "real" lesson.
Given the costs at play and the risk presented, the lesson is that if you have components that are tested with a big surge of power, give them custom test connectors that are incompatible with components that are liable to go up in smoke. That's the lesson. This isn't a little breadboard project they're dealing with, it's a vast project built by countless people in a government agency that has a reputation for formal procedures that are the source of great time, expense, and in some cases ridicule.
The "trust the 28 year old with the $500m robot that can go boom if they slip up" logic seems very peculiar.
On the other hand, it's hard to make these kinds of judgment calls when you're talking about a one-off piece of equipment that's only going to go through this particular testing cycle a single time.
In computing, there are a lot of similar "one-off" operations -- something you to do to the prod database or router config a single time as part of an upgrade or migration.
Sometimes building a safeguard is more effort than just paying attention in the first place. And while we don't always perfectly pay attention, we also don't always perfectly build safeguards, and wind up making similar mistakes because we're trusting the faulty safeguard.
In circumstances like the one in the story, the best approach might almost be the hardware equivalent of pair programming -- the author should have had a partner solely responsible for verifying everything he did was correct. (Not just an assistant like Mary who's helping, where they're splitting responsibilities -- no, somebody whose sole job is to follow along and verify.)
This may be acceptable, but it comes down to managing risks. If failure means the company dies then taking a 1 in 10,000 risk to save 3 hours of work probably isn’t worth it. If failure means an extra 100 ours of work and 10k in lost revenue then sure take that 1 in 10,000 risk it’s a reasonable trade off.
Especially during testing you're often dealing with custom cables connectors and circuits that are different from the "normal configuration".
I would say that the lesson is to do as many critical operations under the 4-eye principle: someone is doing the thing, someone else is checking each step before continuing. Very effective for catching "stupid mistakes" like the one in the article. But again, it is not always possible to have two people looking at one test, especially with timeline pressure etc. So mistakes like these do happen in the real world. You have to make the whole system robust.
I agree with you, but on Earth this is easy. For spacecraft I imagine you can't just use any connector from Digikey
> especially with timeline pressure etc.
If timeline pressure, lost sleep, or rushing jobs not meant to be rushed causes a catastrophic technical error to be made, it is 100% the fault of the person who imposed the timeline, whether that be some middle manager, vice president, board, investor, or whoever. Emphatically NOT the engineer who did the work, if they do good work when not under time pressure.
HOLD PEOPLE LIABLE for rushing engineers and technicians to do jobs that require patience and time to do right.
However, you can't always eliminate timeline pressure. Even if the project is planned and executed perfectly, there will almost always be unknown unknowns encountered along the way that can push your timeline back. As is the case with sending things to Mars there is a window every two years. That's a very real, non-fictitious deadline that can't be worked around.
This is very simple to deal with.
(a) If it's unmanned, rush and launch on-time but don't fault the engineer for mistakes made by rushing. If it doesn't work everyone accept that as a consequence of rushing.
(b) If it's manned, wait until the next launch window and prioritize safety. Period.
Not just that, but to create a situation whereby said person is working unofficial double shifts to get it done, so probably aren't going to be bringing their best selves into the office. If it were my $500 million I wouldn't even care about the name of this guy but would want to have some very robust discussions with the head of their department. Also, "some mistakes feel worse than death" - I get it, but c'mon, it's not like someone actually did die, which is a sadly unfortunate reality of other much less spectacular and blog-worthy mistakes.
I don’t know the details in this case but it could be like this: socket-type connectors are required on external connectors on the spacecraft (to prevent shorts when handling), with a harness in between which will never be removed. The harness would be symmetrical with pin-type connectors.
At some point it is decided a breakout box is required for testing and now you have created an opportunity to plug the breakout box in backwards.
Or the breakout box has a 100 pin connector on one side and needs to connect to 25 pieces of test equipment on the other side. You probably don’t have 25 different connectors to chose from, nor can you possibly demand custom requirements for every piece of test equipment.
Spacecraft are moving more towards local microcontrollers with local diagnostics so this kind of test equipment for every possible analogue signal is decreasing. In the case of motors, they would more likely be brushless now and you would rely on telemetry from motor drivers during both testing and flight instead of having this type of breakout box.
Connectors in aerospace are also following other industries and becoming more configurable at order time, including adding keys so you can have 10x “the same” connector but keyed so they only plug in one place. But it’s still not practical to demand all test equipment is configured like this.
(Same for an 82 year old or any other number..)
My first job out of university, I was working for a content marketing startup who's tech stack involved PHP and PerconaDB (MySQL). I was relatively inexperienced with PHP but had the false confidence of a new grad (didn't get a job for 6 months after graduating - so I was desperate to impress).
I was tasked with updating some feature flags that would turn on a new feature for all clients, except for those that explicitly wanted it off. These flags were stored in the database as integers (specifically values 4 and 5) in an array (as a string).
I decided to use the PHP function (array_reverse)[https://www.php.net/manual/en/function.array-reverse.php] to achieve the necessary goal. However, what I didn't know (and didn't read up on the documentation) is that, without the 2nd argument, it only reversed the values not the keys. This corrupted the database with the exact opposite of what was needed (somehow this went through QA just fine).
I found out about this hours later (used to commute ~3 hrs each way) and by that time, the senior leadership was involved (small startup). It was an easy fix - just a reverse script - but it highlighted many issues (QA, DB Backups not working etc.)
I distinctly remember (and appreciate) that the lead architect came up to me the next day and told me that it was rite of passage of working with PHP - a mistake that he too had made early in his career.
I ended up being fired (grew as an engineer and was better off for it) but in that moment and weeks after it, it definitely demoralized me.
I worked there for ~8 months in total.
Wat? Like serious issues, or minor things that can be improved? Because it's very rare in my place of work that there are no comments on a 'PR'. Something can always be improved.
It sounds as if this team made several mistakes, not just one mistake. It's also not clear if the result of these mistakes was that there might be real damage to the spacecraft, or if the result was just wasted time and hours of confusion about why the spacecraft wouldn't start up.
The first mistake is they didn't realize that the multimeter was not only measuring, but it was also completing the circuit.
That sounds like a real bad idea. But if it was totally necessary to arrange it like that, then that multimeter should never have been touched.
That's not just one guy's error. It's at least two guys at fault, along with whoever is managing them, and whoever is in charge of the system that allows it.
The second mistake is with the break-out-box. They think he misdirected power wrongly into the spacecraft. Then they jump to the conclusion that has generated a power surge which has damaged the spacecraft, because it won't start up.
But they're not sure where the power surge went and what might be damaged. Anyhow they're wrong.
The reason the spacecraft won't start up is just because he took the multimeter out of the circuit before the accident.
I'm still sort of confused about what happened or if they ever really figured out what happened.
He said "Weeks of analyses followed on the RAT-Revolve motor H-bridge channel leading to detailed discussions of possible thin-film demetallization".
Does this mean that they decided that the misdirected power surge might have flowed into the RAT-Revolve motor H-bridge channel and damaged that?
The power absolutely did feed into that circuit, they were trying to decide if it would have damaged it (but a motor driver is going to be able to handle power coming from the motor, so they decided that it probably didn't damage it).
Thanks. I understand that better now. The spacecraft did start up, but it seemed as if it could be badly damaged because they were not receiving any telemetry data
No... I was the senior safety-crit signoff on things carrying human lives. I had to look over pictures of parts broken from a crash and have the potential feeling of 'what-if that's my calculation gone wrong'. My joint that slipped. My inappropriate test procedure involving accelerated fatigue life prediction, or stress corrosion cracking. My rushing of putting parts into production processes that didn't catch something before it went out the door.
It's interesting to read people's failure stories from similar fields but, to me, the ones that people so openly write about and get shared here on HN always come across as... well, workplace induced PTSD is not a competition. It's just therapy bills for some of us more than for others.
But my understanding is the default behaviour of most nuclear weapons (other than Hiroshima-style ones) is "blows itself to pieces without detonating the nuclear part", rather than "vapourises everyone within a mile".
Everything needs to go right for a nuclear weapon to actually blow up with a significant yield.
> He gathered his breath. “This is the most important thing I will ever say to you. The human mind is the ultimate testing device. You can take all the notes you want on the technical data, anything you forget you can look up again, but this must be engraved on your hearts in letters of fire.
> “There is nothing, nothing, nothing more important to me in the men and women I train than their absolute personal integrity. Whether you function as welders or inspectors, the laws of physics are implacable lie-detectors. You may fool men. You will never fool the metal. That’s all.”
> He let his breath out, and regained his good humor, looking around. The quaddie students were taking it with proper seriousness, good, no class cut-ups making sick jokes in the back row. In fact, they were looking rather shocked, staring at him with terrified awe.
-- Falling Free by Lois McMaster Bujold
Does it inevitably come down to that for someone? I mean even if its a detail that a procedure couldn’t have caught, someone is responsible for forming good procedures. I suppose there could be several factors. But it seems like ultimately someone is going to be pretty directly responsible.
Just interesting to think about in the context of software engineering and kinda even society at large where an individual’s mistakes tend to get attributed to the group.
I think this is furthermore almost always true of RCAs, which is why blameless post-mortems exist. It's not just to avoid hurting someone's feelings.
Or many people, or no one directly. Space missions come with calculated risk. So someone calculates the risk that this critical part brakes is 0.5% and then someone higher up says, that is acceptable and all move on - and then this part indeed brakes and people die.
Who is to blame, when the calculation was indeed correct, but 0.5% chances still can happen (and itnwould be a lot)? And economic pressures are real, like the limits of physic?
See Murpheys Law, "Anything that can go wrong will go wrong." (Eventually, if done again and again)
https://en.m.wikipedia.org/wiki/Murphy's_law
Astronauts know, there is a risk with every mission, so do the engeneers, so does management. Still, I cannot imagine why anyone thought it was an acceptable risk, to use a 100% oxygen atmosphere with Apollo 1, where 3 Astronauts died in a fire. But that incident indeed changed a lot regarding safety procedures and thinking about safety. Still, some risk remain and you have to live with that.
I am quite happy though, that in my line of work, the worst that can happen is a browser crash.
Especially when, in prior experience there, asbestos had caught fire in the same situation (O2, low pressure).
Even after the fire, the Apollo spacecraft still used 100% oxygen when in space. The cabin was 60% oxygen / 40% nitrogen at 14.7 psi at launch, reducing to 5 psi on ascent by venting, with the nitrogen then being purged and replaced with 100% oxygen.
> See Murpheys Law...
Indeed. I hope that was a joke.
I believe that it can be quite hard sometimes if you have empathy and don't take the "Once the rockets are up, who cares where they come down? That's not my department" approach."
Perhaps some peer support group (possibly facilitated) for people that build safety critical systems or deal with the fallout. Not all companies will provide good counseling etc.
Perhaps the engineering boards / chartered engineer organizations should provide this and fund it from their membership fee, though that would probably scare people off going to the service as they could be afraid of losing their stamp / chartered engineer status / license.
Perhaps this would be dealt with in the past by getting drunk with colleagues in the pub, though alcoholism (or being impaired at work the next morning) is bad and pubs etc. are less popular now.
Group therapy for doctors and nurses is finally becoming a thing, but unfortunately it is completely dependent on being employed by an organization that cares about it.
> I'm instantly transported back to that moment — the room, the lighting, the chair I was in, the table, the pit in my stomach, ...
I could'nt help but think "that sounds like a trauma reaction". Good on them to be able to use that energy to do better! But also not everyone reacts the same way to trauma nor is it easy to compare such reactions to trauma (for example as a hiring question). I feel there are too many social variables at play
I also want to clarify that I've worked in both aerospace and automotive, and the mention of the word 'crash' in my above comment was referring to work I did in automotive, lest someone tries to start wondering 'which one' with regards to an airframe.
For me, the reaction the the stress of having to make sure I was delivering... and the idea that those things out there, I mean.. put it this way I've worked on enough vehicles that a majority of HN readers will have ridden in something utilizing math that I did or parts that I specified, drew, and released, on a road, at least once in the last 15 years.
I once had potential employers ask that 'how would you respond to this kind of stressful situation' question before and I've actually had difficultly getting my answers across because the real stressful shit I can't even talk about without potentially triggering just a horrible social reaction. Or panic attacks. Or potential legal issues.
The roots cause, and someone correct me if this is not accurate, was that the x-ray tested bolts to hold it down were so expensive, that they had been "borrowed" to use on another project, and not returned, so that when the time came to flip the satellite into a horizontal position, it fell to the floor. Repairs cost $135M.
A cautionary lesson in properly checking how exactly events are connected during an incident. Easy to look at two separate signals and assume they must be causal in a particular direction, when in reality it is the other way around.
When I'm asked to share failures, I'm usually not thinking about "that one time when I almost screwed up but everything was fine", instead, I'm thinking of when I actually did damage to the business and had to fix it somehow.
But, to your point, it is still a failure story that _could_ have lead to a much worse outcome than it did. The fact that it didn’t was mostly due to luck.
This is a refreshingly humanizing article, but is also one written from the perspective of a survivor. Imagine if the rover were actually lost. I asked the question "what would you do if the mission failed after all of this work? How could you cope?" to the folks at (now bankrupt) Masten Aerospace during a job interview, and maybe it was a bad time to ask such a question, but I didn't get the sense they knew either. "The best thing we can do is learn from failure," one of them told me. An excellent thing to do, but not exactly what I asked. This to me stands out as the defining personal risk of caring about your job and working in aerospace. Get too invested, and you may literally see your life's work go up in flames.
I would argue that if we don't charge the process to prevent this kind of catastrophic failure mode then we really haven't learned from the failure.
Incidentally, this happened to Lewicki a few years later when Planetary Resources' first satellite blew up on an Antares rocket: https://www.geekwire.com/2014/rocket-carrying-planetary-reso...
I wonder whether Pete had followed this 1989 general aviation/accident analysis story:
> When he returned to the airfield Bob Hoover walked over to the man who had nearly caused his death and, according to the California Fullerton News-Tribune, said: "There isn’t a man alive who hasn’t made a mistake. But I’m positive you’ll never make this mistake again. That’s why I want to make sure that you’re the only one to refuel my plane tomorrow. I won’t let anyone else on the field touch it."
-- https://www.squawkpoint.com/2014/01/criticism/
(The incident above led to the creation and eventual mandated use of a new safety nozzle for refueling, which seems like a better long-term solution than having the people who've nearly killed you nearby to fuel your plane indefinitely: https://en.wikipedia.org/wiki/Bob_Hoover#Hoover_nozzle_and_H...)
In case of electrical connectors, the connectors are often grouped together in such a way as to avoid making wrong connections. Connectors with different sizes, keying, gender, etc are chosen to make this happen. This precaution is taken at design time. JPL is extremely experienced in these matters. There is probably something else left unsaid, that led to this mistake being possible.
Meanwhile, motor controllers using H-bridge is something that's never boring. I once saw a motor control fail so spectacularly that we were scratching our heads for days afterwards. As always, a failure is never due to a single cause (due to careful design and redundancies). It's a chain of seemingly innocuous events with a disastrous final outcome. But the chain was so mind-bending that we had to write it down just to remember how it happened. Recently, I was watching the Chernobyl nuclear disaster and I got reminded of this failure. Our failure was nowhere near as disastrous - but the initial mistakes, the control system instability, the human intervention and the ultimate failure propagation were very similar in nature. Needless to say, it sent us back to the drawing board for a complete redesign. The robustness of the final design taught me the same lesson - failures are something you take advantage of.
I'd expect that the rover body itself would be bespoke this late in the process, although a parallel test vehicle would be useful, do they have that?).
But in case someone fried the rover's electronics I'd think tearing it apart and replacing them while maintaining the chassis should be doable in 2 weeks, but what do I know?
According to Wikipedia they could have stretched those 2 weeks to around 3 weeks, but after that they'd have missed the launch window.
The usual processes are there to have a near-certainty of a working rover, but under these circumstances I'd think they'd just YOLO it and hope for the best.
But that assumes they've got spare electrical components, or alternatively a better use for the booster sitting on the pad than such an improvised mission.
(I really doubt it was fully tested. But why else have a flight spare vehicle?)
*A notable example of this is in the world of RC cars, where rock-crawlers only very recently have started switching to brushless motors using field-oriented control to deliver acceptable very-low-speed behavior. Until FOC controllers became available, brushed motors offered much better low-speed handling.
You can get good control of brushed motors with just a couple of transistors. Good brushless control means FOC, which really requires a fairly capable microcontroller in addition to all the power electronics for variable-frequency drive. While brushed motors certainly have limitations, those were quite well understood by the early 2000s (to the point here that assessing whether or not damage had occurred was "just" a question of "have these few transistors suffered from voltage applied in an unintended manner"). Brushless motors involve way more components with way more integration required to make then small. Far more complexity and potential failure modes need to be understood.
I see what you mean. Yes, agreed.
I don't think FOC type controllers were anywhere near common back then either, which is needed to run a brushless motor smoothly.
There is just so much more that can go wrong with a brushless setup, vs brushed where you just apply power and that's it.
THE LITTLE VAX THAT COULD https://userpages.umbc.edu/~rostamia/misc/vax.html
Knowing that you are allowed to fail, but are very much expected/required to learn from your failure, makes for rather a good employee, in my experience.
Although the time pressure coming with the upcoming deadline is understandable, perhaps the bigger lessons here is that when you are possibly sleep-deprived, and have already pulled too long a shift, you are bound to make avoidable mistakes. And that is the last thing you want on a $500M mission with at limited flight window.
Recently, I was asked if I was going to fire an employee who made a mistake that cost the company $600,000. No, I replied, I just spent $600,000 training him. Why would I want somebody to hire his experience?
Are you suggesting that a competent person never messes up/makes a mistake?
The most fundamental part of life is learning from mistakes, and today even AI is starting to do this. Mistakes and evolution are what _make us_ human and living.
"Today is the safest day to let you drive my truck, cause I know you'll be extra careful"
If this really manual fiddly process was really the only way they could test the motors, I’d say that’s a big failure on the design engineer’s part.
Your answer is good in the general case, but for the anecdote, the design was clearly bad.
For our databases we have separate credentials, compartmentalized access and disallowed “dangerous” commands. This now seems like obviousness, but we only got this years in. Thankfully, no (major) incidents have occurred to this date.
Many have posted of their failures here so I suppose I could share a couple of mine.
- Pushing gigabytes of records into a Prod table only to realize the primary key was off by a digit, rendering the data useless for a go-live. It had to be deleted by the database admins and reloaded, which took precious hours. I forget why, but an update wasn't feasible.
- A perfect storm of systems issues that lead to all servers in the pool becoming unavailable, causing an entire critical system to go dark. We got it back up within minutes, but harrowing nonetheless.
- Realizing hours before a go-live that a key data element was missing, prompting a client who was now in a code freeze to make a change (they were quite upset). Pretty sure I got an unfavorable review from that, but haven't made the same mistake since.
I’m a firm believer that despite all the short comings of US, what makes it great is there are millions of engineers and scientists working to push the frontier of what is possible, and trillions of dollars in economy to fund that into reality.
NASA is truly an inspiration.
And also the private aerospace industry - SpaceX, ULA, Boeing, Lockheed Martin, Blue Origin, Planet Labs etc.
No other country has that.
Don't blame humans for occasional mistakes, it won't stop them from happening.
I figured this was going to be a story about trying to measure voltage with the meter set up on the 10A current range.
(Also, I know I'm breaking the rule "Please don't pick the most provocative thing in an article or post to complain about in the thread." My defense is this is less of a complaint and more just plain curiosity!)
Really seems like anything which when done improperly could cause millions of dollars in damage, there should be a second person reviewing your setup first.
https://www.popularmechanics.com/space/a4288/4318625/?utm_so...
> I had learned from countless experiences in this and other projects that bad news doesn’t get better with age
That's so true! We tend to sit on bad news and hope that somehow time will blunt them; but if anything the opposite happens.
And
> I still remember the shock when Project Manager Pete delivered the decision and the follow-on news: ‘These tests will continue. And Chris [the author] will continue to lead them as we have paid for his education. He’s the last person on Earth who would make this mistake again.’
We sometimes think people who made one mistake will make another one, and it's better to go with the person who doesn't make mistakes. But that's not the correct approach. People who don't make mistakes are often people who don't do anything.
Edit: I suppose he could've been using alligator leads.
Making an only-faces-the-motor breakout, and a separate only-faces-the-driver breakout, might've been prudent, presuming that they used unique and consistent connectors, for instance a single gender always on the motors. But that's quite an assumption and I can imagine a ton of reasons it might not apply.
Googling him, found it was a junior dev at a FAANG. Oops.
An hour later the contacts repopulated, and all was well. But had to be a white-knuckle time for that poor shmuck.
There are more than a few times where I am scratching my head as to "how could my change have possibly broken this" only to remember a couple of hours later that I had made another change somewhere, or rebooted, or changed a config file temporarily.
I guess it just says that we all need to log everything we do, including removing spare multi-meters, so that by looking over the list we can remember these things.
Love this. Whenever people say “failure is not an option”, I get the sense that they don’t really understand how the universe works. It’s like saying “entropy is not an option”. Uh...
The store manager got a complaint about it like a month later, and she tracked down the gift card number from the receipt. It wound up getting loaded with even more money a few days later, and given to someone else.
Anyway, the whole incident got me fired. And to this day, I always check that I put gift cards into bags (on the rare occasion it comes up in my job as a software engineer for NASA missions).
The videos were recorded because they were legal/financial related and also autotranscripted for text search. I would watch them to trace how the problem occured (I know there are many ways to log).
I was the only developer so it was my fault ha. The app had so many parts and I just used an E2E test to make sure everything generally worked. It's cool Chrome has a fake video feed (spinning green circle/beeping).
I've written a bit about it myself - https://flyingbarron.medium.com/lightning-strikes-92482387ca...
1. Bending pins from trying to insert a connector incorrectly 2. Running a full day of testing but forgetting to set up data recording 3. Accidentally leaving a screwdriver next to the hardware inside a thermal vac chamber in an overnight test
Fun times!
As Fred Rogers said, "look for the helpers".
"The First Bug on Mars: OS Scheduling, Priority Inversion, and the Mars Pathfinder" - https://kwahome.medium.com/the-first-bug-on-mars-os-scheduli...
https://www.latimes.com/archives/la-xpm-1999-oct-01-mn-17288...
> NASA lost its $125-million Mars Climate Orbiter because spacecraft engineers failed to convert from English to metric measurements
Not really sure whether this counts as "engineering" though (accidentally taping over a track), nor would I consider this "worse" than potentially destroying 500M worth of advanced equipment.
I have had couple of scars of mine. I feel like sometimes you become risk-averse and When you are launching new things you will face the fear of failure.
"Phrase not found"
But... how? This is exactly the same story that happened in the book. Since Spirit is older is the book, perhaps this was the real-life inspiration.
One of our responsibilities was to execute nightly checklists to run various processes, do backups at correct times, etc. These processes would be things like running calculations on loans, verifying results, etc.
We had a huge ream of checklists to accomplish this and we were supposed to follow them religiously.
We had two very similar applications, one our core and another a core from another bank we bought with the same application but older version and slightly different config. Consequently, we had two tracks of checklists with very similar steps.
One of those steps was to change the accounting date in the system. The application was a terminal app. We would telnet to the server, log in, then we would execute commands in the menu driven app. To change the date we would have to go to a special menu for super dangerous applications. It required the user to log in again.
Our core system, required logging in, selecting that we want to advance accounting date by one day, entering admin password again, pressing enter, then waiting for about 4 hours while the process ran. Then the process would exit back to the menu where the highlighted option would be to advance the day.
Our legacy system required logging in, selecting the option to advance accounting date, entering admin password, pressing enter, then waiting for two hours after which a popup showed to ask a stupid question where we would just always press enter, then wait for another hour until it exits to the menu.
We quickly figured out, that we can just press the enter key twice on our legacy system. The second enter press would just leave there in keyboard buffer and dismiss the popup. This was very useful for us, as this was the only operation that interrupted what would be the only time during night where we could have a kebab...
One night I made a mistake, and I pressed the enter twice... on the wrong system. When I figured out that I did it I realised the process would exit to the menu and then should ask for the admin password.
But, unfortunately, the application had a bug (or a feature). Once it exited to the menu, it came back in but for some reason it remembered that the admin password has already been entered and started advancing the accounting date again without asking for the password.
Unfortunately, the date was December 24. For entire December 24, the entire bank was unable to process any operations while we were restoring from last good backup (before day close) and then redoing eod operations. Then on December 25 as a penalty, I had to sit for entire day with accounting department observing how they manually entered all of the operations that would normally happen automatically on Dec 24th.
One extra key pressed.
Yes. Because it is know that exhausted people make mistakes. The work is too important to let exhausted people screw it up so you should make sure everyone working on it is well rested.
> your solution is throwing “cheap” grad students at it
Yes? It is testing an electric motor. They can do it. The solution is that you employ enough people so nobody needs to work 12 hour heroic shifts.
In the real world there are budget, personnel and hiring constraints. You don’t get to hire all the people you want. You make do with what you have, and try to push the mission forward, even in suboptimal conditions.
Sir, this is NASA not Kerbal space program ;)
I’m sure it varies by location, but my nurse friends only work 3x12, giving them 4 days off per week. Working 12 hour shifts is much more acceptable when you have more days off than days spent working. They’re virtually unavailable on days they work, but then they’re off traveling or having fun for 4 days, some times more if they combine their days off back to back. My close nurse friend routinely takes week long vacations without actually taking any time off at all.
Then such people are doing the "actual work" and overloaded with tasks, working overtime is a rule rather than the exception. The justification for this is that he was "getting experience" for trying to move up in his career. So all good.
Most people are going to remember that he messed up, rather than that he was working overtime to meet expectations, maybe except for the guy that pat him in his shoulder, he saw enough to understand it.
Here's one they just made up: "near miss". When two planes almost collide, they call it a near miss. It's a near hit. A collision is a near miss.
The first mistake is the $500 million rover is fried.
The second mistake is believing the first mistake.
(or put another way, the first mistake cost $500 million, the second mistake which they didn't realise at the time, saved $500 million)
But you can't explain the second mistake without first explaining the first mistake, hence the title.