A single line of code brought down a half-billion euro rocket launch
jam.dev
jam.dev
Blaming a system failure on a single point like this dooms that system to repeat similar failures (albeit in another element) in the future.
There are numerous testing, quality and risk controls that could've been in place. There are probably even a few people who didn't do their job (besides the one person a decade ago who wrote the 'single line'). The point isn't to pin blame on any one point, but to look at the system (people, processes, technology) and try to understand why the system is fragile enough that a single person's error is able to escalate into a half-billion euro error.
By focusing in on the point of failure, you end up falling victim to survivorship bias [0]. It is how you end up with developer teams swamped with unit-testing requirements and test coverage metrics, but still somehow end up with errors that impact the end-user anyway. It is how you get company surveys that always seem to miss the point, saying that the measures they implemented to improve company culture worked, yet everyone is burning out and miserable.
The mistake was an incorrect specification. A programming tool can't identify that you've made the wrong thing, which is why we need "rigamarole" to validate the spec. That's what the systems engineers are for.
Had they validated more, they would have realized they made the wrong thing.
> the main computer interprets the data as real navigation data and considers it as an indication that the rocket is wildly off-course
Not clear what sort of magnitudes we are talking about but saturation could have worked here and avoided the problem.
But an exception handler could've helped too.
> The code wasn’t necessary after takeoff, it was only part of the launch pad alignment process.
A supervisor for this task could have decided to ignore an overflow fault after launch.
I think people who can write in-bounds code and type correct code with no safety rails have a leg up to write really good code.
The funny thing is that you could pick and choose to attribute any error to a 'single line of code'
https://how.complexsystems.fail/
I cannot read that one without thinking in the descriptions and analyses of disasters like the sinking of the MS Titanic, the Chernobyl disaster, the loss of the Challenger space shuttle, or the Fukushima disaster. Many, many points in the article seem correct for all of them.
The one line of code was the spark, yes, but the catastrophic consequences were due to a series of poorly-designed failsafes and insufficient testing.
Ariane 5 – Flight 501 Failure (1996) - https://news.ycombinator.com/item?id=30556601 - March 2022 (1 comment)
An overflow error costing 500M dollars (1996) - https://news.ycombinator.com/item?id=18939625 - Jan 2019 (20 comments)
The Explosion of the Ariane 5 - https://news.ycombinator.com/item?id=5331474 - March 2013 (58 comments)
There must be others. Anyone?
At Boeing, the backup system runs on a different CPU architecture, with a different program design, a different programming language, and a different team that isn't allowed to talk with the team on the other path.
Yeah that's not restoring.
The mistake the pilots made was not then turning off the trim system (which also overrides MCAS).
The first MCAS incident ended with the pilots restoring trim to normal with the thumb switches a couple times, then turning off the stab trim system, then continuing on and safely arriving at their destination. The media never mentions this.
Inventing your own meanings for words in order to advance an argument is a waste of time for you and me.
> You're trying to tell me a system that brought down 2 airliners isn't that bad because it didn't bring down others.
I'm saying that pilot error was a contributing cause to the two crashes, and the media narrative that the pilots could not have saved the situation is incorrect. Boeing even sent out an Emergency Airworthiness Directive to all MAX pilots after the first crash, detailing the two step recovery procedure (restore normal trim with the thumb switches, then turn it off). The EA pilots did not do that.
There were numerous causes of the accident that all had to combine to result in the crash. One of the causes was the single path design of the MCAS system.
All of the causes needed to be addressed.
> Finding nuance to defend the corporate player that killed hundreds of innocents is disgusting.
I did not defend Boeing. Besides, do you want airliners to be safe or not? If you want to fly safely, you've got to address all factors in a crash, including pilot error.
Cute false dichotomy tho.
>> The worst part? The code wasn’t necessary after takeoff, it was only part of the launch pad alignment process. But sometimes a trivial glitch might delay a launch by a few seconds and, in trying to save having to reset the whole system, the original software engineers decided that the sequence of code should run for an extra… 40 seconds after the scheduled liftoff.
The author appears to be using a different definition of "dead code" than I'm used to. To me, dead code is code that is no longer called by anything else, and has no chance of running. Maybe a more accurate term is "legacy code"?
That's, oh man, that's not how they're stored or how you should think of it. Don't think of it that way because if you think "oh 1 bit for sign" that implies the number representation has both a +0 and a -0 (which is the case for ieee 754 floats) that are bitwise different in at least the sign bit, which isn't the case for signed ints. Plus, if you have that double zero that comes from dedicating a bit to sign, then you can't represent 2^15 or -2^15, because you are instead representing -0 and +0. Except, you can represent -2^15, or -32,768, by their own prose. So there's either more than just 15 bits for negative numbers or there's not actually a "sign bit."
Like, ok, sure, you don't want to explain the intricacies of 2's complement for this, but don't say there's a sign bit. Explain signed ints as a shifting the range of possible values to include negative and positive values. Something like
> With 16-bit unsigned integers, you can store anything from 0 to 65,535. If you shift that range down so that 0 is in the middle of the range of values instead of the minimum and your 16-bit signed integer now covers everything from -32,768 to +32,767. Anything outside the range of these values and you’ve run out of bits.
Not a downvoter, but: your concept of "shifting the range" is also misleading.
In the source domain of 16-bit numbers, [0...65535] can be split into two sets:
[0...32767]
[32768...65535]
The first set of numbers maps to [0...32767] in 2's complement.But the second interval maps to [-32768...-1].
So it's not just a "shift" of [0...65535] onto another range. There's a discontinuous jump going from 32767 to 32768 (or -1 to 0 if converting the other direction).
And actually, we don't know if the processor used 2's complement or 1's complement -- if it was 1's complement, they would have a signed 0!
I think they'd have to say "remapping" the range? On the whole, I think OP did about as well as you're going to do, given the audience.
We can infer it used two's complement, and absolutely rule out one's complement or any signed-zero system, because the range is [-2^15,2^15) and with a signed zero you can't represent every integer in that range in 16-bits, you have one too many unique numbers.
We know the signed int they're talking about can't be the standard 1's complement because its stated range of values is [-32768, 32767]. If the representation were 1's complement the range would be [-32767, 32767] to accommodate the -0. It could be some modified form of 1's complement where they redefine the -0 to actually be -32768, but that's not 1's complement anymore.
Every negative number has 1 as it's first bit, every positive number (including 0) has 0 as its first bit. Therefore first bit encodes sign. The other 15 bits encode value. They may not use the normal binary encoding for negative integers as you'd expect from how we encode unsigned integers, but you cannot explain every detail every time.
https://zjkmxy.github.io/posts/2021/11/twos-complement-2-adi...
The most common situation in which it crops up is when dealing with quantities that require fractional units/arithmetic of some commonly discrete unit of measure. For example, you implement some complex logic to do request sampling, and in your binary you convert the total number of active requests to a float, add some stuff, divide some stuff, add some more stuff, multiply it again, then convert back to an int something like “number of requests that should be sampled.” Because floating point operations are non-associative, non-distributive, and commonly introduce remainder artifacts, you can end up with results like sampling 1 more request than there are total requests active, even when the arithmetic itself seems like that should be impossible.
This is also common when dealing with time, although typically the outcome is not that bad. Despite time having a simple workaround of just changing the unit of measure (eg using milliseconds instead of seconds) and using int operations on that, because people don’t know why they shouldn’t use floating point operations in this case, they don’t always reach for it.
The worst is when some complicated operation is done to report a float (or int converted from a float) as a metric. In the request sampling example, that would likely be noticed quickly and fixed. But when the float value looks reasonable enough and doesn’t violate some kind of system invariant, it can feed you bad data for a very long time before someone catches it.
I noticed some odd behaviour recently when using ruby to save to Postgres where the handoff between the two systems introduced imprecision in the saved value. Didn’t get to dig into it because it wasn’t a priority but it’s definitely an annoying unanswered question.
A simple rule of thumb is to try to avoid using floating point values at all outside of contexts like scientific simulations. For basic situations, you can almost always use either a library (for big numbers, decimals, fractions, etc.) or express your logic with ints by using established patterns (make the unit of measurement smaller/bigger, explicitly round up or down, etc.). Any time you take something that is an int 99% of the time, convert it to a float for something, then convert it back to an int, you are doing something wrong.
If you just want to fix the odd behavior: adjust the schema so that you only work with whole numbers. For example, in the database schema, you can use DECIMAL instead of REAL/DOUBLE columns, or use two columns to specify the ratio of two integers (for example num/denom in frame rates in various video containers/codecs). In the application code: work only in whole numbers (e.g. cents, satoshis) instead of fractionals, using bigint or string types as applicable instead of e.g. double.
https://docs.oracle.com/cd/E19957-01/806-3568/ncg_goldberg.h...
https://jvns.ca/blog/2023/01/13/examples-of-floating-point-p...
An oceanographic research ship is doing gravity and magnetic surveys off the coast of Brazil.
Suddenly, data acquisition software crash!
byte day_of_year;It will always be a single line of code. The nature of most programs is to execute commands in a sequence. Eventually you hit one that fails.
Hell, you could reduce it to even be less than a line of code. It could be a single variable. A single instruction. It could be a couple bits. A couple bad 1’s and 0’s in memory blew up a multibillion dollar rocket launch.
Oh, excellent possible interview question? "Write some code that reliably converts the full range of possible 64 bit floating point values to a 16 bit signed integer. What are the issues you'll have to deal with and what edge cases might arise?"
I’ve interviewed 8 web devs over the last two weeks, each with years of experience, each asking $115k+, and more than one couldn’t figure out how to take an array of objects each with a ‘category’ property and output an array with the unique values of ‘category’. (this is one line of code)
Only one could successfully wire up a <select> populated with the list of unique categories and then filter the original list based on the selected category.
By the way, these interviews were done over Zoom with screen sharing with the candidate able to use their own dev environment and browser of their choice.
The interviews were allotted one hour and all candidates took longer than the scheduled time.
It’s been beyond depressing. I’m ready to just start asking for FizzBuzz again.
The first issue is that we're taking 64 bits of data and trying to squeeze them into 16 bits. Now, sure it's not that bad, because we have the sign bit and NANs and infinities, but even if you toss away the exponent entirely, that's still 53 bits of mantissa to squeeze into 16 bits of int.
The second issue is all the values not directly expressible as an integer, either because they're infinity, NAN, too big, too small, or fractional.
The only way we can overcome these issues is to decide what exactly we mean by "converts", because while we might not _like_ it, casting to an int64 and then masking off our 16 most favorite bits of the 64 available is a stable conversion. That might be silly, but it brings up a valid question. What is our conversion algorithm?
Maybe by "convert" we meant map from the smallest float to the smallest int and then the next smallest float to the next smallest int, and then either wrapping around or just pegging the rest to the int16.max.
Or maybe we meant from the biggest float to the biggest int and so on doing the inverse of the previous. Those are two very different results.
And we haven't even considered whether to throw on NAN or infinity or what to do with -0 for both those cases.
Or maybe we meant translate from the float value to the nearest representable integer? We'd have a lot of stuff mapping to int16.max and int16.min, and we'd still have to decide how to handle infinity, NAN, and -0, but still possible.
Basically, until we know the rough conversion function, we can't even know if NAN, infinity and -0 are special cases and we can't even know if clipping will be an edge case or not. There's lots of conversions where we can happily wrap around on ourself and there are no edge cases, and there's lots of conversions where we have edge cases, but we can clip or wrap, and there's lots of conversions where we have edge cases and clipping/wrapping."
Step 1: Enumerate all possible inputs. <--- this is the important part
Step 2: Map each input to something in the output domain.
Step 3: Are you using Ada??? C++ is *right out*.
...it looks like the proper name for what is needed is a "non-injective surjective total function".Maybe I have no business writing floating point code.
There should be layers upon layers of safeties to prevent this dumb thing from happening. The computer should know the position, orientation and velocity of the rocket at any point in time and new signals should be interpreted in the context of what the computer already knows and in context of what other sensors are telling. It is not like the rocket can turn itself around in 1ms and if it does there probably isn't much it can do anyway.
This suggests to me the problem is not the bug, it is the overall quality of development.
It’s easy to judge with hindsight though. Every rocket development project that has ever happened on the planet has had unintended explosions and accidents, and they are always staffed with brilliant people. Layers of safety at some level might only make the problem harder, more code just adds more complexity and failure points. Since there is an uncountable number of ways for something seemingly innocuous to break a rocket, and until we’ve all tried making rockets, it’s probably best to take away the lesson that fundamentally making rockets is highly prone to catastrophic failure.
What I meant was other signals -- other information about state of the rocket. There is lots of sensors in these things if only so that you can figure out when stuff goes wrong.
> It’s easy to judge with hindsight though. Every rocket development project that has ever happened on the planet has had unintended explosions and accidents,
No, that's bad excuse. Accidents do happen but can only be excused when a reasonable effort to prevent it has been taken.
For example, Challenger disaster resulted in such a huge shakeup in Nasa exactly because reasonable precautions were NOT taken vs. other accidents which were truly unforeseen and were due to lack of knowledge/experience.
The question is about taking reasonable precautions. It is reasonable effort to design a system that will drive multiple half a billion dollar rockets with some level of care to ignore absolutely idiotic signals.
I do it on my home projects and at work with non safety critical applications. Why can't they do it for such a critical project?
They can, they did, and they still exploded a rocket! Again, hindsight is 20/20. This question isn’t very reasonable to ask this way (with incredulity) until you’ve successfully built several rockets yourself.
Imagine discussion after Challenger disaster.
The Commission: So why did it blow up? Could anything have been done?
The Nasa: These things just happen! Hindsight is 20/20. Also, you don't have enough experience to point out our engineering problems until you exploded couple of them yourself!
This is important because assuming that something dumb happened is part of the problem too, and that’s what your comments above attempt to communicate. Pretending like it was easy to avoid is to be intentionally ignorant of the fact that nobody ever has avoided this problem, not in rocket launches, not in web development, not in cars or video games, or in any code of any significant size. Safety critical engineering has to absorb this fact deeply, and nobody can walk into it thinking, well @twawaaay did it with their home project, so all we need is duh multiple layers of testing and redundancy. Yeah, they absolutely had multiple layers of testing and redundancy, they had everything you’ve ever thought was a good idea for writing safe code, and then 10x more than that, it might be worth reading more about the history before jumping to such conclusions.
Also this FP to integer does not smell any better. Anybody who's been interested in programming for any length of time will learn this is just a bad idea asking for trouble.
If I was tech lead for the project I would definitely make sure there is static analysis that prevents these kinds of conversions from happening.
The levellers job was to smooth out any waves that might have acquired during the rolling process - almost like a clothes iron. The gap between the rolls needed to be adjusted by hydraulically positioning backup rolls that are even able to bend those work rolls across their width (maybe 3000mm). As you are always intending apply a huge amount of force anyways, to achieve the desired results, it was a mix of metallurgical driven algorithms and hard limits to doing the "setup".
While there was always an operator that had to accept the setup before the run, there was always the risk of hitting the machines surfaces too hard, and straining components, and maybe causing a prolonged as expensive outage. Obviously the biggest risks were when there were changes or even experiments by both engineers and metallurgists. It was fun times as a quite junior engineer, and think there were a few times when over zealous setups resulted in some big noises. But I don't think I broke anything fortunately.
Even worse, why was I completely unsurprised, nay expecting this to be the case when I clicked?
(to the article's credit, it didn't quite start in the typical "George was walking his dog home when he noticed something wrong" fashion...)
This is so unbelievably untrue. I've never seen code anywhere that waits to fail before doing the right thing.
This is exactly why I think exceptions are mostly useless, someone has to anticipate the problem, so why not write something that works right the first time. There are cases where exceptions can happen, but I don't think floating point arithmetic should be considered one of those cases.
Nobody writes 100% correct code. Ever.
These stories of gnat-brings-down-empire are, to me, always missing the point. There are going to be bugs. The hard part is in creating management around the actual software engineering such that this is not a problem.
In this case the syntactic sugar helps deal with passing data across function boundaries and up the stack in a way that can make the code more readable.
Every function could handle returning its own error states, and then higher level functions could aggregate and check and return the error states of every function it calls… or you might find it saves a massive amount of boilerplate to just use try/catch!