NASA has a list of 10 rules for software development
cs.otago.ac.nz
cs.otago.ac.nz
https://spinroot.com/gerard/pdf/P10.pdf
The original clearly describes that the coding rules primarily target C and attempt to optimize the ability to more thoroughly check the reliability of critical applications written in C. The original author clearly understands what they're doing, and explains lots of other ways to verify C code.
For what it's worth, the rationales in the original all make perfect sense to me. Perhaps this is because I learned C on tiny systems? I learned C for hardware for implanted medical devices, and our lab did similar kinds of guidelines.
Writing C code for a safety-critical system running on an RTOS really humbled me. Felt like I should make more than I did relative to peers slinging code using 748mb of ram in a browser tab. ;D
Way better than plain old C, yet there are some weaknesses on the armour, and we know better since AT&T's Cyclone project.
Because, you know, all the crusty idiots writing software before you didn't know about the wonders of syntactic sugar and automatically installed dependencies.
Ada Web Server (AWS) [0][1] is Free and Open Source software, but that framework doesn't seem to get much use, and it doesn't inspire confidence that [0] mentions SOAP but doesn't mention JSON. I'm not aware of any proprietary/payware Ada web server solutions.
[0] https://www.adacore.com/gnatpro/toolsuite/ada-web-server
• Airbus Chooses GNAT Pro Ada for Development of Unmanned Aerial System https://news.ycombinator.com/item?id=24488986
• https://www.adacore.com/press/a350
• https://www.adacore.com/press/airbus-selects-gnatpro-for-vsr...
Also before that, design things so that errors cannot happen; for instance use C's type system, as weak and clunky as it is, to make it is hardly possible to pass out-of-range parameters. There's a trend these days to bash-at-will C for its "unsafe" nature, but a significant part of the issues are "between the keyboard and the chair", that is not doing the right thing because it is tedious - that's probably the main defect of C.
As for C’s unsafety — it is a poor engineer who doesn’t account for human factors. A language that doesn’t mitigate PEBKACs, and which lacks affordances to make safe code less tedious, is intrinsically unsafe.
There are a couple of talks by Dr Holzmann on youtube regarding JPL high reliability software development process: "Mars Code" is the one that I remember: https://www.youtube.com/watch?v=16dQLBgOwbE well worth the watch I think.
Were they able to use Ada or even Modula-2 more widely, much of that won't be needed.
Yes, this is because space engineers don't like so much unpredictable things as Garbage Collectors, which used on nearly all functional languages by definition.
In pure C, one could write code with 100% static allocation, so any step could be checked and ensured to run in very exact time limit. - Typical GC could give instability larger to O(c^n) from size of garbage.
Second reason, in most air-space development companies, main are air-dynamic (hardware) engineers, and all others considered as lover priority (even when practically, now software costs more than half of typical plane), and most hardware engineers don't understand functional programming.
Originally publicly published in the IEEE Computer journal (doi:10.1109/MC.2006.212):
* https://en.wikipedia.org/wiki/The_Power_of_10:_Rules_for_Dev...
> If the rules seem Draconian at first, bear in mind that they are meant to make it possible to check code where very literally your life may depend on its correctness: code that is used to control the airplane that you fly on, the nuclear power plant a few miles from where you live, or the spacecraft that carries astronauts into orbit. The rules act like the seat-belt in your car: initially they are perhaps a little uncomfortable, but after a while their use becomes second-nature and not using them becomes unimaginable. while true; do wget $NetOpWibby_website || sudo shutdown; doneWhy five?
In my initial comment I was thinking of four main western countries (there are others), each with multiple court cases that hammered home the core human rights violations inherent in anti-miscegenation laws and forced, often deceptive, abortion policies.
Had the GP commenter here simply asked for an example or an expansion I'd have provided that .. but the "five examples" demand was just .. odd.
They've wandered off with no reply so I suspect that might have been the limit of their rhetoric .. such as it was.
People be fuckin'.
Software is critical for many business, even if no-one dies, millions may be lost, and drive the company into insolvency.
But here it’s being used as a familiar pain point the author assumes everyone deals with.
I wonder if companies would do that today without heavy incentives. I can’t imagine for example a VC backed company doing that.
https://www.forbes.com/sites/brentdykes/2025/01/28/the-data-...?
> I wonder if companies would do that today without heavy incentives.
Didn't Tesla give away all of their patents at some point?https://www.youtube.com/watch?v=iPgDgNtOouo
Some more background:
https://www.volvocars.com/en-ca/news/safety/let-your-head-re...
SaaS - seatbelts as a service, available in 3 packages:
- Starter: 4 tightenings* per day, $2.99/mo
- Standard: 6 tightenings* per day, $5.99/mo
- Premium: unlimited** tightenings, $9.99/mo
*Once exceeded, belt will remain slack. Check your local laws. We will not be held liable
**Subject to a fair usage policy. Policy may change at any time. Please check before driving
All plans support a maximum of 2 occupants. Please subscribe to additional plans for more occupants.
It's a little crazy to me that people are perfectly comfortable going 80+ down the freeway with no belts on in the back. Like the two seat backs are enough. It reminds me of the "no smoking sections" in restaurant that were sectioned off by a half wall.
Old UK safety video on this: https://www.youtube.com/watch?v=TWLmoeoHrP4
Are we supposed to believe that the front seat will somehow move forward from the force of the rear passenger hitting it? And this force will be so great that it will crush the front driver’s skull against the steering wheel? Is that really the take-away here?
If so, that’s a PSA about poor engineering and design of the driver seat and less about rear passenger seat belt safety.
> It is estimated that if all rear seat belts were worn, 120 deaths and
> 1,000 serious injuries could be prevented each year. Back seat
> passengers are three times more likely to die in an accident if they
> are not strapped in, according to the AA.
>
> The organisation says each year more than 50 people in the front seats
> of cars are killed after being hit by back seat passengers who were
> not wearing seatbelts.
edit This 2018 article from the RAC also says it's real: https://www.rac.co.uk/drive/news/motoring-news/drivers-warne...That said, and more on topic: this isn't so much about laws as it is attitudes; plenty of people even today don't bother to wear seatbelts.
> It’s interesting how the seat belt comment sets the time this was written.
I was surprised by this comment, as the PDF version that I read was undated.During my search, I found there is a whole Wiki page dedicated to this paper! Ref: https://en.wikipedia.org/wiki/The_Power_of_10:_Rules_for_Dev...
That Wiki page says 2006.
What we learn from that is that even well written code cannot guarantee everything and you need physical fallbacks like end-switches that remove power from the servos if activated (and whoever thought pulse length is a reliable way to set servos is probably wrong).
But if by digital you meant some hypothetical servo that needs to receive its data as bytes with a checksum — yeah that would work, as long as the thing can do a graceful shutdown on powerloss. But I am not aware of such servos (although I wouldn't be surprised if they existed, on that project I was just a programmer).
Thanks for the correction.
2. Be efficient
3. Have a plan for every edge case you meet
1. The more you buy the more you save
2. Black leather jacket
1. setjmp/longjmp is exception handling
2. exception handling is good
and I take serious issue with that second premise.
Also the loop thing obviously means to put a max iteration count on every loop like this:
for (0..N) |_| {
}
where N is a statically determined max iteration count. The 10^90 thing is silly and irrelevant. I didn't read the article past this point.
If I were to criticize those rules, I'd focus on these points:
* function body length does not correlate to simplicity of understanding, or if anything it correlates in the opposite way the rules imply
* 2 assertions is completely arbitrary, it should assert everything assertable, and sometimes there won't be 2 assertable things
but it needs to be in the language so that the tricky code is write once.
In .NET and Java it costs at least 100x as much time to throw an exceptions as to return an error code.
Other languages may have cheaper exceptions, but for many mainstream languages, you pay a big performance price for exceptions.
That price is often offset by other positive things, so the tradeoff is made willingly and with eyes open.
The conjecture is that if checks are always executed while exceptions are free when not thrown.
That said, predicted branches are also fairly free so arguing about this in the abstract is pointless.
Though profile guided optimization will probably catch this. And it might be that even without, compilers still optimize around this.
Don't know about .NET but in Java it is not throwing exceptions that is heavy but filling in stack trace during construction of a Throwable. It can be made much more performant using constructor that disables stack trace: https://docs.oracle.com/en/java/javase/21/docs/api/java.base...
And wouldn’t the value end up on the CPU the same if it was hardcoded or not?
Perhaps there could be some verification done with hardcoded loop bounds for making sure things work in real time. I’ll read the article before commenting further. Edit: Yeah, the hardcoded bounds is to give some sort of guarantee as to how long a function will take.
Since they're independent, the 2 computers don't actually have to run the same software, I believe during Mars entry descent and landing the standby compute element runs a different less sophisticated but easier to validate version of the EDL code to take over if any fault is detected while the primary software is running. (I was going to do a quick check on dataverse.jpl.nasa.gov to confirm that but it seems to be down)
Also I think a few years ago on the Mars Curiosity rover (2012) there was some corruption in the flash storage on one of the computers that prevents the full flight software from being loaded on to it, so instead it runs a stripped-down version of the code with very limited functionality to function as a lifeboat in case the fully-working computer ever fails. https://ieeexplore.ieee.org/document/9843266
Still, if you're going to use words like "better protect" instead of "protect", I'm not sure I can say you're wrong...
For basic software protection you may want to depend on a watchdog timer and filling unused memory with NOP slides to trap the processor until the timer reboots. If you have hardware controlling something more risky like explosive bolts, you may want stronger assurances that the hardware won't fail by adding lower level redundancies.
From the NASA document:
> Rule: All loops must have a fixed upper-bound. It must be trivially possible for a checking tool to prove statically that a preset upper-bound on the number of iterations of a loop cannot be exceeded. If the loop-bound cannot be proven statically, the rule is considered violated.
> Rationale: The absence of recursion and the presence of loop bounds prevents runaway code. This rule does not, of course, apply to iterations that are meant to be non-terminating (e.g., in a process scheduler). In those special cases, the reverse rule is applied: it should be statically provable that the iteration cannot terminate. One way to support the rule is to add an explicit upper-bound to all loops that have a variable number of iterations (e.g., code that traverses a linked list). When the upper-bound is exceeded an assertion failure is triggered, and the function containing the failing iteration returns an error. (See Rule 5 about the use of assertions.)
Wrt exceptions, you didn't elaborate. Exception handling can be a hot button issue with passionate opinions. My comments are general in nature; they are not specifically directed toward you or the Zig language.
IMHO, exception handling is inherently neither good nor bad- it is merely a mechanism; it has upsides and downsides, and the context in which they are being discussed / analyzed / used should be the guiding factor.
C provides no bounds or overflow checking. In the context of a system which must not fail, you are then forced to use some combination of rigorous assertions, overflow flag checking, and return value verification. It's cumbersome and ugly.
Rust provides both bounds checking and overflow checks (in debug mode at least). Asserts are still available, but are not required in as many instances as in C.
Neither of those languages provide what we would consider to be exceptions or exception handling. An argument could be made that if you squint the right way that Rust's resume after panic is a form of exception handling, but that's a semantic debate, and it would certainly be considered non-traditional if it were to be categorized as such. Wrt C, it had never occurred to me that setjmp()/longjmp() could be used as an exception handling mechanism. I have always seen them used as a context-switching mechanism for e.g. tasking (green-threads).
In languages that provide exception handling, both runtime and user defined exception conditions are handled by the runtime itself (oftentimes behind the scenes). This mechanism allows one to provide a catch-all lexical scope for handling exceptions, which can be cleaner and more ergonomic from a programmer perspective and can be simpler to reason about (although this is certainly debatable). It also provides an opportunity to handle conditions that might be unknown or unforeseen at the time the code is being written. Defer semantics in a language without exceptions might be considered a mid-ground approach.
The usual complaints with exceptions are: (1) that they essentially constitute hidden control flow and hidden code execution that happens outside of the plain reading of the source; (2) that you pay the performance penalty for checking / verifying the exception conditions at every line of code (also see #1); (3) that the conditions considered "exceptional" by most implementations are instead just run-of-the-mill error conditions that should be handled explicitly; and (4) that there is no well-defined structure or convention for where and when to handle exceptions and when to pass them up the stack (again, see #1) which results in a spaghettification of sorts.
C++ (optional), Java, JS, Python and C# (among many others) all provide exception handling and are all mainstream.
My general rule of thumb (which is always subject to situational variance) is that application code benefits from robust exception handling while systems level or performance critical code should not use exceptions, or at a minimum should be very judicious with usage.
Exceptions can be abused and misused like any other feature, but the reduction in repetitive manual error checking (see Go) can be a win for many teams.
YMMV of course- we have all been dragged into the 7th circle of hell at one time or another, and programming features are a lot like liquor; once you've gotten sick on one, it's near impossible to go back.
It seems like if there's 1 assertable thing then it's trivial to also assert an inverse condition. Before arguing that it would be redundant, remember that in the presence of random bitflips and other anomalies caused by radiation exposure, logic may not operate as deterministically as expected.
My own personal approach in JavaScript is to avoid defining functions in a nested manner unless I explicitly want to capture a value from the enclosing scope.
This is probably due to an outdated mental model I had where it was shown in performance profiling that a function would be redefined every time the enclosing function was called. I doubt this is how any reasonable modern JavaScript interpreter works although I haven't kept up. Since the introduction of arrow functions (a long time ago now in relative terms) their prolific use has probably lead to deep optimizations that render this old mental model completely useless.
But old habits dies hard I guess and now I keep any named function that does not capture local variables at a file/module scope.
A lot of the other notes are interesting and very nit-picky in the "technically correct is the best kind of correct" way that older engineers eat up. I feel the general tone of carefulness that the NASA rules is trying to communicate to be very good and I would support most of them in the context that they are enforced.
I use variable argument macro's to implement debug print functions for debugging. And I just leave them in the code as a form of documentation. Looking at the debug print statements tells you a lot of about what the code is doing.
> The use of trampolines requires an executable stack, which is a security risk. To avoid this problem, GCC also supports another strategy: using descriptors for nested functions. Under this model, taking the address of a nested function results in a pointer to a non-executable function descriptor object. Initializing the static chain from the descriptor is handled at indirect call sites.
> On some targets, including HPPA and IA-64, function descriptors may be mandated by the ABI or be otherwise handled in a target-specific way by the back end in its code generation strategy for indirect calls. GCC also provides its own generic descriptor implementation to support the -fno-trampolines option. In this case runtime detection of function descriptors at indirect call sites relies on descriptor pointers being tagged with a bit that is never set in bare function addresses. Since GCC’s generic function descriptors are not ABI-compliant, this option is typically used only on a per-language basis (notably by Ada) or when it can otherwise be applied to the whole program.
> For languages other than Ada, the -ftrampolines and -fno-trampolines options currently have no effect, and trampolines are always generated on platforms that need them for nested functions.
From: https://gcc.gnu.org/onlinedocs/gccint/Trampolines.html
You can actually do it in C, too! It just requires setting up a few macros, and getting it right requires a fairly decent understanding of the underlying hardware. (via TARGET_CUSTOM_FUNCTION_DESCRIPTORS).
The processors I program for don't have no execute stacks so the use of a trampoline is of no consequence.
Also my memory was that no execute stacks were heralded as the solution to stack smashing attacks. No execute stacks problem solved. And turned out if you're vulnerable to stack smashing no execute stacks won't save you.
So I feel we threw out something good, nested functions for some cargo cult security. And since no one uses them the reactionary luddites on WG14 are free to hope they go away.
In some cases inline functions (if they fall within V8 optimisations) can be re-used more cheaply than the overhead of non-inline alternatives like binding or currying
[1] e.g. https://nodis3.gsfc.nasa.gov/displayDir.cfm?t=NPR&c=7150&s=2...
>> The MISRA guidelines for Rust are expected to be released soon but at the earliest at Embedded World 2025. This guideline will not be a list of Do’s and Don’ts for Rust code but rather a comparison with the C guidelines and if/how they are applicable to Rust
/? Misra rust guidelines:
- This is a different MISRA C for Rust project: https://github.com/PolySync/misra-rust
- "Bringing Rust to Safety-Critical Systems in Space" (2024) https://arxiv.org/abs/2405.18135v1
...
> minimum of two assertions per function.
Which guidelines say "you must do runtime type and value checking" of every argument at the top of every function?
The SEI CERT C Guidelines are far more comprehensive than the OT 10 rules TBH:
"SEI CERT C Coding Standard" https://wiki.sei.cmu.edu/confluence/plugins/servlet/mobile?c...
"CWE CATEGORY: SEI CERT C Coding Standard - Guidelines 08. Memory Management (MEM)" https://cwe.mitre.org/data/definitions/1162.html
Also important and not that difficult, formal design, implementation, and formal verification;
"Formal methods only solve half my problems" https://news.ycombinator.com/item?id=31617335
"Why Don't People Use Formal Methods?" https://news.ycombinator.com/item?id=18965964
Formal Methods in Python; FizzBee, Nagini, deal-solver: https://news.ycombinator.com/item?id=39904256#39958582
SAST and DAST tools can be run on_push with git post-receive hooks or before commit with pre commit. (GitOps; CI; DevOpsSec with Sec shifted left in the development process is DevSecOps)
While the criticism of rule 3 is right in that there is a dependence on the compiler, it is still a prerequisite for deriving upper bounds for the runtime by static analysis on the binary. This is actually something that is done for safety-critical systems that require a guaranteed response time, based on the known timing characteristics of the targeted microprocessor.
It's definitely that. Static stack size analysis in the embedded world is a long-standing paradigm and recursion and function pointer indirection defeat that.
If the point of rule was that NASA was arguing that recursive constructions are always harder to reason about than iterative ones, then obviously NASA is wrong.
[1] There are some other related rules like "no alloca()".
> Rationale: Simpler control flow translates into stronger capabilities for verification and often results in improved code clarity. The banishment of recursion is perhaps the biggest surprise here. Without recursion, though, we are guaranteed to have an acyclic function call graph, which can be exploited by code analyzers, and can directly help to prove that all executions that should be bounded are in fact bounded. (Note that this rule does not require that all functions have a single point of return – although this often also simplifies control flow. There are enough cases, though, where an early error return is the simpler solution.) [0]
Folks spent a while (centuries?) iterating on the standard page and character sizes. It makes sense to me that what we landed on wasn’t solely due to the limitations of paper, but also the limitations of humans.
The criticism is (mostly) super contrived and totally misses the wisdom behind why some of these rules were made in the first place. A lot of the author's points are very reminiscent of the same classic rebuttal against criticisms of the C language for being insecure: "It's perfectly fine if you're good and don't make stupid mistakes." That's just not a very mature view of working with groups of human beings.
Most of these rules are designed to reduce errors that are difficult for humans to see by making the code more readable, deterministic, and avoiding situations that can lead to unintended behavior that is subtle in its true complexity. Creating a series of "Gotchas" where it perhaps negates that idea in an obscure situation doesn't really mean that the rules don't tend to produce code that is more reliable and auditable.
Some of these rules really do seem kind of anachronistic, but... then again there's still a lot of old FORTRAN code and the like running on NASA hardware.
From that perspective, the avoidance of recursion is more compelling. Plus, Fortran didn't support it...
It's a little hard to duck out to Mars to reboot the lander after you bricked it with a recursive function that never exited.
setjmp() and longjmp(), are a poor way of handling exceptions, as any cleanup code won't be executed. Of course, following the spirit of these rules, one would not have resources that need cleaning up but still.
The main issue, that is never mentioned, because it will be different in each application, is what to do when something goes wrong. Say an iteration limit is exceeded, or the fixed resources allocated at startup are not enough.
I see that is repackaging of low level rules and not some „magical 10 development rules” that would apply to your yet another CRUD app.
Not to bash article just informing people who would write „software engineering is not real engineering” - it is real engineering and there are norms. Just because someone did not hear about standards and rules and practices doesn’t mean there are none.
Argument is only that there is a lot of proper engineering in software development - whether it is implementing memory safety by default in new languages or having organizational processes. All is there and I only dislike opinions that say software development is immature new field. It is not.
Article was also about NASA organization rules so pointing to MISRA was argument „it is. It only NASA and those are not so special rules”.
Things like spinlocks or CAS (Compare-And-Swap) are elegant and safe solutions for concurrency, and AFAIK you can’t really limit their upper bound. Others in the thread have pointed out that those are more guidelines than rules - still, not sure about this one.
- https://www.fastcompany.com/28121/they-write-right-stuff (pdf: https://www.eng.auburn.edu/~kchang/comp6710/readings/They%20...)
11. Use strict typing for all scalar types. Do not mix imperial and metric units.
Value types to the rescue. I use them liberally.
> This does a bounded number of iterations. The bound is N^10. In this case, that's 10^90. If each iteration of the loop body takes 1 nsec, that's 10^81 seconds, or about 7.9×10^72 years. What is the practical difference between “will stop in 7,900,000,000,000,000,000,000,000,000,000,000,000,000,000,000,000,000,000,000,000,000,000,000,000 years” and “will never stop”?
... yes, you're very smart and have found a way to technically satisfy the requirement while building a broken program
Clearly the upper bound rule is to make it easy to reason about worst-cases so setting a very high one isn't really the gotcha they seem to think it is. It just means you can look at the high bound and go "this is too high, fix it"
How do you implement a state machine, like a byte code interpreter? You need some way to dispatch functions based on opcodes. You could argue it has no place in mission-critical code, but well-understood, time-bound state machines (like Plan9-style regexes or BPF) can actually greatly increase code readability.
Also: Greenspun's tenth rule.
int const N = 1000000000;
for (x0 = 0; x0 != N; x0++)
for (x1 = 0; x1 != N; x1++)
...
for (x9 = 0; x9 != N; x9++)
-- do something --;
Rule 1 would disallow this. Probably other rules as well, depending on the details for your code analysis tools.Also the article seems to imply the rule about recursion was about knowing your program would terminate, but the rationale from the rules says it's about verifying execution is bounded. That's a somewhat different concept. With recursive calls you have the stack to consider.
Unless you work on commercial aircraft avionics. NASA flying a probe into Mars makes them look dumb. Flying an aircraft into a mountain haunts people for the rest of their lives.
I don't think they are saying that if you write some kind of service program at NASA, it cannot have an infinite loop for taking requests. Or if you write a line-by-line file processing utility, it has to cap out at a maximum number of lines so that the loop is statically bounded. Or that if you write some tool to process syntax, perhaps a compiler, that you cannot use recursion.
These are not embedded, safety-critical things.
But, as a C#/C++/Typescript developer, I don't have the first clue how any I might go about implementing those rules.
Anybody know of patterns for this kind of safe programming that I can use?
Does just using a language like Zig automatically make my code 'safe'?
Anybody know if generating code from a proof checking language like Lean would satisfy all these rules?
You can overcome lot if you invest a lot for type system, but that depends on the developer.
Rule 1: Don't write any code that can make the ship go ka-boom.
Rule 2: We need to consistently use the f'ing metric system and SI units - and don't you forget it.
Rule 3: ... You haven't already forgotten about rule #2, have you?
typedef struct { float value; } meter;
typedef struct { float value; } feet;
Using strong typing, you can avoid passing values in wrong units.
https://github.com/mpusz/mp-units
which does all sorts of checking, allows arithmetic with proper accounting for units and so on. But yes, the basic notion is wrapping things in structs which can't be simply assigned to each other.
Maybe the answer is that strong typing should somehow continue outside of the individual programs and be embedded in file formats as well?
> The assertion density of the code should average to a minimum of two assertions per function.
I would actually say the opposite in modern languages like Rust. Assertions are a sign you have failed to encode requirements in the type system.
Maybe. You still shouldn't write unsafe code without asserts.
Gotta appreciate the things that the compiler takes care of for me.
While I wish the stop-the-world GC pause were more predictable and took less time, I’ll gladly pay the cost for safety.
Zig seems interesting, too. It feels like a better C, and I’m closely following its development as it marches toward 1.0.
So I would just prepend the document with "Rule 0: Try to avoid writing critical code in C."
These NASA principles are more about enabling better possible static analysis of the code and ease of someone else, maybe decades later, debugging or pushing changes to something likely on another planet.
Also you have to remember space based computing lags well behind terrestrial computing because of the radiation hardening. They are often still dealing with legacy systems that might be 8 bit with very limited memory, they were still in the hardware expensive engineers cheap mode until well into the nineties, if not later. Rad750s run at 400 mips and were, and maybe are, preferred choice of processor.
Even then, the essence (code clarity, robustness, tooling/static analysis, etc) is a good guideline for general code. You can play fast & loose with prototypes/MVPs, or one-off scripts, but once it's time to bring it to production and maintaining it long-term, you will be grateful to yourself for keeping stuff clean.
In theory yeah, but I've very rarely encountered it in the wild.
It's hard to use right in any code that has cleanup to do, without building a lot of infrastructure for resource handling.
bueller? bueller???