NASA JPL C Coding Standard [pdf]
lars-lab.jpl.nasa.gov
lars-lab.jpl.nasa.gov
longjmp banning is also slightly questionable (although I can see why because it is very easy to do wrong). I use it inside of my code as part of an STM implementation (so begin_tx() setjmps[1], abort_tx() longjmps; its faster than manually unwinding with if(tx error) { return; } spam in deep call stacks.)
Using longjmp for this makes writing code much easier (no needing to error check every single tx function call), so less chance for bugs to slip in.
1: The only ugly part of that is begin_tx() is a function macro, which I prefer never to use in code that is executed; I tolerate it in "fancy template-like generator" setups, though.
Edit: Andrew, thanks for pointing out my ignorance, I've never coded FORTRAN, and have been thinking it's a semicolon this entire time. What fitting irony. Sort of makes for an even better story now. >_<
if it's correct than it was actually a decimal point instead of a comma.
But you'd never see something like that in flight control software. Simplicity begets correctness. Correctness begets safety.
When the software is flying a rocket ship, I'm okay with it being 10% longer but 10% safer.
Writing code checkers for these sorts of rules is a really interesting exercise and it helped me grow a lot as a programmer! I went from having no exposure to formal languages, parsing, and grammars to actively playing around with these concepts to try and help build more reliable software. It was a humbling, challenging, and incredibly rewarding experience.
Sometimes, a rule is extremely simple to implement. For example, checking a rule that requires that an assert is raised after every so many lines within a given scope is just a matter of picking the right sed expression. Other times, you really need an AST to be able to do anything at all.
A rule like "In compound expressions with multiple sub-expressions the intended order of evaluation shall be made explicit with parentheses" is particularly challenging. I spent a few weeks on this rule! I was banging my head, trying to learn the fundamentals of parsing languages, spending my hours diving into wikipedia articles and learning lex and yacc. The grad students at LaRS were always extremely helpful and were always willing to help tutor me and teach me what I needed to learn (hi mihai and cheng if you're reading!). After consulting them and scratching our heads for a while, we figured we might be able to do it with a shift-reduce parser when a shift or reduce ambiguity is introduced during the course of parsing a source code file. This proved beyond the scope of what I'd be able to do within an internship, but it helped me appreciate the nuance and complexity hidden within even seemingly simple statements about language properties.
Automated analysis of these rules gives you a really good appreciation of the Chomsky language hierarchy because the goal is always to create the simplest possible checker you can reliably show is able to accurately cover all the possible cases. Sometimes that is simple as a regular language, but the next rule might require you to have a parser for the language.
For what it's worth, this is only one of the ways the guys at LaRS (http://lars-lab.jpl.nasa.gov/) help try to improve software reliability on-lab. Most of the members are world-class experts in formal verification analysis and try to integrate their knowledge with missions as effectively as possible. Sometimes, this means riding the dual responsibility of functioning as a researcher and a embedded flight software engineer, working alongside the rest of the team.
If anyone's interested in trying out static analysis of C on your own, I highly reccomend checking out Eli Bendersky's awesome C parser for Python (http://code.google.com/p/pycparser/). I found it leaps and bounds better than the existing closed-source toolsets we had licenses for, like Coverity Extend. At the time, it had the extremely horrible limitation of only parsing ANSI 89, but Eli has since improved the parser to have ANSI 99 compliance. Analyzing C in Python is a dream.
Another thing that happened around this time is getting licensing for Coverity and other tools, and introduction and promotion of static code verification, even for non-flight software.
Here's an interview with Gerard:
Hmmm....how about this? If the code is parenthesized enough, then the precedence and associativity of the operators has no effect on the shape of the parse tree. So, if you take the expression and repeatedly make random changes to the operators and parse it, and you keep getting the same shape for the parse tree, it is sufficiently parenthesized.
Smart.
The following grammar accepts input like 3 + (5 * 7) but rejects 3 + 5 * 7.
The key is that, if your expression doesn't start with a parenthesis, you know that all the operators at that level have to be the same. (I assume that sums or products of several things like 1 + 5 + (2 * 3) are permitted without parenthesizing further.)
Also, tool choice matters. ANTLR is an LL parser and Yacc/Bison are LR parsers; IMHO with LL it's much easier to understand what's going on. This grammar would need substantial rewriting for Yacc to deal with the fundamental differences between LL and LR parsing.
(edited to deal with HN markup issues related to asterisks and fix implementation bugs)
grammar parencheck;
prgm : expr EOF ;
expr : atom ( (PLUS poratom)*
| (TIMES poratom)* )
| '(' expr ')' ;
poratom : atom | '(' expr ')' ;atom : INT | VAR ;
PLUS : '+' ;
TIMES : '*' ;
INT : ('0'..'9')+ ;
VAR : ('A'..'Z' | 'a'..'z' | '_')+ ;
Note that it says "after task initalization". What this really implies is that you must use O(1) heap space, and you should get what you need ahead of time.
Having to deal with out of memory conditions 4 seconds after main engine turn-on is not a fun party. Neither is blocking on malloc() so you can prepare your struct course_adjustment to send to another task.
In addition, a lot of the rules are created to make the task of reading source code easier, both for humans and machines. I remember Dr. Gerard Holzmann once half-joked in a meeting that he wanted to disallow any declaration of pointers except at static initialization. I sort of thought he was joking, but then he assured me that it was a serious consideration. He reminded me of the gravity of the situation and explained that $2 billion of public funds were on the line.
Disallowing pointer indirection would make the task of certain automated analysis techniques much, much simpler to perform. Adding a pointer indirection can really conflate matters sometimes.
That said, this system still uses a dynamic memory system - its somewhat non-optional when you use an operating system which has to maintain memory for stacks and its own resources.
These guidelines just forbid you to use the memory allocation routines after the task initialization phase to make memory allocation and usage completely predictable.
The other limitation in a lot of these systems is the underlying virtual memory system and some times there isn't one. Memory fragmentation issues are a huge problem when you have a couple kB to a few MB of physical RAM and a limited VM subsystem.
Also, C is not the only horse in town. In fact, I hear that Ada is actually pretty popular in space-flight. Other missions have successfully leveraged FORTH even. Using a Lisp read-eval print loop from many millions of miles away once saved the Voyager mission.
One really interesting point of view on this topic is Ron Garret's essay called "Lisping at JPL," available at http://www.flownet.com/gat/jpl-lisp.html .
Among the big players in spacecraft flight software, there seems to be a divergent east-coast/west-coast preference for C and C++, respectively. In my estimation, this is the reason: there is a wide variety of target hardware and OSes and the need for FSW to be reused across all of them (embedded linux, VxWorks, QNX on PowerPC, SPARC, intel architectures). In terms of development environments and compiler toolchains, only ISO C (and to a slightly lesser degree C++) is supported by all of them.
Edit: Various instruments on spacecraft may be programmed in Forth or other nifty languages, for example, and there's a growing effort to make some of the more "interesting" challenges in spaceflight (autonomy, fault management, guidance-and-control, etc..) to be coded in custom domain-specific languages or other scripting languages like Lua.
This is an overstated feature of C, pointers were already available in systems programming languages before C was created.
No it's not, and no it doesn't.
No dynamic allocation is a pretty standard precaution for safety critical software. It requires careful coding and design, but it eliminates an entire class of runtime errors and makes it relatively easy to put an upper bound on memory usage.
And then in the request path, be miserly about what you are willing to do. "Modern" web servers like nginx seem to be designed this way. Node JS's HTTP parser makes a point of not allocating any memory.
I think you'd be surprised how far this pattern can go.
would probably make it hard to do anything really interesting
The Rockbox music player firmware has the same restriction and it doesn't seem to prevent it from running Doom or decoding FLACs.
Why not: if (!c_assert(p >= 0)) { return ERROR; }
I am not defending the bans, but find it strange that people would be surprised by them.
yes, i think a lot of replies here are speaking past each other, making valid and not actually contradictory points, but not really replying to each other.
if you read the entire doc it is clear that they are pushing c in a very safe, but somewhat unusual direction. there's no dynamic memory for example, so the main reason that most of my c code uses goto - to free memory on failure in a "catch" - is irrelevant.
taken as a whole, it's not what i would call normal c use, and i don't think it's very useful for most other people as a guideline, but it is internally consistent and, for the specific use case, reasonable.
gotos have a pretty important niche in error handling blocks. You can see them all over the linux kernel, arguably one of the biggest (and most successful) C projects out there.
With C I have never used a goto, sure you can compact your code, but is it managable later on by somebody else and by avoiding goto's you also tend (at least I have found it to be so) to get more structured, easier to follow code. Also smaller functions albiet more of them as well I'd say from what I have experienced.
Also remember a goto may be fine for what the program is to do today, but what about down the line and changes. In that as much structure and in that control is the ideal.
Some might say if you want to code goto's then code elsewere in assembler.
It is harsh, but there again so is space (sorry had to say it).
goto has its uses. It shouldn't be used indiscriminately, but (especially since C does not have tail-call optimization) it is the best approach sometimes.
Embedded C compilers typically do not have TCO.
Yes and yes. Never have used or needed a `goto` in C for these. Could you provide an example of how it would be useful?
* Especially on embedded hardware, which may only have a few KB for the stack.
Another use case for gotos is handling cleanup on error, when writing code that has to be fault-tolerant. Here's a rough example, generalized from a VM I'm working on:
typedef struct thing {
int id; /* object ID */
int buf_sz; /* current size of the buffer */
char *buf; /* internal buffer */
foo *f; /* some other thing that needs alloc / init */
} thing;
thing *thing_new(int id, int buffer_size) {
thing *t = malloc(sizeof(*t));
foo *f = NULL;
char *buf = NULL;
if (t == NULL) goto cleanup;
buf = malloc(buffer_size);
if (buf == NULL) goto cleanup;
t->buf = buf;
f = foo_new();
if (f == NULL) goto cleanup;
t->f = f;
return t;
cleanup:
/* Avoid leaking memory if any part failed. */
if (f) foo_free(f);
if (buf) free(buf);
if (t) free(t);
return NULL;
}
I don't tend to use goto by hand much outside of that particular idiom, but it's common enough that it should be recognized, and the equivalent with ifs and multiple returns would be much worse: "if not B, free A; if not C, free A and B; if not D, free A, B, and C; if not E, ...". rules(struct stateful state) {
switch (state->foo) {
case ....
}
}
sensors(struct stateful state) {
state->button1 = ....
}
motors(struct stateful state) {
switch(state->foo) {
case ....
portBar = ....
}
}
main() {
for(;;){
sensors(state)
rules(state)
motors(state)
}
}
Simple state machine, no gotos are needed. Works fine for non-simple ones too. And the first rule of fault tolerant code (particularly for embedded) is never use malloc.I agree that goto is useful for error handling, but not parsers or state machines. Anyway, I'll play your bait-and-switch. Here is error handling without goto.
typedef struct thing {
int id; /* object ID */
int buf_sz; /* current size of the buffer */
char *buf; /* internal buffer */
foo *f; /* some other thing that needs alloc / init */
} thing;
thing *thing_new(int id, int buffer_size) {
do {
thing *t = malloc(sizeof(*t));
foo *f = NULL;
char *buf = NULL;
if (t == NULL) {break;}
buf = malloc(buffer_size);
if (buf == NULL) {break;}
t->buf = buf;
f = foo_new();
if (f == NULL) {break;}
t->f = f;
return t;
} while (0)
// clean up
if (f) {foo_free(f);}
if (buf) {free(buf);}
if (t) {free(t);}
return NULL;
}
Now normally I would never do that. More typically I would actually use the loop. Because I would not be malloc-ing, I would be trying to initialize a piece of hardware over SPI or i2c. The do...while would be replaced with a for(i=maxtries; i; i--) loop. After maxtries, the loop terminates and the peripheral is shut down.Turns out, you actually can! instead of "break toplevel;", you just write "goto toplevel;". There's another minor change, in that you have to put the name at the end of the scope, rather than the beginning of the scope, which is why people tend to name it after the next block (e.g. "goto cleanup;" in this case).
Still waiting for someone to show me a state machine that absolutely needs goto.
There is nothing that absolutely needs goto (including error handling), because Turing completeness does not require goto. so you might wait forever; I'm not sure what it is that you guys are arguing about with respect to state machines.
(of note, I keep waiting for TCO proponents to show me an example in which the guarantee of TCO in scheme makes the world so-much-better. The claim always comes up, and every example I've seen so far requires at most adding two more lines in Python, and no adding of TCO)
Case in point with examples from the NASA JPL coding standards for C:
* no direct or indirect recursion What is it, FORTRAN-77? Some algorithms are way easier to implement recursively whereas the iterative algorithm can be much less straightforward and buggier. Think sorting: it’s easy to prove that the recursion is finite and that the implementation of the algorithm is correct. Do they use sorting in NASA or is it prohibited by this rule?
* no dynamic memory after initialization FORTRAN-77 again! While dynamic memory management can be challenging in real-time systems and the generic malloc/free implementation is not acceptable, it doesn’t mean that statically pre-allocated fixed-size memory is better. It inevitably leads to brittle code ripe with excessive memory use, bugs like static buffer overruns, and sometimes even inability to use dynamic data structures like linked lists. To work around this restriction, a developer can construct a linked list structure in a statically allocated memory, but doing so is essentially equivalent to creating your own dynamic memory manager which is more likely to be poorly implemented than a good dynamic memory manager. Instead of denying the use of dynamic memory they should develop memory managers with acceptable performance characteristics.
* The return value of non-void functions shall be checked or used by each calling function, or explicitly cast to (void) if irrelevant. Given that there are a lot of library functions in C that return some error code rarely useful, this rule leads to code littered with (void) casts: “(void) printf(…)”, “(void) close(…)”, etc. Along with the littering the rule doesn’t make the code any more robust because it encourages to use (void) casts to ignore error codes and therefore error codes will likely be ignored rather than handled correctly.
* All functions of more than 10 lines should have at least one assertion. This leads to littering code with assertions in those functions that don’t necessarily have anything to assert and that are accidentally longer than 10 lines (for example, due to mandatory parameter validation checks. I hope parameter validation checks are not assertions, are they?).
* All #else, #elif and #endif preprocessor directives shall reside in the same file as the #if or #ifdef directive to which they are related. This is just a bizarre rule. What developer puts #ifdef in one file and #endif in another? Unless of course he’s drunk or high but I hope that’s not how NASA develops its software.
* Conversions shall not be performed between a pointer to a function and any type other than an integral type. Wait, pointers to functions should be converted to which integral type? They are a number of integral types: char, short, unsigned long long. Which one do I choose? Why not void* or intptr_t?
* Functions should be no longer than 60 lines of text and define no more than 6 parameters. Finally a good rule. But what does the explanation say? “A function should not be longer than what can be printed on a single sheet of paper in a standard reference format with one line per statement and one line per declaration.” Printed on a sheet of paper? Is this still how code is reviewed in NASA?
And before you say "these coding standards are for a special kind of software that runs on space flight control systems," embedded devices these days are more powerful than desktop computers ten years ago. Embedded sortware grew beyond draconian restrictions a long time ago and it's much closer now to non-embedded software.
Let's not forget that NASA did use Lisp in their systems and they were able to solve pretty difficult problems remotely with help of Lisp REPL (http://www.flownet.com/gat/jpl-lisp.html). Lisp code certainly can't be subject to any of the restrictions from these coding standards, which is another indication of how irrelevant these coding standards are for producing robust software.