Resources about programming practices for writing safety-critical software
github.com
github.com
It seems like all the resources here are concerned with trying to whittle C/C++ into an appropriate choice of tool, rather than choosing a different tool. It seems like a 1980s-1990s mindset.
Units can help verify that a formula within a program is correct (velocity v = 0.5(metric_acceleration)9.8 t;cout<<v.to_metric();) won't compile, for example.
But it won't help with Program 1:
velocity_imperial v = 0.5(metric_acceleration)9.8 t* t;cout<<v;
Program2:
velocity_metric v; cin>>v; BurnFor(doSomeRocketScienceToCalculateEngineBurn(v))
A much more robust method (especially considering that this table was prepared in advance) would have been to have someone else to independently duplicate the work and compare results.
> you don't trade away stability for performance on a safety-critical chip!
In a sane world, no. But we don't live in a sane world. We live in one where (on one system I worked on) it turned out a plane having an overheat (not fire) on one engine would cause the overheat on the other to fail to report to the pilots (corrected or I'd name and shame). And that wasn't just an issue of language, but of sound (or unsound in this case) logic. No one sits down and develops these things correctly. 10k lines of code (at most) on that project and most of it was just cobbled together in an ad hoc fashion.
For how much longer?
The machines running self-driving cars aren't tiny little processors running single threaded code. They're basically full server racks worth of compute with multi-core cpus, gpus, and who knows what else.
The current approach of "use crufty-but-trustworthy hardware and never do anything too complex" doesn't scale to the next generation of "embedded".
One of the major hurdles to real Level 5 systems is proving their correctness.
There's the rub. They don't meet the standards of safety-critical code, but they are safety critical.
I'm not sure how this challenge will be addressed, but I doubt the answer is "write everything in C". That approach works when your code is relatively simple, but doesn't scale when the code is actually extremely complex.
https://hackernoon.com/five-skills-self-driving-companies-ne...
A team of software engineers can then implement that with static analyzers and run time checkers during development to give confidence to an implementation.
Can you prove stability of a deep learning system across all possible operating conditions it may encounter?
I was thinking more of self driving though.
The main problem is that C and C++ won the political war and many devs don't even bother to look elsewhere.
Actually even C++, in spite of its safety improvements over C, has issues to cater to embedded devs.
Thankfully the IoT of shame is already making people aware that another path has to be taken.
There are legitimate gripes about C/C++, especially in a space with hostile actors an unknown inputs, but that example was particularly weak.
"Better" / richer / more refined (aiming/aspiring at least!) type systems than C/++/Java/C#/Go/etc aim to make it ever-more "powerfully convenient, effortless, and free-of-cost to denote eg. such different units as uniquely distinct (yet 'compatible' when explicit conversion is finally expressly (and visibly (and searchably)) called for) types" that all source code cannot possibly mismatch accidentally without the compiler catching it --- however the issue remains that we don't, and possibly can't have type systems that also enforce such styles (rather than just relying on a developer's/team's own discipline & resolve to adhere to such convention even in the face of deadlines & budget/schedule pressures) unless we actually totally forbid primitives like lone (semantics ambiguous) ints, lone (likewise) floats etc. Similar to "bool blindness", there's the general issue of "primitive/scalar-type blindness" always lingering. Probably some PhD candidate will write a Haskell extension for such a scenario some day soon.
Though of course the next thing the deadline-driven developer will do is construct a single "semantic" int type used for all different semantic sorts kinds and types of "ints"......
https://blogs.msdn.microsoft.com/andrewkennedy/
https://channel9.msdn.com/Blogs/Charles/Andrew-Kennedy-F-Uni...
That this is a mistake that is possible to make in any language, but nevertheless this is also a problem that could be caught by a sufficiently good type system, if it was put to sufficiently good use. In fact, part of the justification for allowing user defined literals in C++ was precisely to make it easy and convenient to avoid mistakes like the Mars Climate Orbiter. Bjarne Stroustrup himself used that as an example in at least one talk he gave about C++11.
int temperature_a;
int temperature_b;
or even: Temperature a;
Temperature b;
you would have: Celsius a;
Fahrenheit b;
And an assignment between a and b would be an error, as Celcius and Fahrenheit are different types. Going along this path then, if you can extend the type system enough: Yards width;
Yards length;
Acre size = width * length; /* this would be okay */
Yards width;
Yards length;
Yards height;
Acre money_bin = width * length * height; /* error */
/* as acre is a measure of area, not volume */
Think of the type of calculations that the Unix command 'units' does, but built into a language instead of an application.You could wrap your value in a typed data structure:
enum LengthUnit {
Feet(f64),
Meters(f64),
}
And provide conversion functions between them, then only operate on one type, `LengthUnit::Meters` and throw errors if `LengthUnit::Feet` is passed in.I'm using Rust syntax here because it's fresher on my mind, but you could do the same with Haskell, OCaml, F#, etc. IIRC, OCaml would optimize away the outer structure so you wouldn't have much/if any performance hit. Presumably compilers for the other languages could/would do the same.
EDIT: for formatting and clarity.
So as weird as this may seem to you, that mindset is applicable much more recently than you might expect.
1. There was a spacecraft (MCO) and a module that was sending some data from the Earth.
2. The module was delivered late when MCO was on its way for 4 (!) months already before that staff manually calculated the needed data.
3. Some teams switched into "defensive mode" not willing to communicate and fixing the problem when it was clear.
https://mitpress.mit.edu/books/engineering-safer-world
This list is barely scratching the surface of safety-critical system engineering, but it's a start.
Ideally you should read it before starting your project, since it deal with the product specification/gathering requirements phase, which is your starting point in safety critical systems.
[1] https://betterembsw.blogspot.com.br/2010/05/test-post.html
I guess testing an over-the-update for a car that was build by ann OEM and thousands of suppliers must be quite a task.
If someone dies because of a preventable bug in your software, shouldn't that be considered manslaughter?
Obviously you formed a corporation in order to shield yourself from legal action (among other things). Fine, so you personally don't get charged with manslaughter. But in that case the corporation should be charged, and if convicted should be sentenced to the corporate equivalent of 25 years in jail. That would be a strong enough incentive to care about software. Of course, it never works like that in real life.
Is this as ridiculous as it sounds to me or is my outrage misplaced somehow?
As Hoare so elegantly described at his Turing award speech, regarding Algol compilers, back in 1981.
"Many years later we asked our customers whether they wished us to provide an option to switch off these checks in the interests of efficiency on production runs. Unanimously, they urged us not to--they already knew how frequently subscript errors occur on production runs where failure to detect them could be disastrous. I note with fear and horror that even in 1980, language designers and users have not learned this lesson. In any respectable branch of engineering, failure to observe such elementary precautions would have long been against the law. "
Full speech here:
http://www.labouseur.com/projects/codeReckon/papers/The-Empe...
Also people don't realize, but by using a linter you basically don't write the code in C, but in "safe C". It's like a different language.
However, the JSF project has been reported to have lots of software defects.
I haven't read anything that differentiates between these two possible scenarios:
1) Poor engineering, execution, etc.
2) The bugs expected in this software project. When I think of it this way, I'm amazed it ever was completed (but maybe I'm thinking about it the wrong way):
* Meet the specifications of not only three U.S. military services but also militaries and other entities in multiple national governments (with all the politics, compromise and complexity that involves).
* Invent and implement technologies to provide capabilities so bleeding edge that few people will imagine some of them for years, if not decades. There are no prior designs; nothing like it has ever been done. Part of the point is to exceed competitors' engineering capabilities by as much as possible.
* Integrate these technologies into a massive system of systems, arguably the most complex system in the history of humankind.
* The system is human-rated.
* Performance is the highest priority; there is no making easy compromises of performance for safety: Human lives, the outcomes of battles, the fates of nations, and the course of history may depend on performance.
* Accomplish this in secret, greatly restricting your access to outside resources. Will this work? You can't publish a paper and get feedback, or make a presentation at a conference.
* Accomplish this in coordination with thousands of suppliers in many countries.
* Because it's hardware and very expensive, your ability to iterate is limited. My completely amateur guess based on the above is that it's a massive, decades-long waterfall-style project.
You went a bit overboard there. There's plenty of systems probably more complex that work fine on a daily basis. They were usually designed centrally, though.
Also, can you distinguish between those two scenarios (which was my main point)?
Far as amazing examples, I'd go for the centralized, five-9's systems like NonStop or the decentralized ones such as OpenBSD before the failure we're discussing.
I think there is a miscommunication. OpenBSD is as complex as the F-35? I think the OpenBSD team would be very disappointed if that were true! OTOH, it would make the ~10 minutes it took to install it on the laptop to my right very impressive.
EDIT: What I meant to ask in the GP was, do you see a way to distinguish between whether the F-35 is poorly executed, or whether we're seeing/saw the normal bugs for a project as described above (even admitting hyperbole about complexity, which I'm not sure of, it's still quite a project).
> where the profits of the contractors are assured through corruption
Not a major point, but these issues are too important to let pass these days, IMHO: How much of whose profits are assured through what corruption? I'm sure some of that goes on, as in all large institutions (including the large companies mentioned above), but I'm not ready to assign corruption to all or most defense industry profits with a broad brush.
Also, I rarely hear much attention given to reducing it. For example, the GOP in Congress was working to take procurement authority out of a central DoD office, where it was put to prevent corruption and to put real procurement experts in charge, and into the hands of generals and admirals, people without procurement expertise (commanding fleets and fighting wars is a much different skill set) and with a bad track record regarding corruption in recent history (look up the Fat Leonard scandal, as one example off the top of my head). What I struggle to remember at this moment is whether that change passed; it was a pet project of McCain's.
But don't worry, C++ is a "COTS Industry Standard" so you can bet there were no overruns on the F-35. /s
Yes, including the piece I directly touched. It was deemed to minor to be worth the cost of rewriting.. Hell, when I left the program we still had an arthritic VAX to build on, should we ever need to rebuild the code.
As for your last line, see my comment elsewhere about blaming the carpenters instead of the tools. :-)
http://www.militaryaerospace.com/articles/2013/10/software-c...
Reading the JSF coding standard, it sounds like all libraries must be DO-178B: "All libraries used must be DO-178B level A certifiable or written in house and developed using the same software development processes required for all other safety-critical software. This includes both the run-time library functions as well as the C/C++ standard library functions. [10,11] Note that we expect certifiable versions of the C++ standard libraries to be available at some point in the future. These certifiable libraries would be allowed under this rule. "
http://www.militaryaerospace.com/articles/2013/10/software-c...