Safety-critical realtime with Linux
lwn.net
lwn.net
I wrote a couple articles on the general idea:
The Fukushima and Deepwater Horizon disasters show how other industries fail at applying the concept. Reading the sequence of failures in those just makes me grind my teeth. For example, venting the hydrogen overpressure, good idea. Venting it into an enclosed space with sparking electrical equipment, spectacularly dumb. (And the list goes on and one with both disasters.)
Nobody has taken it to heart and applied it pervasively like the airframe industry.
edit: reference to the article- https://lwn.net/Articles/540368/
I'll show myself out.
I worked on HA (Highly Available) system with Linux inside Xilinx's Vertex Pro PPC. It is redundant system with multiple fault detections and switch over if any subsystem detected failure.
There was one 250 ms hard real time requirements: If I am a slave and don't detect the master 's UDP ping for 250 ms. I will assume the master has failed somehow, I will start action and take over control as master.
The sub-system did trigger from time to time while the master is alive and working perfectly OK.
Eventually I figured out that one of the system API was using > 250 ms time. (Forget which one now, that was > 10 years ago.) I have to profile very carefully and redesign the code to get around that API.
(Btw my favorite acronym is STONITH.)
Its better to have an odd quorum, to break tie-breakers.
Of course, a segmentation fault is usually a symptom of pointer misuse, which means your code is likely to also suffer from corruptions.
If the system gets in a state where it is in a reboot loop, that is obviously bad, but those are more tractable to fix. In any event, I believe that FAA requires semi-formal analysis for software the failure of which would cause loss of life, but I might be wrong.
https://en.wikipedia.org/wiki/DO-178B
https://en.wikipedia.org/wiki/DO-178C
There's an ecosystem of products developed this way, products such as static analyzers to help develop them, and even companies assisting in any part of the lifecycle. Market shifted toward reusable components for things such as RTOS's and middleware. Examples include INTEGRITY-178B, LynxOS-178B, Alt's OpenGL + driver for Radeon, partitioning Ethernet/networking/filesystems, Esterel SCADE, SPARK Ada, and so on. Lots of good stuff came from this.
On a tangent note, it's why I'm for software regulation where the useful subset of assurance methods that make those great become part of the requirements for commercial software. Defects will go down and predictability up across the board. Things will get more expensive with development slowing down a bit. That will probably be a good thing given Worse is Better effect market has been doing.
I went into more detail on regulation prior results and some predicted ones here:
https://www.schneier.com/blog/archives/2016/11/regulation_of...
There were also various mechanisms to prevent the surfaces from going too far (cannot move full travel at 500 mph, it would rip the airplane apart).
This is all very serious stuff, and was worked over by a lot of people imagining every perverse thing that could happen, and going through the list of things that had happened, to ensure it is safe.
The track record of the 757 in service shows how effective this is:
https://en.wikipedia.org/wiki/Category:Accidents_and_inciden...
None of them resulted in recommendations for design changes.
We did a foam board stretch build of Model Airplane couple years ago.
https://www.flitetest.com/articles/ft-guinea-build
After we flew a few times, we added a servo to the cargo bay to be open and drop small parachute from it.
After than, my son kept complained that he would loss control for 1-2 seconds while the plane is in the air from time to time.
After research on the net, eventually I figure out it was the addition servo. Each servo in that air plane use about 1 Amp of current depend on condition. The 5th servo actually drew enough current that causes the remote receiver to reboot while the plane is in flight.
That was a fun bug to figure out.
The RC receiver is actually very good design which can be reboot and be operational in ~1 seconds.
A safety critical system, absolutely must be able to handle seg faults without compromising safety. Any system that requires that a seg fault will not occur is not a safe system.
It would seem to me that a worst case scenario could easily cause slowdowns of many orders of magnitude. You could mitigate some of them by careful manual memory layout and hardware specific tricks like hardwired TLB entries, but still be left with a lot of uncovered stuff.
I would think with that strategy, you'd be servicing "slice-interrupts" more than anything else.
I'll stick with writing ISR's that are only a few commands and do my work outside the ISR's, as standard in industry now.
Now, on a related topic, if they're discussing getting the system to a 10us timing, that would be useful in using Linux as a 3d printer controller directly, rather than an Atmel or STM32 chip. But those requirements of what needs turned off seems pretty onerous, unfortunately.
Doing actual work in an interrupt routine is a non-starter in a real time environment.
Presumably you can sort of brute force it if you have defense level budgets, but it's a seriously bad situation.
[1] http://epickrram.blogspot.co.uk/2015/09/reducing-system-jitt...
I wasn't as aware of the issue of safety-critical systems as I should have been until I was inside a couple industrial companies where PLC's were everywhere (for this very reason). The thing that interests me now about this is how hard I see netconnected PLC's pushing into industrial applications, mostly because everyone in industry is on the edge of their seat for IOT to hit so they can use and abuse the data (instead of waiting for service call to pull data like they used to, why not just use an LTE-modem PLC, for example?) Do you see where I am going with this? Safety-critical industrial applications <sil4 are increasingly more vulnerable, and it's not from lack of realtime response to stimuli. In the end, using linux in realtime just seems to exacerbate this particular angle on the issue that I see. It does make me wonder about the implications of microkernel design vs monolithic in such applications though.
http://www.nfpa.org/codes-and-standards/all-codes-and-standa...
https://en.wikipedia.org/wiki/IEC_61508
https://webstore.ansi.org/RecordDetail.aspx?sku=ANSI%2fRIA+R...
https://www.iso.org/standard/69883.html
https://webstore.iec.ch/publication/22797
https://en.wikipedia.org/wiki/Comparison_of_real-time_operat...
It feels like LWM has bad access control, and someone abused that to post an article that shouldn't be free.
In this case, the link was posted by Jon Corbet, who is LWN's main editor and developer.