Toyota's firmware: Bad design and its consequences
edn.com
edn.com
https://news.ycombinator.com/item?id=6615457
engineers are often not aware of basic principles of fail safe design. I mentioned Toyota, and this article confirms it.
Not mentioned in this article is the most basic fail safety method of all - a mechanical override that can be activated by the driver. This is as simple as a button that physically removes power from the ignition system so that the engine cannot continue running.
I don't mean a button that sends a command to the computer to shut down. I mean it physically disconnects power to the ignition. Just like the big red STOP button you'll find on every table saw, drill press, etc.
Back when I worked on critical flight systems for Boeing, the pilot had the option of, via flipping circuit breakers, physically removing power from computers that had been possessed by skynet and were operating perversely.
This is well known in airframe design. As previously, I've recommended that people who write safety critical software, where people will die if it malfunctions, might spend a few dollars to hire an aerospace engineer to review their design and coach their engineers on how to do fail safe systems properly.
Yes, it's unnerving.
Once the clutch plates are no longer in contact the engine can do whatever it wants but none of that makes it to the wheels. And then you can get the car out of gear to coast, turn the whole car off and then back on.
There's no substitute to having a human being in the loop to execute high-level "executive" functions, especially when things aren't done to the very high standards of aerospace. You know, like cars.
There exist "manual transmission" that lack an operator controlled clutch. DCTs are very much manual transmissions, but many of them rely on a TCM to disengage the clutches.
There are all sorts of advanced DSG / dual-clutched transmissions in many newer cars, but even some older cars like the MR2 and Smart cars have a http://en.wikipedia.org/wiki/Electrohydraulic_manual_transmi... : quite literally a manual transmission where the clutch has been actuated electronically.
If you have computer-assisted hill starts, collision avoidance, or computer assisted braking [through ABS, certain traction & stability control systems, etc] your computer has control of your brakes.
If you have range-assisted cruise control, or early collision alert, your computer likely has total control over your throttle.
If you have parking assist, lane departure warning, lane following assist, or electric power steering etc: your computer has control over your steering.
If your car has a DCT: your computer has control over your shifting _as well as both clutches._ -- Meaning you have _no mechanical interface_ to disengage the motor from the transmission. This is similar to your Prius example: the shifter is not mechanical.
On many new cars: there's no ignition key to remove. You likely have a smart "keyfob" that simply needs to be within X-feet of the car, and then you have a push button ignition.
I'm sure there's some override for the button [push and hold for three seconds], but it's still going to go through some electronics to figure that out.
---
The only bit that scares me is that all these systems potentially share the CAN bus with the horrendous "Infotainment" systems that every manufacturer loves to install. shudders
I have to admit though I'm surprised that some of those accessories even require a bus as they're typically operated by simple switches.
I guess door locks, wipers, etc. make some sense: with the prevalence of central locking, as well as wipers that sense rain and things like that.
Windows are a bit odd, as they just need a very simple switch. I suppose having all the accessories on a common bus must simplify the wiring harness though.
---
I'm curious how TESLA integrates their in-dash system with the rest of the car.
I thought their in-dash console could control some rather safety-critical "preferences" of the car ... for instance I thought you could adjust the level of regenerative braking.
(Other manufacturers are also offering "sport modes", etc -- though they're not always controllable through the dash.)
They certainly don't require a bus in principle. However, if you want to understand both the business case and some very good engineering reasoning behind the use of data buses to control simple automobile accessories you should have a look at the wiring harness for an early 90's Mercedes or similar luxury car. Demand for fancy features drove the number of wires running to and fro to an unmanageable level. Having a mile of wires in a car is expensive for many reasons I'm sure you can imagine, as well as each connection point becoming an opportunity for something to go wrong.
I'm also curious how electric cars in general will manage safety systems going forward, especially in light of this Toyota court decision.
The explanation always seemed obvious to me: Non-technical driver experiences UA. In a panic shifts into neutral. Hears engine scream. Thinks, "The car is still going!" and doesn't realize that despite the sound, the car is actually slowing down.
I know this to be a common misconception in my experience. People commonly misinterpret high revs with "going fast".
But I had not considered that this correlation is only intuitive to me because I own a car with a manual transmission.
But then, people can react irrationally (or not at all) in a full on panic.
Also pilots get proper training to handle their vehicles, car drivers not so much.
Fair enough, but as far as I know not a valid point in this specific case.
>>Braking is a software function
Yes, but a completely separated system. The odds of both software systems failing simultaneously is getting in the hash collisions domain...
>>most Americans drive automatic (software) transmissions.
Again a seperate system, and you still have a neutral position.
Are you sure? It seems sensible to me that accelerator inputs would factor heavily into the braking system, so it seems very sensible to me that an unexpected condition in one could translate over to the other - as the article noted the only way to undo one unexpected acceleration condition was to completely remove your foot from the brake pedal. Sounds like cross-over to me...
>you still have a neutral position.
... which is likely just a software input to the transmission computer.
In all likelihood the only non-electronic failsafe is the handbrake, which I still think is a direct mechanical connection in almost all cars.
Not on my Nissan Leaf. The brake lever is a switch that turns what I assume (based on the noise) is a small motor to engage the rear brakes. As far as I can tell, everything is electronically controlled. Brakes, accelerator, "ignition", "transmission" (both in quotes because the Leaf really has neither), parking brake. If there's a firmware failure, there's not a mechanically-operated fail safe to save me.
It isn't, really, in a lot of newer cars. You have a brake override system [1] that can reduce the power output of the engine by various means.
[1]: http://auto.howstuffworks.com/car-driving-safety/safety-regu...
Of course, if the computer is busted, I'm assuming the long button press will be sent to /dev/null.
That is what sucks about cars with power steering, they don't get any of the benefit a car built with no power steering does.
Getting that boat of a minivan around the next corner in traffic was both entertaining, dangerous, and probably the best upper-body workout I got that year.
The point is, most of the options you've listed there really may be computer-mediated in modern cars. (And yes, I've heard that there's a strong correlation between unintended acceleration and older drivers, and that a lot of those cases really are driver error. But I don't think you're making that case here.)
All of these functions are electronic on the Prius (and indeed on every hybrid car that I know of that's on the roads). The balance between regenerative and regular friction braking in particular requires quite a bit of computer code and calibration to get right.
The drivetrain spins the electromechanical motor at all times, adding drag. The drag is not just from the added mass of rotation, but also a dynamic resistance caused by electrical properties of the motor being varied in different ways so that the computer can achieve either regeneration (by temporarily changing modes to allow the motor and circuitry to act as a generator, usually during a coast downhill or to a stop), or additional braking (by electrically braking the motor, using the energy stored in the batteries, to add further resistance to the drive train at the cost of heat generation and range reduction).
If a check engine light that has to do with the electromechanical subgroups of your prius comes on (indicating a fault) those systems are disabled, meaning that the car is more or less non-hybrid during those times. Braking will feel stiff, and the car sluggish, but it is by no means dangerous to drive (unless you consider the new learning curve for the cars' performance profile to be dangerous, which it is.)
Also : Your emergency brake is indeed fully mechanical, but on newer models they may be released electromechanically via a command, i'm unsure. I haven't worked on one since the second generation.
p.s. you forgot a sub group. Your steering rack is also electromechanical. One of the first of its' kind in production. Meaning, if you ever experienced a total blackout, your steering would, too, become much more resistant. This , however, isn't considered to be a safety hazard, because at speed the steering rack does little to assist the driver. the forward momentum takes care of that. The steering assistance is mostly there for parking lot situations.
(source : I was at one time a toyota technician, and my back still remembers the recall on first generation prius battery packs, they weighed 124lbs and were way awkward to remove.)
Vehicle tests confirmed that one particular dead task would result in loss of throttle control, and that the driver might have to fully remove their foot from the brake during an unintended acceleration event before being able to end the unwanted acceleration.
Every one of those approaches you suggested are, in many modern cars, fully software driven. And the article even shows an example of how a bug in the software can only be resolved through the exact opposite of what a rational person would do in a crisis.
I think the only actual mechanical failsafe left is the handbrake. Please tell me that's still sacred...
So you're probably right, it's a stretch to call the handbrake something useful in emergencies when in reality it probably wouldn't perform that function.
source: I have tested this at ~30-40 mph and nothing about that experience leads me to believe that it would be safer if I had been going faster, and at full throttle.
I don't drive cars, so i have no idea if that would conflict with normal uses of the handbrake. Perhaps it could have a position beyond the normal brake-engaging position that did this? So that if someone panics and yanks on it as hard as they can, they get the result they probably want.
Yuck!
>Luckily most the features you listed are usually implemented in separate ECUs.
What concerns me would be how the systems handle unexpected inputs.
In the article it notes that the only way to end one unexpected acceleration event was to stop using the brakes. I'm not sure if the vehicle in question has separate controllers, but if it doesn't that's a real concern that unexpected input from one tickles a bug in another.
The gearshift in the Prius is totally electronic, and it does not allow you to switch into neutral if you're traveling above a certain speed.
I don't think this is correct (at least for 1997 - 2009 models). Could you offer a citation?
Edit: this made me curious, so I did some cursory research and found this:
http://answers.yahoo.com/question/index?qid=20100204083409AA...
According to one commenter, to shift into neutral when driving you can do one of the following:
1. Press the park button
2. Shift into reverse
3. Hold the shifter in the neutral position for 3 seconds
A video posted by a different commenter shows the driver holding the shifter in the neutral position for not quite 3 seconds, but still longer than is required to shift the car into other gears (and if my memory is serving me right, longer than is required to shift into neutral when stopped).
In any case, the most obvious way to shift the car into neutral did not work for us, and it's unlikely a panicked driver would think to try any of the methods listed above.
Yes, better driver training could have made some of these faults less serious. Either by fully braking properly, switching into neutral, or other techniques. That doesn't excuse the faults though.
http://answers.yahoo.com/question/index?qid=20091121085512AA...
http://www.cliosport.net/forum/showthread.php?623742-06-Clio...!
Aircraft investigation is really good at overcoming hindsight bias and looking at human factors in a more objective way. What seems absolutely logical for you to type, having read about these incidents before, might not be as obvious to a driver who hasn't read about unwanted acceleration but is suddenly experiencing it.
Software that ensures safety like this really ought to be mandated to be open-source.
That part I assume is on purpose, because even a computerized system could have "stop immediately" as the default policy when the emergency lever is pulled. Would be interesting to read the analysis that led to the decision. My guess is that it's because on-train issues are statistically the most likely emergency situation a passenger would signal (heart attacks, fights, etc.), in which case continuing to the next station (typically 30-90 seconds) where emergency staff can meet the train and access it, rather than stopping in the middle of a subway tunnel or elevated rail segment, is the most sensible policy.
http://www.huffingtonpost.com/2013/10/04/cta-blue-line-crash...
The tripcocks, I hope, are still connected directly to the brakes.
I don't know the failsafe design of the 787, but I have faith that Boeing, the FAA, and the aerospace engineers know what they're doing with failsafe design.
But you're right in that if I was actually working on the 787, I would have no such faith, and would verify the designs I was responsible for.
In 1989, a DC-10 operating United flight 232 suffered an uncontained engine failure which damaged the tail and disabled flight controls, resulting in 111 deaths (and could have been more, but in a freak of chance, a DC-10 flight instructor was on board and was able to assist the crew in landing, which may have made the difference for the other 185 people on board). In 2010, an Airbus A380 operating Qantas flight 32 suffered an uncontained engine failure which damaged a wing, disabled a hydraulic system and braking systems, and disabled some flight controls while starting a fire. It resulted in... zero deaths.
In the early 1990s, Boeing 737-200 and 737-300 aircraft had a variety of uncommanded rudder movement issues, resulting in at least 157 deaths. In 2012-2013, the 787's battery fires resulted in... zero deaths.
In other words, the main difference between "then" and "now" is exactly the opposite of the usual arguments against modern aircraft development: more recent aircraft, when they have serious issues, result in fewer injuries and deaths than older aircraft when they experienced serious issues.
This track record of improvement gives justifiable faith that modern aircraft development is safer.
There was an in-flight fire, on an ANA 787. It made an emergency landing and the plane was evacuated. No lives were lost, and the airframe was not lost.
For those that like finding out, "What's the formal name for that?"... it's called a kill switch:
https://en.wikipedia.org/wiki/Kill_switch
For further reading, see the related "Dead man's switch":
I would suggest they should not be called "engineers" then. And in many countries, they're not. Part of the problem is that the tech community includes a lot of different people. Some are programmers, some are program managers, some went to engineering school, some are licensed engineers (in some other discipline). In the US, these are all commonly called engineers. Sadly, I think a lot of web programmers just don't know the true scope of the software industry and its practices.
If you want to design/build a bridge, you need a state license and insurance. The software industry isn't regulated like that. Anyone can design the software that controls a car. That's probably OK since web apps are non-critical systems. But I can't help but wonder if net security wouldn't be better if more programmers had better training in recognizing and improving the total impact of a system.
It is only the reputation of the company and potential damages in a lawsuit such as this one that put pressure on the car manufacturer and web-app startup to test their code in depth. Actually, I do wonder how much the US auto safety regulations are involved with firmware--or do they just test the macro behavior of the car?
Wat?
The sample tests I looked at had none. The GRE exams I took had none. The engineering courses I took never mentioned it. I don't recall ever seeing an engineering textbook discussing it. I've never seen it brought up in engineering forums or discussions about engineering disasters.
And, I see little evidence of awareness of it outside of aerospace - Toyota, Fukushima, and Deep Water Horizon being standout examples of such lack. You can throw in New Orleans where hospitals (and everyone else but one building) put their emergency generators in the basement. And in a NYC phone company substation was entirely destroyed because a vital oil pump was in the basement that got flooded during Sandy.
Failsafe design flaws are not uncovered by testing code.
Failsafe systems are designed not by "the code works therefore it is safe", but by "assume the code FAILS". Regardless of how much testing is done, you still ASSUME IT FAILS AND ACTS PERVERSELY. Then what?
(Note that acting perversely is hardly farfetched in these days of ubiquitous hacking.)
I've really been enjoying the posts you shared relating to your time in aerospace. I think there are a lot of lessons the entire software industry should learn from...
The computer control systems were dual, meaning two independent computer boards. The boards were designed independently, had different CPU architectures on board, were programmed in different languages, were developed by different teams, the algorithms used were different, and a third group would check that there was no inadvertent similarity.
An electronic comparator compared the results of the boards, and if they differed, automatically locked out both and alerted the pilot. And oh yea, there were dual comparators, and either one could lock them out.
This was pretty much standard practice at the time.
Note the complete lack of "we can write software that won't fail!" nonsense. This attitude permeates everything in airframe design, which is why air travel is so incredibly safe despite its inherent danger.
Thank you very much for taking the time to share and answer my question!
I think the guiding principle of engineering regulation would lead one to believe that software controls for a car should be covered by licensing, but that this has not occurred in practice due to the regulations not keeping up with technological change.
"Safe Systems from Unreliable Parts"
http://www.drdobbs.com/architecture-and-design/safe-systems-...
"Designing Safe Software Systems"
http://www.drdobbs.com/architecture-and-design/designing-saf...
Still though, sort of horrifying. (Hopefully I hit some sort of safety or fallback mode that can only occur on boot!)
It is with a mixture of fear and amusement that I observe the workings of my car's firmware. (I can hear the interrupts fire in my stereo system when I change the volume through the steering wheel control and the music skips, some buffer wasn't full! I almost have a 100% repro worked out. :) )
Edit:
Having worked in Firmware for some time now, I can confirm that the skill level of many embedded developers is not what I'd call stupendous. Now of course most people are, on average, average, but embedded seems like a special case.
The thing is, embedded systems have exploded in complexity in the last few years. No longer are software projects worked on end to end by just a handful of engineers, rather embedded engineering teams are being forced to learn the lessons about properly scaling up software engineering that developers in other areas learned long ago. A project with 16KB of space for code could be written by 2 or 3 developers sitting next to each other, and it was reasonable to keep the entire program state in one's head.
Now days? You can get Cortex M4 boards that look a damn lot like an actual computer. Sure you don't have much RAM, but the code complexity is way up there. You aren't just talking over an I2C bus to a couple of peripherals anymore!
On top of this, more and more features are being shoved into cars through the use of software. I talked to a developer of wiring harnesses for one of the major auto manufacturers, he described to me exactly how the auto companies see software as "the easy part" of things, which means they get the short shaft in terms of resources (test, dev, time, budget, etc), but are expected to bear most of the load of new feature work. (After all, it is so cheap to do it in software!)
FWIW, this developer said he has gone back to purchasing older model cars, he won't buy any of the cars running his own team's firmware.
My outsider's impression is similar to your insiders - the complexity of external systems is catching up with embedded systems, but perhaps not the practices and talent.
I don't lease/purchase Mercedes vehicles anymore.
I think that may have been an overzealous algorithm giving the car a blip of throttle to stop it from stalling when it is idling.
An engine still needs fuel and air to run when even when there is no load on it and this is controlled by a computer. This computer targets some rpm for idling. If the rpm drops substantially low, due to for example a 'mechanical' glitch like dodgy fuel, then the computer may over compensate with too much throttle, causing the revving you experienced.
:)
a) Brake pedal is fully pressed b) Vehicle speed is at zero
If these conditions are met, and you're attempting to prevent a stall, you may want to consider pushing the transmission into neutral to prevent "unintended acceleration". Transmission is drive-by-wire, so could be done by issuing a command over CANBUS.
These cars run in the millions in production numbers and are driven everyday, all day, by various types of drivers in various conditions. Where's the big firmware mess? Ignoring the limited case of this Prius issue, which seems to have a lot to do with floormats being stuck to the accelerator, I'm just not seeing it.
I think there's a problem of visibility here. If you've worked in, say, fast food, you might not want to ever eat it. Software is the same thing. You get to see how the sausage is made. That doesn't mean that the sausage is unsafe. Or that older methods, you didn't witness or were part of, were better.
You drive the rear wheel drive V8 that weighs 2x my car's weight and I'll drive mine with front drive, ABS, traction control, and incredible MPG. There are a lot of reasons to own a classic car, but safety and efficiency aren't those reasons.
That a piece of software is widespread and works good most of the time simply can't be used to say whether the software is "safe". To say anything about that, you have to look at the process used to develop the software, and see whether that process has a good success rate of predicting and uncovering software bugs.
If Toyotas code is as bad as this article makes it out to be and the certification standards are so poor, I won't be surprised if we see more cases like this in the future, although they will still be rare.
Actually, a lot of newer cars weigh more, not the other way around. I believe that this is mostly due to additional safety devices. For example a 1967 Mustang weighs 2500-3000 lbs[1], while a 2014 Mustang is 3500-3700 lbs[2]. A 2014 Toyota Prius is 3000 lbs[3], while a 2014 Corolla weights 2800 lbs[4]. A 1980 Corolla weighed 2000 lbs[5]. As you see, that rear wheel drive V8 weighs the same as the Prius. I think the improvements in MPG mostly come from improvements in engine efficiencies, not weight savings. I don't know what kind of car you drive, but I'd be surprised if it only weighs 1500 lbs.
That being said, I mostly agree with your other points.
[1] http://auto.howstuffworks.com/1967-1968-ford-mustang-specifi...
[2] http://www.motortrend.com/cars/2014/ford/mustang/specificati...
[3] http://www.toyota.com/prius/features.html#!/weights_capaciti...
[4] http://www.toyota.com/corolla/features.html#!/weights_capaci...
[5] http://www.ultimatespecs.com/car-specs/Toyota/5260/Toyota-Co...
That said, its also a little unfair to compare a 4 door sedan to a 2 door sports car. Modern cars are still heavier, but the difference is a bit more sane.
I think my dad's Cadillac when I was growing up was over 5,000 lbs. The 70s and 80s certainly had heavy cars, but heavy and death-traps and super shitty mpg.
I have no idea what actually went wrong, it's been working swell for 6 years after that now. My next car is going to be manual though, I no longer fully trust that sort of thing.
Was this outside of warranty? Because I've paid _hundreds_ for transmission control modules of all sorts: Dodge, FORD, Honda.
The "chips" they would be referencing are likely surface-mount not socket-mount on anything much newer than 1995. (I think my 1993 EEC-IV has a socket mount for the main processor but I haven't had it apart in ages.)
There is no way a brand new TCM is $15 before insurance; and I just don't see what "chip" they would be replacing that costs $15.
Gotta love Honda transmissions though; they've has had their fair share of rather interesting transmission reliability issues. My personal favorite comes from the Acura TL.[0]
[0]: http://en.wikipedia.org/wiki/Acura_TL#2000
"...as the third clutch pack wore, particles blocked off oil passages and prevented the transmission from shifting or holding gears normally. The transmission would slip, fail to shift, or suddenly downshift and make the car come to a screeching halt, even at freeway speeds..."
I honestly don't know, the car was a 2002, this was in 2007. I bought it used a year before (from by father). IIRC it had 115k miles on it when I bought it. I've been able to get another 60k on it since then.
I've never quite bought the "chip fried" line. My suspicion at the time was that this was a frequent issue that they were trying to keep somewhat under wraps (hence the near-free repair), but now I am wondering if there was a firmware issue that they fixed by reflashing something.
That or they couldn't find the issue and it resolved itself when they reset something...
I wish I had pressed for more details, but I was mostly just happy the transmission hadn't killed itself.
I know :-(
I have worked in firmware for almost my entire software career and I agree completely. The main problem is that most embedded software programmers started out as EE's (like me) and never learned how to architect or build complex software systems. The ones who took the time to learn CS & SE produce noticeably better code and with their knowledge of electronics can do amazing things, but the rest... well, I still remember trying to explain to an EE developer during a code review why "#define ZERO 0" was still a Magic Number!
What the embedded industry needs is more systems engineering. Not taking an EE and saying, "Poof, you're a systems engineer now!", but actual, interdisciplinary engineers who understand systems development concepts, requirements analysis, safety analysis, and everything else that goes along with it. Its primarily an aerospace thing right now, but it really should expand to the broader field.
[1] In fact, at the moment I am a pure software engineer.
Personally I'd be spun out in a ditch about 5 times a day without traction control, I tried turning traction control off in my car once, and it turns out the way I drive is basically 100% dependent upon traction control functioning!
I feel like something has to give soon. I don't think we can sustain the exploding complexity of software systems too much longer. The cognitive load will be too much to bear to develop an "average" software system soon, unless we come up with some really fundamentally different concepts to manage complexity.
Maybe "systems" are the problem. Systems are almost by definition infinitely complex. Maybe there's a better model that reduces the "system" effect? Could it be a functional approach?
Most of the crappy projects I've worked on weren't crappy because the developers didn't know what they were doing. They were crappy because they were old code bases that were recycled over and over and over again through small hardware iterations, and no team working on a single iteration can ever get permission from management to do the necessary maintenance (refactor, redesign, whatever) -- they are just required to get their little port finished fast. Everybody makes the smartest hacks they can on top of the crap they were given, and each step of the way the system gets worse.
Most companies don't start from scratch -- which is almost impossible anyway, given the low quality of documentation from chip manufacturers -- but start from whatever example code base or framework the manufacturer provides. These are invariably bad, and were developed like I mentioned above (hacks on hacks with each hardware iteration)... and then the hacks for your specific project start.
Then the actual development cycle is pretty much dictated to be the 'waterfall' method since it's tough to be 'agile' with hardware. Proper debugging tools, at the level you would get with desktop software development, are either unavailable or cost more than the company is willing to invest. Proper code analysis tools are the same. Continuous deployment and automated testing are practically impossible unless you have tens to hundreds of millions of dollars to spend on infrastructure.
And there's never enough time for any of it.
And there's never enough documentation for any of it.
At this stage, embedded has reached the complexity of enterprise software running on desktops/servers, but without the process and tools needed to make that actually work.
(I am not saying hes verdict is biased, I am just stating he has cause to find Toyota guilty).
Furthermore was Toyota's code really of such poor quality relative to the rest of the industry, and relative to the economic realities of the market. I mean it's all good and well to demand aerospace and medical device quality code, but would the average consumer be willing to pay $1000 per LOC for his speed control in his car? I very much doubt it.
But your point is taken, it might well be feasible to demand aerospace quality code from the automakers.
No amount of quality can guarantee the computer cannot fail.
If you're implying that Michael Barr somehow stretched the truth of what he found, and did so under a legal oath, then I think you're implying that he perjured himself, which is a pretty serious allegation.
That doesn't mean it's the truth though. I too read through that looking for something more damning than I found:
The memory wasn't ECC and the code didn't do anything to mitigate that risk, but that doesn't prove that a memory fault occured or even tell us what the likelihood is. Toyota screwed up the stack depth analysis. But AFAICT no stack overflow condition was found. Apparently some other stuff was found (probably with a static analysis tool) on which the article doesn't elaborate.
These are bugs, and certainly don't make me feel good about the system. But they're not a specific finding of fault either.
It's easy to overreact after a tragic incident (and I have full sympathy with the families), and to demand more stringent certification, but it comes at a cost as well, the price for the software goes up.
NASA: "The NESC team examined the software code (more than 280,000 lines) for paths that might initiate such a UA (unintended acceleration), but none were identified."
Everyone has come across code which no-one in the company wants to touch because it's too complex. However, that same code may have been in production for decades, proving fit for purpose by example.
That said, I'm not sure I'm entirely against the idea of companies being sued for failing to write sufficient quality code. Could do wonders for the industry!
Until it isn't anymore because one of the critical but undocumented (and possibly even unintended) assumptions no longer applies. See also: Ariane 5. I can easily see that happening in this case - all they'd have to do is shrink the available stack space to free up RAM elsewhere and suddenly critical variables would intermittently get corrupted, possibly only under circumstances that didn't happen in testing.
Unintentional RTOS task shutdown was heavily investigated as a potential source of the UA. As single bits in memory control each task, corruption due to HW or SW faults will suspend needed tasks or start unwanted ones. Vehicle tests confirmed that one particular dead task would result in loss of throttle control, and that the driver might have to fully remove their foot from the brake during an unintended acceleration event before being able to end the unwanted acceleration.
EE's often have little software engineering taught as mandatory subjects in college, and often only get to learn good SW engineering practices on the job, or via milspec/aerospace certification, if ever.
No disrespect is intended, it's just my experience after working with a lot of otherwise very talented EEs.
Usually a computer engineer or software person works with the EE who developed the board level hardware + ASIC engineers.
Of course, there exist EEs who are talented at writing software, but I wonder how often EEs are tasked with writing production code.
http://en.wikipedia.org/wiki/Conway's_law
For example, if the firmware is controlling a device that does not typically include higher-level computer systems. For example, an ECU originally designed for a non-hybrid vehicle, there will likely be no CE/CS people involved.
And, no, most weren't "EEs who are talented at writing software" :)
In laymen's terms, we call that "burning" :-)
But formal methods, which typically are based on computer-checked proofs, can help you to eliminate certain possibilities. They are severely underestimated and underused because of our dependence on C.
The formal methods do not 'fail' as such. They just fail to prove anything boyond the proven property. Such properties can very well (and often do) include failure (of whatever kind) of parts in the system.
Apart from AI, which as an approach to embedded systems is almost by definition the opposite, I only see one way forward from the mess we're in now: formal methods.
Well, okay, so the article kindly submitted here takes the position of personal injury attorneys who just won a trial before an Oklahoma jury. And perhaps that is the correct factual position about what Toyota did and what Toyota should have done instead. (Disclosure: I am a lawyer, so I had law school courses that trained me to think on both sides of issues that go to litigation.) When "unintended acceleration" cases were first mentioned in the news media, including one case that occurred here in Minnesota, I was very wary of buying Toyota cars, and bought other brands instead. But as our previous cars wore out, we bought Toyotas, and Toyota vehicles are what we drive for all our driving now. I notice that both cars we bought have very clear warnings near the floor mats about attaching those securely and not using any floor mat that isn't attached securely. Toyota, from this point of view, seems to be acknowledging that floor mats used to get jammed up against accelerator pedals in a way that made cars hard to control. We have not had any problems with our vehicles. The news stories about unintended acceleration in Toyota cars seem to have diminished. Perhaps whatever was bad about the former designs has been fixed.
Process issues can be solved, given motivated management and diligent engineers. Whether the effort to solve them was done and how that effort goes becomes a cultural thing, and I can't speak to the Toyota SW engineering culture(s?).
I could be convinced that this court ruling is erroneous, and that the unintended acceleration issues can be entirely accounted for by floor mats and driver error. But in this case I think we should be thankful for bad floor mats and driver error, as they've brought to light very fundamental flaws in Toyota's firmware engineering processes.
If there had never been a single unintended acceleration in a toyota vehicle it would not have been through robust engineering but instead through luck. And we need our vehicles to be safe by design, not through happenstance.
I especially appreciate your comment because my childhood best friend (an electrical engineer who designs safety-critical systems) thinks this way. He started out in avionics, and was the lead designer for the avionics system for a commercial airliner that so far has a very good safety record indeed, and then he moved over to the medical device industry. In his work, "zero defects" is the only standard, and fundamental understanding of how a system works, from the level of subatomic physics on up, is his approach to design with no hidden flaws. That approach is not easy, but he thinks that is the appropriate approach when human lives are at stake.
It's telling that even at a big company like Toyota which is fairly risk averse and is well known for its commitment to quality and safety they are still capable of churning out pretty crappy software that is hugely important. Software dev is still a pretty hard problem overall, we live in an era where there have been lots of successes, but failure is still common and the consequences of failure can sometimes be severe.
I'm not saying Toyota should be allowed to provide this shoddy piece of software in a critical subsystem, but I very much think a) other vendors software will be just as crappy and b) this feels like the court longing for reasons to fault Toyota on something that was still very likely user error, not software misbehaving.
Besides, why would bugs cause more problems when the driver is elderly (http://www.forbes.com/2010/03/26/toyota-acceleration-elderly...)?
Toyota claimed the 2005 Camry's main CPU had error detecting and correcting (EDAC) RAM. It didn't.
Unintentional RTOS task shutdown was heavily investigated as a potential source of the UA. As single bits in memory control each task, corruption due to HW or SW faults will suspend needed tasks or start unwanted ones. Vehicle tests confirmed that one particular dead task would result in loss of throttle control, and that the driver might have to fully remove their foot from the brake during an unintended acceleration event before being able to end the unwanted acceleration.
In other words, they used non-error-correcting memory, and the investigation found a code path that would lead to the observed behaviour if a single bit flips.
"Might" is a weasel word you use when you don't have proof. Are they saying that the the system is non-deterministic? Seriously, you can say it "might" cause the problem, or you can run tests that cause the problem. Even if it is non-deterministic, you could run 1 gazillion test and get a percentage. Then you could figure out, based on the amount those system get used, how often you'd expect acceleration to be uncontrolled.
And it still doesn't explain why it mostly happens to the elderly.
He found massive failures in all of the safety systems, and successfully demonstrated that a single bit flip could cause the task responsible for controlling the gas/fuel mixture to stop running, preventing the driver from decelerating the car. The safety mechanisms in the car would entirely fail to catch this, and at this point Toyota wasn't using error-correcting RAM, so it's not entirely implausible.
He found many possible buffer overflows, stack overflows, race conditions, and unsafe casts that could lead to memory corruption or logic errors. He went on at length about bigger-picture design flaws in the way that their failsafes were implemented, rendering them often useless. They explicitly ignored error codes from the operating system which indicated that things were going wrong, as well as from their own code which was warning them that the CPU was overburdened and necessary tasks my not have been completed.
He testifies that Toyota has no real bug tracking system, no consistent code review, and had countless violatings of both their own safe coding standards, and other standards which they had had contributed to.
The corresponding Reddit discussion at http://www.reddit.com/r/programming/comments/1pgyaa/ may also be of interest.
absence of any bug-tracking system
wow. and this is toyota.
[edit: http://en.wikipedia.org/wiki/The_Toyota_Way http://en.wikipedia.org/wiki/Toyota_Production_System ]
Toyota's rep has been well earned, there's plenty of literature in operations about this, as well as popular trade coverage (i.e., http://www.nytimes.com/2013/10/29/automobiles/japanese-autos...)
Their software process wasn't given the same attention, and they've been burned by it. Embedded systems are really tricky, and a good process (design through end-of-life) is key.
This is a perfect example of software development done without management maturity, process maturity, and with inferior technical tools. That large scale systems for life-critical services are written in ASM/C is horrifying. That management did not enforce certification compliance is horrifying. That the correctness process did not account for tin whiskers or ECC memory is horrifying. That the engineers violated MISRA (which evidently they attempted to adhere to) is less horrifying, but still bad.
Let's just say that this is an incredibly readable discussion on how to do safety critical software wrong in many, many, many ways. Everything from using binary blobs to using gratuitous amounts of globals (> 10,000 ?!?!) to not having an issue tracker(!!?!?!).
At any rate, this document counterindicates buying a 2005 Camry.
I was in stop-and-go traffic and went from accelerator to brake, and the vehicle started revving it's engine to try accelerating. (gears were grinding trying to shift to compensate for the slammed brake). The situation ended when i lifted my foot off the brake and pumped it.
I took it to the toyota service and they said it was due to me pressing both brake and accelerator at the same time. I was skeptical of my own user-error, so I later tried to recreate the issue (by pressing brake+accelerator together in various combinations) but was unable to repro it.
Now that I see this article, it makes me realize it's probably the firmware.
Why does the industry use something like C? Even with the MISRA guidelines, it still seems like the wrong tool for the job when you have things out there like Erlang, which are designed for fault-tolerance and concurrency from the get go. I can imagine the requirement for hard real-time operation might be an issue (not sure enough about Erlang specifically to say if it does or doesn't address or hinder this, so I'll leave that to someone else).
What's that industry moving towards, software wise, to address this? Is Ada an alternative?
C (even MISRA) just seems like a horribly risky thing to use in this scenario.
Specific to Erlang: http://www.erlang.org/faq/implementations.html (see 8.9)
It's easy to forget how small the RAM budgets still are in some parts of the embedded world.
My friend had a 50s Ford when I was in high school. The throttle control was a lever and rod attached to a couple of springs.
The point is there being multiple independent ways to stop the car. Running everything through the same central computer makes for a completely unsafe system.
1993 Mustang's have a plastic ratcheting mechanism that picks up slack on their [very heavy] clutch cables.
The teeth on the plastic ratchets eventually wear out (duh....)
If you're FORD, you design it so that this piece fails catastrophically, releasing all tension on the clutch cable. (Instead of, say, hitting a bump-stop that keeps some minimum amount of tension on the clutch cable.)
No tension on the clutch cable = no clutch. No clutch on the OEM T5 w/ a stock-spec clutch = good luck with the next downshift.
(For anyone interested: firewall adjuster + aluminum clutch quadrant = sweet deal.)
---
However I'd still argue the purely mechanical systems are _objectively better_ in that they lend themselves to preventative maintenance.
Had I thought to look: I would've seen the _very_ worn teeth on my clutch quadrant. I knew the clutch cable was old, rusty, and starting to bind in the sleeve, which probably led to the wear of my clutch quadrant. As for a throttle cable: not only can I feel it binding in the pedal, but it's fairly easy to visually inspect a throttle cable or spring for faults, rust, etc.
In addition mechanical fixes are all radically simpler. If my throttle is sticking, I lubricate the cable and replace the return spring. If your computer controlled throttle is sticking: you hope it's (A) an actual bug and not intended behavior, (B) a bug that the mfr. is aware about, (C) a bug that a patch is available for, (D) that the ECU flash can be applied free or inexpensively.
Say you're an enterprising embedded electronics engineer and you wanted to fix it yourself? If you try to modify an automotive computer: you've just tampered with an emissions control device. Your car is no longer street legal in the United States.
Purely mechanical systems by their very nature are perfectly transparent. Proprietary computer software is almost always a "black box." -- Due to federal regulations though: you have absolutely no legal way to repair or replace it on a street-driven vehicle. You are stuck with software you cannot see, understand, or control.
But consider that a good mechanic could figure out the failure in the analog world, and this Toyota issue has required deaths and 9-figure lawsuits to get close to an answer.
Unless you had cruise control [remember when that was an _option_?]... then there's a servo or equivalent motor in-line with the cable and springs.
---
Even with a digital system: there's no benefit. It's the same mechanical piece, it has to be, that's how combustion engines work: you open a throttle plate, the vacuum sucks in more air, the computer compensates w/ more fuel, compress that, spark it, you have more power.
You will always have a throttle plate and some sort of lever; and since you want it to fail closed you best have a return spring somewhere.
All you gain by adding a digital control system is [perhaps] clever software control and [certainly] additional complexity.
Is it worth it? Perhaps... there are some _very_ cool systems developed as a result of drive-by-wire. Lane following assist, parking assist, assisted hill starts, manual transmissions with automated clutches, collision avoidance, etc.[0]
I would definitely argue for more _transparency_ in this technology, but I don't know that I could argue that the technology itself is a bad thing.
[0]: http://www.youtube.com/watch?v=ridS396W2BY "Volvo Trucks - Emergency braking at its best!"
Get your manufacturer to pay a respectable firm to certify their software rather than just their mechanical hardware.