https://www.theverge.com/2021/7/26/22595060/tesla-chip-short...
https://www.theverge.com/2021/7/26/22595060/tesla-chip-short...
See the Toyota accelerator-gate lawsuits. From https://en.wikipedia.org/wiki/2009%E2%80%932011_Toyota_vehic... :
> However, on October 24, 2013, a jury ruled against Toyota and found that unintended acceleration could have been caused due to deficiencies in the drive-by-wire throttle system or Electronic Throttle Control System (ETCS). Michael Barr of the Barr Group testified[30] that NASA had not been able to complete its examination of Toyota's ETCS and that Toyota did not follow best practices for real time life critical software, and that a single bit flip which can be caused by cosmic rays could cause unintended acceleration. As well, the run-time stack of the real-time operating system was not large enough and that it was possible for the stack to grow large enough to overwrite data that could cause unintended acceleration.[31][32] As a result, Toyota has entered into settlement talks with its plaintiffs.
Not following software best practices(i.e. ISO 26262) can expose your customers to unnecessary risk and leave your company vulnerable to lawsuits. Tesla may learn a very expensive lesson one day, or they may get lucky. Time will tell.
---
Depending on the design of your system and software stack, switching SoCs can be anywhere from a massive effort to a simple recompilation. Even switching architectures could be trivial for some components(i.e. a UI written in HTML5 making REST API calls) or a nightmare for others(ECU, ABS or other real-time algorithm written to use bare metal modules on the SoC without an OS providing abstraction).
That's a matter readily solved by a PID controller made of a few dozen 74XXX parts. Possibly made n-times redundant.
Even a more fancy LP solver probably wouldn't need a fully fledged computer.
Electronic fuel injection is a _great_ example. Much of modern fuel efficiency stems from the fact that an engine can accurately control fuel in the cylinder based on a bunch of factors. Not only does this system require physically fewer materials, it uses less gas (in turn saving production effort).
What I mean under a computer is a Turing complete machine. A circuit just doing some LP solving, or PID is not necessarily a computer.
Most of the functionality of these chips can be represented in some form of hardware. However, that hardware is often much more expensive, significantly less flexible, and likely significantly bulkier.
Jadon Cammisa has a great series on this stuff. Here's just one episode on one topic, drive by wire throttle control: https://youtu.be/gKsCHx5NOMM
It's very close to what you are taught in the "How to write a PID controller" class without much extra creativity.
I mean, I could implement a PID controller without any 74xxx logic at all: a basic op-amp, a handful of resistors and capacitors and boom I'm done!
But there's a reason that engineering doesn't do that these days and it's not that engineers don't know how. My PID control built from an opamp? In 2021 that's almost always a stupid idea. We use computers because doing it digitally has many advantages. We don't use 74xx (or any other logic family) for this because microcontrollers are far more flexible, allow for behavior changes without reworking hardware, allow inventories to scale (the same component can be used in hundreds of products) and more reasons I won't detail.
The reality is that in 2021, often the simplest, most robust and cost-effective way to do something trivial is by using a computer to do it.
Very much not. Complicated MCUs are a very risky single vendor supply, and as we see now, people are now paying for an engineering overkill.
74xxx are available in every shape, and form, from dozens of suppliers.
Repeat for manifold air temp, exhaust temp, emissions, rev limiter, etc., and as you mention, make that 3x for redundancy. Now you have hundreds to thousands of components, in a high noise, vibration, and temperature environment, and you have something which is about as efficient as an 80s car. And you can't as easily simulate it cause it's analog.
Or just use some digital chips and simulate all the tuning.
Even if they had perfect readable John Carmack tier coding they still would have lost.
If/when Tesla gets hauled into court over its software flaws, any lack of adherence to best practices will make the company appear negligent.
That's your opinion. They have shown data where the information from radar was of a lower quality than what their vision stack was providing (e.g. lack of vertical resolution). Watch their AI day webcast for examples. Their current vision-only stack has been validated with LIDAR ground-truth data.
So far their biggest problem is that they allow flicker. That's why sometimes the software picks the wrong lines. (Of course this is a very hard problem. Our brain conveniently smooths over sensory changes for us, because that's how our everyday reality is. Things don't flicker in and out of existence, nor does a car suddenly appear as a different thing, then switches back.)
And if it turns out the sensor(s) failed, it has to be able to handle that too.
Unless the car loses power and catches fire. Then you can't open the fucking doors [0].
[0] https://www.washingtonpost.com/business/2019/10/23/man-died-...
This is literally just the basic fact of the situation. If you want to disagree with literally every state safety agency in the world, then you are the one who needs to come up with real evidence.
But every automatic I've ever driven (not many - I prefer a clutch) moves forward unless you're pushing the break in. In the rest state, it's in motion.
The problem isn't Toyota, the problem is a broken system.
The Toyota case was an instance of the car's engine control software unexpectedly commanding acceleration not requested by the user.
In principle it could happen on a manual car as well, but most of Toyota's vehicles in the U.S. are automatics.
They are very irritating to drive behind. Some seem to keep pressure on the brake pedal, and the brake lights stay on as they accelerate. When they slow it takes longer to realise they are braking as the lights have been solidly on.
In first gear I can reach 8 km/h, but it's possible to get all the way to fifth gear while idling (takes a long time though).
In street driving, the main application is being ready to brake while driving normally. It is much faster to brake with the left foot hovering over the brake pedal than by moving the right foot from the throttle to the brake. Of course covering the brake pedal with the right foot precludes driving normally on a flat road or uphill.
This should of course be practiced in low-risk situations before being attempted in high-risk situations.
If bad things happen, the tendency is to blame someone else.
The PCM (engine controller) that a university audit got access to had 12,000 global variables. Their conclusion was it was not possible to prove either way.
There was a definitely real issue with floor mats getting the pedal stuck. I don’t know enough background on it past the audit that was done.
There was a plausible root cause - a bit flip in the CPU. It was proved such flipping the right bit would cause uncontrolled acceleration that characterised each event.
I don't agree Toyota was guilty of shoddy work, or cost cutting. At worst, there were guilty of working some fairly shitty, hard to support code. (Someone below mentions 10,000 global variables.) If anything that shitty code just proves how dedicated Toyota was to testing something until until all the safety related bugs are gone, because despite some horrid kludges like replying on watch dogs resetting the thing when it got stuck and a reset so fast most people wouldn't notice it, despite the intense scrutiny of the code no one every found a safety issue with it. Toyota was well aware of a electrically noisy area like a engine bay could cause bit lips, and defended against it. Every variable was stored in two places in RAM, and on use they were always read and compared, and if they differed it was reset.
Sadly for Toyota they used a proprietary ECU and OS provided by NEC. NEC didn't let their customers look at their proprietary stuff. NEC didn't defend against bit flips. And it turned out a bit flop in a task ready bit could cause that task to never be scheduled again. If that task was supposed to turn off the cruise control acceleration, then bingo.
I have no idea what changes Toyota made in response to this mess, but if I was to hazard a guess it would be when it comes to software, they demand their supplies open their source to them.
Trying to assess whether thousands of global variables are still playing nice during a major rewrite to accommodate a new chip would be definitely be difficult and time consuming!
Not that a fast-and-loose approach is ideal, either.
Writing reliable software in the context of an organization is hard.
Almost every place we would have a constant (tuning parameters, etc) we instead have a configurable value with the likely default declared in code, but all overridable in config. Managing config can be a hassle, but the number of times we've merely had to tweak a value rather than roll a new build pays for itself every day.
We have thousands of such "global variables"... of course they are read-only, so aren't used to share state.
If Toyota is doing that: nbd. If they are using them to share state ... god have mercy on their souls.
It's a matter of interpretation I guess, but I would not classify a wrapper object (like .Net's ConfigurationManager singleton) as "10K global variables" even if the accompanying config contained 10K items, or if the ConfigurationManager backing store was preallocated in the data section of the binary.
Which to me on an old _accelerator_ design is suspect because that's alot of heap for the micros that would have been used back in the day. This isn't a Linux system. It was at best a micro with like 8KB of RAM. (I'm not actually sure what it is but there's no way it was impressive).
> Other egregious deviations from standard practice were the number of global variables in the system. (A variable is a location in memory that has a number in it. A global variable is any piece of software anywhere in the system can get to that number and read it or write it.)
However, that was from the article author explaining the testimony, and not the testimony itself. It's not totally clear whether the variables were writable.
I would have guessed they were statically allocated rather than heap allocated, but what really matters is whether they were `const` — and they probably weren't `const` because otherwise this testimony wouldn't have rung true:
> "And in practice, five, ten, okay, fine. 10,000, no, we’re done. It is not safe, and I don’t need to see all 10,000 global variables to know that that is a problem," Koopman testified.
The changes in Tesla vehicles are absurdly fast. They disassembled original Model 3 and then China Model 3, and since then the refresh 3 and the Y.
Recently they showed the new Model 3 LFP (Lithium Iron Phosphate) and the electronics and how it was integrated is already different again.
A lot of this era comes down to this. You have a legacy industry with tons of regulations, then a new guy comes up, steps on all the lines, and convinces people they're smarter for doing more for less. (up until regulations come back in the equation).
If anything, GM had more problems and failures of their EVs so far. So the claim that Tesla is ignoring liability and doing much higher risk development is literally just an accusation based on nothing.
In actual fact, GM is spending 10s of billions buying back every single Bolt and they have been forced to shut down the line. At the same time Tesla had no such issues.
So, how about these people actually prove or substantiate in some way that the new comer ignores regulation or quality. Specially when, they themselves have a worse record.
The traditional players are already doing a bad job trying to write safety critical software the way things are now.
The biggest challenge is when you need to update your firmware to use a microcontroller from a different manufacturer. Often times these chips are specifically chosen due to the set of hardware functionality they offer, and the firmware is written to take advantage of that from the start. The two are coupled.
Now you are forced to use a different chip, and the firmware that was written for a specific set of hardware now has to be modified for a new chip that may have a different feature set. Things fine print on how things like Analog to Digital converters becomes extremely important.
Either way, you're left with something where the saturation behavior changed, or it's no longer possible to read out a value without risk of corruption since they assumed it can be treated as write-only, or some other hard to find and debug problem.
Swapping out chips is an exercise in testing and risk management, and never should be done without care, even before we start talking about safety critical applications.
There is also a huge lack of improvement in MCU designs. Why doesn't my SPI bus autonegotiate everything? Why can't the serial port do the same? This should all be available in hardware as a standard feature of the ports. What a waste of decades of engineering talent to have to repeatedly troubleshoot everything at the bit level.
They also do a lot more in house than their competitors. That means they can optimize their development processes, and align with chip suppliers across different components. If you have to work with a multitude of suppliers that each ship their own hardware and software, life is a lot more complicated. Changes take years in such an environment. Even a simple thing such as over the air updates to software is still science fiction in a large part of the industry. VW famously struggled with doing that for the ID.3 only managing their first updates fairly recently.
Of course, Tesla is still affected by supply issues as well. They can't switch suppliers every quarter.
(In a box at 120°C for a day is roughly equal to a thousand days at room temperature).
Smarter not longer. Write in a memory-safe language and you don't need to pay people to maintain spreadsheets of memory allocations...
I’m sure Tesla is more agile than that. For better or worse sometimes…
The article quotes Intel at 16nm, which is, what, 10 years from cutting edge process node? At this point mature OEMs and Big Auto should have been on a refresh process to auto-migrate forward the chips. The semiconductor industry is 40-50 years old now. And new generations produce better chips (cost, power use, performance).
As other comments say, it smacks of laziness and lack of forecasting.
Tesla has other advantages, since they are so much more vertically integrated, they can probably manage migrations a lot more easily and centrally.
Big Auto is more like "Big assemble OEM parts". OEMs make all the components, Big Auto just wants to screw them into place and wire them together. It's part of why they suck at software that integrates things: They don't make any of the components, and a dozen different departments order them from a hundred suppliers, so getting interfaces/protocols/specs is a lot more time consuming than for Tesla.
Why not even use ARM, X86, etc? It'd save them money at this point. And there wouldn't be a supply risk.
I mean, it's an entire MASSIVE industry that's comprised almost entirely of truly commodity technology, and so you get all the perverse outcomes that happen when trying to "differentiate" despite being commodity by inventing ways to be "special" to drive lock-in. Interoperability only really benefits the OEM, and the OEM isn't the only part of the chain.
That said, we're working with some groups in some automotive companies who are breaking this mold rapidly.
Shameless plug WARNING: we (https://auxon.io) are hiring for engineering, marketing, and BD if industrial & operational technology is your thing.
BMW, Bosch, Continental, Daimler AG, Ford, General Motors, PSA Peugeot Citroën, Toyota, and Volkswagen.
I always rooted for an older competitor GENIVI which (very handwavy in a general sense) boiled down to "make dbus great again". I just always have a soft spot for anything that replaces CORBA and DCOP. More or less GENIVI had the same players as AUTOSAR.
I believe AUTOSAR suffers from trying to do too much; GENIVI's "all you need is a compatible bus" is the minimum you need so simplicate and add lightness and that's what should win; doesn't really matter if that guy runs QNX and that other guy runs FreeRTOS as long as the busses talk; whereas AUTOSAR tries to own and control everything top to bottom which just ends up stifling any productivity.
However, GENIVI/Adaptive-AUTOSAR tends to serve a different function in the vehicle architecture. It's mostly cockpit and less control/platform. AUTOSAR Classic is still the champ on the MCU side of the house.
I've never heard of anyone dying due to ECU failure either..I'm not sure how that would even happen given all the critical systems on a car are mechanical first with electronic assist. So you can't lose steering, braking, etc. The worst that happens is you lose power, which is about the same risk as a subpar standard transmission driver stalling out (and can also happen in a number of different ways). Can you expand on how an ECU failing might kill you?
I would love if someone could elaborate on why I'm wrong instead of drive by downvoting. This isn't reddit.
Take for example the Boeing 737-MAX. Eh, its bigger, needs a little more elevator movement to simulate the older model, just flex the software so it can wiggle a bit more what could possibly go wrong?
Likewise, remember the VW Diesel emissions "scandal". You can nearly kill a company without actually killing anyone.
So, the specified ECU chip (which is no longer in stock) could output 40 mA to the gas tank vacuum solenoid so we spec'd the solenoid to draw 30 mA on the coldest day of the year, usually it draws much less. Got a substitute chip only rated to 20 mA usually it'll work fine, what could possibly go wrong? Until it burns out and the vacuum solenoid fails open and nationwide millions of gallons of "excess" gas evaporate per year from sitting cars. Harmless on an individual scale but on a nationwide scale its a lot of ozone... Insert yet ANOTHER $35B recall to replace all the ECUs for willful emissions violations ...
I'm just saying its not binary where either people die and it kills the company or people aren't even harmed and nothing bad happens at all. Plenty of "company killer" situations where "what could possibly go wrong" got the F-around and find out treatment.
All that shit's still not in stock anywhere. And remember that ARM, x86, etc. only defines the processor architecture. The I/O is far more important in a control application and I/O is anything but standard across platforms.
For one product we ended up buying a vendors entire supply and redesigning the product to use it, crossing our fingers that the supply chain would have unwound itself by the time we ran out of inventory.
But I'd also like to throw in:
- 16 years or so ago, the US side of the Auto industry was trying to steer towards PPC. Motorola/Freescale's presence in the market probably had a bit of a play in this, as well as ARM's then-status of 'not quite powerful enough to be future proof.' A StackOverflow post from 2012 [0] seems to indicate they did indeed standardize, whether they moved on from there is another question.
- Most Carmakers probably -don't- want to change designs outside of a refresh. There's a few reasons for this, including both external (updating documentation to repair network) and internal (when I worked at a place that did software/services for a US automaker, we had to submit pages of documentation/paperwork to change a couple of items placement/wording in a dialogue box.)
[0] - https://electronics.stackexchange.com/questions/26035/whats-...
Only the libraries for I/O need to be rewritten.
In some cases, there are drop in replacements (Gigadevices STM32 clones)
I seen a few dashboard where there is an STM32 sitting connected to a single button, whose only duty is to register a key press, and burp something on the CanBus in response.
B) If you did have a $20 microcontroller for Canbus control of a window motor driver but it saved $19 in extra wiring/labor (driver controls usually route to all windows), provides reliable debounce, automated one-push-to-open, and allows things like holding keyless lock to close all windows, it's totally worth it.
Single chip system, with built in analog comparators, timers, signal capture and PWM generation module, serial UART, etc. I.e., a whole "system on a chip" with the external world interfaces already built into the chip. Changing one of these for an ARM CPU with external analog comparators, timers, PWM modules, serial UART's, etc., is a redesign effort, not just a "change the chip" effort. If the clock speed possible via the internal clock oscillator is sufficient, then this chip needs only +5V and ground (plus programming) to be able to interface to and control some external analog or digital equipment.