As winter approaches, here's a story about why hardware is hard
twitter.com
twitter.com
A fellow new hire and I were tasked with fixing a machine that ran in a clean room where semiconductor wafers were made. On weekends while the line was down, we would go in, crank up one section of the line and let waste wafers travel through the line where very rarely they might get stuck in a multi-lane machine that would etch the wafers with some sort of acid.
The machine had over two dozen asynchronous motors, actuators, pumps, sensors and so forth. All generating interrupts and I/O events that were sent to a computer that ran the whole line and controlled all the machines.
We couldn’t slow down the machine, it had to run at full speed. The program controlling the machine was thousands of lines of assembly language—everything was assembly language, including the homemade OS that ran the computer running the line. It took like an hour for us to bring up the line and two more hours to see the machine do something strange.
The computer running all this had no user interface other than a some front panel switches and some panel lights that would reveal 16 bits of it’s 128K of memory at a time. This was in the 1970s before Ethernet had been invented.
It felt a bit like those escape room events where you know there is a solution, but you don’t know if you will ever get out. Without my coworker cracking jokes about our plight, I’m not sure we would have ever triumphed over that stupid machine.
I remember one episode where they figured the wings would freeze and not move so the pilots lose control and would crash.
They reengineered the part to not freeze during extreme temperatures.
Edit: just realized their stated fix is to remove two resistors. Did they actually test the full effects of that change? Seems like there's a decent chance of a third head-scratcher a few months from now...
Surprised, but not surprised, that everyone doesn't do at least some stress testing.
I was expecting something more complicated, like electromagnetic interference from a passing vehicle. Or ESD (static discharge).
Not to mention that the standard checklist for solving "why is this joint having problems" is "is it cold or does it get worse when we make it cold"....
> just realized their stated fix is to remove two resistors. Did they actually test the full effects of that change? Seems like there's a decent chance of a third head-scratcher a few months from now...
As a half-decent analog guy I can actually believe this. I can imagine a lot of ways something like this could fix a problem, though of course I haven't seen any schematics here.
However, if they were working with a COTS component and took some resistors off that, I doubt it was actually the right fix.
0 - https://en.wikipedia.org/wiki/Failure_mode_and_effects_analy...
On one hand I’m surprised they didn’t do that, but on the other, I have no time constraints and I’m very, very uncertain of what I’m doing most of the time. I knew I wanted to have accurate readings and reliable performance so my garden doesn’t die due to malfunctions. So, I goofed around and made sure stuff was right. I’d do the same with code in my spare time, but my employers have me cut corners all the time. People could criticize me for it, but it’s not as though I don’t know better.
It might be similar for this team. By the time the bugs strike, it’s not clear what’s been properly vetted or who knows what about which components. Debugging becomes harder because the initial spec and how well it was met is no longer clear.
I’m totally guessing as I don’t really know their team, process, or hardware in general.
I also discovered the ruggeduino line of arduino boards in the process of component elimination, which are pretty cool. Overkill for my use case, but I hope to have a use for one some day. I’m thinking of making a robot shop vac and metal-picker-upper, and the ruggeduino would be great there I think.
If they develop a business need to avoid similar issues before they happen, then the answer would be yes. The work they unwittingly skipped (like design for environment) and struggled with (diagnosing an environment problem) are the daily tasks for many people, so they obviously don't have one of those people.
-old hardware lore
For those unfamiliar with this issue, or who forgot: The power regulator IC on the Pi 2 was a flip-silicon slab that was otherwise uncoated or encased. When engineers were taking pictures of it, they likely were using their cell phones and otherwise loose lighting under fluorescent or sunlight. When high intensity Xenon bulbs were used for flash photography, however, intense IR light would hit the silicon, excite something, and cause the power supply to go wobbly kneed, drop out like a Silicon Valley startup founder, and you'd have a bad day.
This failure was only figured out when someone happened to compare two flashing lights (one Xenon and one a bike light) that they figured out what was going on.
Engineers miss things all the time, even those who are supposedly really good. They are good, but humans err, because to err is human.
Prime example: I diagnosed an arcade cabinet that only had issues on sustained playtime during the summer, but never during the winter. It was fine on the off day that there was a lot of cloud cover, and only seemed to get worse after a shade awning was taken down: before the removal, it would intermittently sputter back to life during the later afternoon. Nobody knew what was going on.
I spent an afternoon watching it fail. At one point, I eventually went and grabbed an IR thermometer and pointed it at the cabinet. It was registering well into 100F. It was also outright refusing to work.
I rolled down the window blinds, pointed a fan at the poor machine, and waited for a while. Miraculously, it woke up after a while and started working.
I traced it to the three canned voltage regulators. The voltage regulators for the system were strapped to the metal chassis of the CRT, which contained the system's power supply. The machine happened to be right next to a window, and the cabinet had been painted black by the previous owner. As the daytime sun would wander over, the black paint would soak up the UV and absolutely cook the regulators. Once the sun had passed behind the shade, it was quickly cool enough inside the chassis to bring the regulators back into spec, which caused the machine to start working again. When the shade was removed, it only went back up once the sun was no longer cooking it like a roast chicken.
I believe my recommendation was "Replace the side paneling with woodgrain and move it out of the sun."
The engineer that designed that arcade machine probably never thought it'd be in a hot, New Mexico university faculty lounge next to a window that faced the sun. I remember going back about a year later and sure enough, it had moved about three feet and was "rarely, if ever broken" now.
Engineers miss things all the time.
One of the most memorable corollaries to Murphy is:
"Mother Nature always sides with the hidden flaw. Mother Nature is a bitch."
And so there are many problems we [programmers] don't have. For instance, if we put an if statement inside of a while statement, we don't have to worry about whether the if statement can get enough power to run at the speed it's going to run. We don't have to worry about whether it will run at a speed that generates radio frequency interference and induces wrong values in some other parts of the data. We don't have to worry about whether it will loop at a speed that causes a resonance and eventually the if statement will vibrate against the while statement and one of them will crack. We don't have to worry that chemicals in the environment will get into the boundary between the if statement and the while statement and corrode them, and cause a bad connection. We don't have to worry that other chemicals will get on them and cause a short-circuit. We don't have to worry about whether the heat can be dissipated from this if statement through the surrounding while statement. We don't have to worry about whether the while statement would cause so much voltage drop that the if statement won't function correctly. When you look at the value of a variable you don't have to worry about whether you've referenced that variable so many times that you exceed the fan-out limit. You don't have to worry about how much capacitance there is in a certain variable and how much time it will take to store the value in it.
All these things are defined a way, the system is defined to function in a certain way, and it always does. The physical computer might malfunction, but that's not the program's fault. So, because of all these problems we don't have to deal with, our field is tremendously easier.”
— Richard Stallman, 2001: https://www.gnu.org/philosophy/stallman-mec-india.html#conf9
But reproducibility can be vague and sometimes, when you're under pressure, you can be quick to point to something and declare "aha! that's the root cause!" and be totally wrong.
We hope that fixing one piece of code will solve three or four exhibited problems. But it's often more like, change three or four pieces of code to make one problem go away.
As Heinlein is purported to have written, "If it's not one thing, it's two things."
An archive.org search find a few examples, like this 1979 Harlequin short story https://archive.org/details/romanticshortsto00harl/page/30/m...
If Heinlein did use a phrase like it, I expect my searches would have found it.
It doesn't appear to be a common saying, so I'm curious how you acquired the association between it and Heinlein. It doesn't seems like a common misquote people end up spreading.
[1] I've only read "Tramp Royale" up to the point where they left the US, I haven't read the "stinkeroos", nor his posthumous novels, nor most of what Wikipedia lists under "Other short fiction", nor a couple more non-fiction publications.
Why do you associate it with something written by a famous author?
I don't see anything which suggests its a well-known phrase. Google and DDG together found fewer than 100 pages using that quote. Most appear to have been spontaneous creation. One attributed it to "a Norwegian friend."
(This all assumes my search terms were meaningful.)
By Any Other Name, Spider Robinson:
" And McLaughlin rescued the moment, in that split second before Higgins’s control would have cracked, doing his prizefighter imitation. “Aw Jeez, Tom, that hard cider. If it ain’t one thing, it’s two things. Go ahead; we’ll keep your shoes warm.” "
for completeness, it's a play on the more common expression, perhaps best summed up in the modern age by Snoop Dogg in Pump, Pump:
"if it ain't one thing it's a MFing nother"
While reading the thread, a red flag[1] was immediately raised in my mind when:
> We couldn't reproduce it, but we did come up with a theory for why it was happening.
...going right into mechanical subsystem redesign. Surely a cursory review would have challenged such a reactionary proposal: What meaningful steps were taken to falsify the prevailing theory?
There's something implied about discipline when this vacant QA Tester Hardware/Software engineering position description[2] bundles verification/validation test roles on the design/development/production/field support fronts with the following caveat:
> Initially, you'll be the only QA engineer and will perform active testing of new product releases in our lab and in the field at construction sites.
Also, non-rhetorical question: Selenium for industrial hardware test automation...is that really a thing in the wild?
[1] https://twitter.com/tessalau/status/1604018887603138561
[2] https://boards.greenhouse.io/dustyrobotics/jobs/5373908003
And worst of all, for me, there is no money in hardware. At best you make a trinket that requires a $9.99 subscription to really get use out of. At worst you make a cool trinket, get forced by pricing to make it in China, and then end up just having the idea stolen and reproduced to be sold for 1/2 the cost.
Ok rant over.
And their SoC's of course, but ain't no lone engineer spinning up their own phone, much less their own SoC.
I'm really talking about hobby projects. Hobby swe you can actually make a product, sell it, and make some income. Hardware? Maybe you can make something, but make some money? lol.
By the time we had rolled out the coupler "fix" to all robots, the weather had warmed up enough across the country that the issue didn't reoccur. We thought we had fixed it, when actually spring fixed it."
At first they correlated a possible root cause and then after learning from that mistake they finally understood the root cause.
I've seen it happen many times where people with not enough time and knowledge to debug a huge system had to resort to shotgun debugging. IME taking the time to understand always ends up 1) solving the problem and 2) saving time and money.
This is especially true when the problem is actually caused by two or more root causes.
There are ISO standards to tests for temperature and humidity resilience and just as you should test for EM immunity and emissions, environmental testing is just as important, especially for industrial hardware.
It’s possible that in this case the customers were using the product outside its specification but when you are designing an electromechanical device there are tons of things that can go wrong once you’re put of the narrow band of environmental comfort.
Grease and lubricants can seize or liquefy, metals expand, humidity affects corrosion and heat convection, rubbers and seals can harden and contract, electronic components can overheat or change characteristics or prematurely age,…
I don't have an environmental chamber, but I wish I did. Everything I build goes into the fridge for at least a few hours (which means it gets tested at extreme humidity as well as extreme temperature, for better or worse.)
I should put boards on a hot plate at ~60C for a similar length of time, but I didn't do that recently, and I paid heavily for that bit of negligence. Probably wasted 100-200 person-hours at the factory, having them rework a NOR flash part that was fine all along but didn't like being inadvertently overclocked by 2x once the board reached operating temperature in the test area.
In a core dump all the IO (network connections, files, etc) is closed while a debugger gives you access to the functioning environment.
https://www.nytimes.com/2022/12/14/world/europe/ukraine-russ...
Turns out it was the cold. Now when I take a trip out for coffee & coding, I boot up and let it sit for a while before starting my work.
There is no mystery or surprise here. It’s basic functional qualification. You buy or make a thermal chamber and cycle release versions of your device before you ship one. This isn’t some uncatchable mystery, they just didn’t test adequately.
Depending on the product size and cost you may also do this to every individual robot off the line. This is not uncommon.
This isn’t “hardware is hard” this is “we thought it was software with screwdrivers.”
This is a general problem with all these California-based companies and inventors (especially SV inventors cranking out crowd-funded bike stuff.) They seem blissfully unaware of things like cold weather, water, dirt/mud, and road salt...or combinations of them. I laugh at all those stupid fucking delivery bots because they'll fall apart anywhere there's snow, and get completely stuck on the slightest bit of ice.
For many years, driving a Model S in heavy rain would cause water to get into the drive unit via either seals or vents that weren't sufficiently designed to keep water out. It "totals" the drive unit, causing corrosion of the motor control boards. And Tesla denies warranty claims on such repairs, because of course they do - just like they did on the windows that randomly shattered in parked cars.
Raise your hand if you've owned a car that had problems with water ingress issues affecting its transmission. Or windows randomly shattering.
What's that? Nobody? Exactly.
I have owned a couple of Philips bread makers. Basic ones and expensive ones.
If you make sourdough with them (ie bread) the coating is stripped off the bowl and the stirrer corrodes.
They will deny replacement and claim you sprayed something acid on it. Yes, fermentation is acid, but they don’t believe bread would damage their unit.
When you run a hardware startup, you can only hope for an experienced team that would do everything by the book and implement best practices from the very first production unit. Reality is: that’s a luxury for most hardware startup teams out there.
Typically, there’s a frantic rush to get your device to market that you simply skip, or more likely don’t even have time to think about stuff like climate chamber cycling.
One thing I’m almost sure of: these guys have learned something — the engineer’s way. Good chance there’s a guy there googling climate chambers to ask the CEO for a budget to buy one.
And that, right there, is the difference between a company oriented around myopic management vs a company oriented around robust quality.
Any company trying to build a quality reputation would spec this stuff out AT THE BEGINNING — what are the operational requirements, what loads will they see, in what environments will they run, etc??? Then spec every component, and test the whole lot against those requirements. Sure, this is more like the dreaded "waterfall" vs "agile", but the result is a quality product from the start that has far fewer of these problems (because they did this whole test & fix routine at the prototype or Alpha test stages), rather than showing up with stories like this of how they recovered from customer-reported problems.
If you're telling your customers that they're the Alpha testers because they get early access, fine. If you're selling it as a finished product, then we know your company isn't prioritizing actual quality.
Sounds like Tesla... although it is public knowledge there that you are still alpha testers years after a model was introduced.
> When you run a hardware startup, you can only hope for an experienced team
If you run a hardware startup and fail to acknowledge that places like Alaska or Ontario exist, you fail before even getting close to merely inexperienced. The most charitable word I can think of is "myopic".
I suspect I would be disappointed by the performance of an iPhone that could operate in >125F temperatures (temperatures which I have worked in outdoors for several years)
I’m undecided if they lacked imagination. I work with a lot of electronics that go in oil wells and we have to make different models for different temperature ranges. We usually have to sacrifice a lot of functionality to gain high operating temperatures.
I think it’s possible Apple decided not to serve the markets which need high operating temperatures, rather than simply didn’t even think about the possibility.
It affected me, when I worked outdoors in Saudi Arabia and UAE and even Houston. But I probably would still buy a high performance, high battery life 13 Pro Maxover a hypothetical lower performance 13 ExtremeEnvironment edition.
Operating at 125F ambient would require a whole lot of sacrifice on the power envelope and much greater overall size of the product. I like big phones but my current Pro Max is about as large as I’d like a phone to be.
I was born and raised in a place where the annual temperature range can be 90°K. From -55°C in the winter to +35°C in the summer. That is not even an extreme range, there are well-known places with large populations (>10MM) that can experience >100°K differentials.
Perhaps more importantly, temperature changes of 40°K in less than 24h are not uncommon. Daily deltas of 25°K are experienced several times a year.
If a company builds hardware, and they don't factor in a routine impact of thermal effects, they really have no excuse. Choosing to ignore these can be a valid product development or marketing strategy, but not being aware of them is nothing short of myopic.
My point is that turning things that are unknown unknowns (from the point of view of the company) into known and checked for possibilities takes concerted effort, time, and money. It's easy to look at a problem after it's found and post-facto determine how to find the same problem quicker next time.
I agree that temperature is very basic, low hanging fruit. Especially for a device that seems to be aimed at operation by construction crews. But regardless of where you draw the line on testing environmental factors, you have to draw it somewhere. And so you will still end up with unknown unknowns that escape your QA or debugging process, sending you down the same path of needing to question your assumptions to figure out what's going on.
(Also you're not really giving them the benefit of the doubt here with this assertion that they didn't take temperature into account at all. It seems that they at least looked at the part temperature ranges. And it'd be courteous to assume that they did power dissipation and temperature rise calcs. What they didn't do was component or integration testing at varying temperatures.)
The difference is we did it before the units reached customers.
:)
It reads to me like they didn't properly spec/source components that were appropriate for the weather conditions these robots are likely to see. The fact that this issue was reproducible at refrigerator temperatures is even more shocking. 39F (taken from a photo) is not very cold.
Yeah that's pretty inexcusable. As a comparison, common "industrial" semiconductors are often specified from -40C to 125C.
Range Min Max
Commercial 0 to + 70
Industrial 0 to + 85
Automotive -40 to +105
Extended Automotive -40 to +125
Military -55 to +125
But that's for the parts themselves, and just working at that. Weird things can and do happen at temperature extremes. If you're very lucky, they're documented in datasheets/manuals. If you're lucky, they're known enough to be in white papers or industry presentations, or known to one of the companies that do this stuff for a living. If not... well, hope someone's heard of it, or you've got a huge testing budget....That is a total waste of time in a startup with a small number of units shipped.
A component behaving out of spec due to temperature excursions simply isn't that common nowadays. If my system is mostly ADC to digital to DAC (standard for robotics controllers), testing for temperature is a waste until I'm shipping significant volumes.
There is a video of one of the slightly famous YouTubers who has a high voltage thing that fails at the altitude of his lab. The manufacturer did check it for function at the elevation of Denver, but his lab is higher than that. There are limits to how much engineering effort you should put in until you get an actual failure. (Maybe someone can link the video as I can't remember it at this point.)
You can waste infinite engineering effort covering all possibilities. Or you can ship the thing and fix the failures. "Good engineering" is about balancing the two--you need to ship, but you don't want to have too many failures in the field either.
A $5k temperature chamber running over night or weekend would have caught this.
You are right. People are commenting here as if the company didn't know what they were doing. I suspect there is more to the story then what has been shared in a tweet. My guess is that the deployment was done in a place with a temperature differential well outside the component specs/tests and so simply was not known/tested for.
They wasted resources on fixing what they "thought" was the problem. That money would have paid for some nice chambers and other equipment.
Next thread: "We learned our robot doesn't work when sitting in a shipping container at -40C for three months because customs was being difficult."
Thread after: "Hardware is hard! We learnt about EMC compliance."
If they didn't account for temperature, they certainly didn't account for moisture, and these boards probably don't have any conformal coating on them. I guess its good for the customer that its a robots as a service, but these things are going to end up failing spectacularly in the field. These things are outside at construction sites, they needs to account for the environmental conditions.
Leave it to Californians to forget that winter exists.
This isn't rocket science.
Now, if you happen to be OK with selling a product that isn't reliable, now's a good time! I still need some more toys for the holidays.