The Power of Power Cycling
growthalytics.com
growthalytics.com
Yesterday, I had to power-cycle a CNC milling machine. The machine (a Tormach 1100) started the spindle with the spindle RPM about a tenth of the value specified. This resulted in pushing an end mill into the workpiece (it wasn't spinning fast enough to cut properly), snapping off the end mill (like a drill bit) which went flying off, and damaging the workpiece. This is not a huge deal (it's about $25 in damage, and everybody wears safety goggles around milling machines) but it shouldn't happen. We discovered that all spindle speeds were far slower than they should have been; entering a new spindle speed did cause the spindle drive to change speed, but the speeds were about an order of magnitude low. The spindle speed control is digital and computer-controlled.
Power-cycling the machine (which runs Windows Embedded Standard, a Windows XP derivative) cleared the problem, and I was able to run several jobs successfully. But someone else had reported a spindle overspeed on Friday. Something has gone badly wrong in spindle speed control on that machine.
The "Recommended Best Practices" for this machine tool actually says "Reboot controller once a day to force both Mach3 and Windows to restart."[1] That's not good.
[1] http://www.tormach.com/uploads/883/SB0036_Mach3_Best_Practic... (a PDF file with the wrong suffix)
I recall being told about a high end cnc machine that power cycleled and moved the tool to home straight through the work piece.
That doesn't mean that individual components can't be power cycled on fault without affecting the rest of the system. That, of course, means writing the system in a certain way that allows that (allows fault isolation). You can do it with Erlang, OS processes, containers, separate hosts, etc.
In other words there is nothing that prevent the company to auto-power-cycle some component that crashed. Report the error and then issue a fix later, while you still get to use the machine.
"Disclaimer: If you have a recurring problem, power cycling will not fix your root cause. "
At some point I must have decided the whole thing was too long and that sentence didn't make the cut.
I personally would not go as far as to say:"it's not acceptable to power cycle to fix problems".
I see it more like:"not all problems can be fixed with power cycling".
Sure they might be the only vendor, or someone inside your organization might choose them anyhow despite your protestations. But that doesn't change the fact that using Windows for these kinds of things is foolish.
Just because big companies do it doesn't make it smart, does it? I mean, if it did, then startups wouldn't be able to exist would they? Startups are able to be a thing because big companies sometimes do stupid stuff.
Most of the time these machines run a small microprocessor which receives commands over some port and interprets them as it is instructed to. This is what's called a "hard realtime" system as it always responds within a certain amount of time, guaranteed, provably as per the design.
What Tormach is doing is eliminating the dedicated gcode interpreter hardware/controller and performing those operations strictly in software, on a program running on a PC. There's some utility to that, but pretending that it's as good as having a dedicated, realtime gcode interpreter is not honest.
I would guess Windows is used for the higher-level functionality (GUI, possibly format conversions), while Mach3 does the lower-level stuff that requires precise timing.
Adding a second OS makes sense as it makes it way easier to keep the real-time stuff real-time.
Using XP embedded for the top layer shouldn't be that big of a risk. It may have lots of known exploits, but you can remove lots of the attack surface, and an alternative GUI may not have seen much security auditing.
- "failure is not an option": you don't get to reboot your rocket controller when it has a floating point error, or your Therac when you're doing an X-ray. This involves investing a lot of time, effort, review, and formal validation into getting it right. It's completely incompatible with Agile and RERO.
- "s--- happens": the cost of failure is small and you can just accept it, issue refunds/apologies and move on. Or show the fail whale. This is much easier and cheaper. This environment moves towards powercycling and redeploying software/EC2 instances/Docker containers whenever something happens. You monitor observed reliability and make a commercial decision as to whether it's unacceptable and you need to fix some bugs.
Almost all of HN works in the second area.
Power cycling is different. A system that expects you to power cycle often because it consistently fragments or cannot dynamically update/reread its configuration is broken and a major annoyance. In many ways, systems that require lots of power cycling are as such because they're designed in a way antithetical to crash-only. It's a failure to enforce boundaries and separation of concerns.
OS processes/Erlang processes/Separate Machines(Containers) can reliable do that. Because they isolate fault propagation.
In a shared memory system if something crashes you don't know what the state of the rest of the system is. Maybe on bad client wrote over the memory of other 999999 clients. So just restarting that one thread is not safe.
Another overlooked aspect is to do crash-only right with recovery you need stable storage. That is where you can persist a known good state so you can restore from and continue. That might be hard or easy depending on the environment.
> You monitor observed reliability and make a commercial decision as to whether it's unacceptable and you need to fix some bugs.
Yap. But "you" here is the developer not the end user. When you go to buy socks on a website, it is not your job to monitor and restart their back-end system or switch IPs. They should be doing you are just buying socks.
> Almost all of HN works in the second area.
Because the first is very hard. NASA does it, medical device manufacturers do it. Critical security modules have it. It is very expensive though.
It turns out rocket controllers get power cycled too.
At the systems level you often have a mix of both classes (like your Xray example, typically). And there is often a third path "failure has a well defined mitigation path" in the mix, also.
I mean, it shouldn't happen, but if "should"s were horses I'd have a ponyburger for brunch.
There's so many examples I could provide: you are a software engineer, the program you develop works well for a day and then crashes. You can either investigate the problem, find there's a memory leak and fix it or Power Cycle your program.
-> If you design a product, don't go for the easy stuff, don't power cycle just to get back to a well known scenario. Investigate and fix.
I often reduce how I classify the importance of bug reports I get that I can't reproduce. It's often not efficient to work on something with so little information to go on. Further information on the bug or a reproducible error case vastly help.
Then the engineering team never gets to fix the problem, and your support team starts eating up all the resources because they get constant calls from the customers asking about the same problems.
(thankful for not being in that job anymore!)
1 0 * * * kill -HUP $(cat /var/run/my_server.pid) >> /var/log/my_server/error.log
You think I've kidding but I've done this in production. Now I'm more likely to use god/monit.In fact, not joking at all: this is how I make Finder not slow down. Every 4 hours, I restart it via crontab.
Even though I agree with this, there is a fundamental difference between computer science and software engineering, and in actual software engineering there will often be hacks, and bad state. Even if sometimes it's not yours, but a library you're using.
I agree with you, but this article deals with a more "real life, pragmatic" approach, versus your thesis which is true, but a little more idealistic
Any company that expects its developers to take pride in their work should track crashes and strive to produce software that doesn't routinely need to be restarted.
The reason power cycling is such a useful tool is that most computer systems are running in an environment where the things outside of it's control way outnumber the things within it's control.
Power-cycle related story: my machine crashed two weeks ago while playing a movie and my 3TB hdd showed only 2TB on _SMART_ after power reboot! After googling around, letting the hdd cool off and playing with wd tools, several power reboots later my 3TB partition is back but not without loosing a couple of MB of the size of the disk at the end of it. Puff, just gone, not deleted, gone! My old partition is now bigger than the 3TB drive.
Smart reports a 3TB disk again, and I'm on the look for a spare drive.
No, really, not all power cycles are equal.
The push for making broad use of consumer hardware running production environments to save (a veritable fistfull of) money has only made this problem worse.
It makes me wonder if we shouldn't be adding a module to the Linux kernel to occasionally reload blocks of code from disk.
Why reboot it in the first place? If the HW has ECC, the SW does not rot unless it is poorly written.
And the OP assumes that once 'off' has been achieved that "on" is always a possibility. Forget the car. Think airplane that has lost instruments, but the engines are still working. Turning things off might move you from bad to worse. Maybe you stick with the partially-working machine rather than risk bricking things.
Or that to do this is a simple operation with no other side-effects. As one of the other comments here points out, power-cycling is probably acceptable for consumer products, but with systems that are more mission-critical, should really be closer to a last-resort method.
The reason why power cycling works at all is because the machine's state is inconsistent with the program's, and power cycling gives the machine and the program a chance to start over from scratch.
So many programs out there, especially so in embedded devices where power cycling is common, don't really have an explicit state machine model of operation. As such, the debugging of these types of errors is near impossible. If you have an explicit state machine model of operation, and you have a bug that is remedied by a power cycle, you can quite easily trace the bug back to a specific state transition. Of course, this becomes more untenable with increases in system and hardware complexity, but on the level of a driver or a single executable, an explicit state machine model works wonders.
I must have spent several hours writing and re-writing that function, and I still couldn't get the tests to pass. I went over all the code character by character, debugging, etc, and I just couldn't figure out where it wasn't working. Eventually, I just nuked all the work I did that morning, rewrote that entire routine from scratch, and it worked perfectly.
I yanked the power cord, put it back in. The steppers re-homed themselves and the bottle made it to the bottom.
Free Coke.
'A novice was trying to fix a broken Lisp machine by turning the power off and on.
Knight, seeing what the student was doing, spoke sternly: “You cannot fix a machine by just power-cycling it with no understanding of what is going wrong.”
Knight turned the machine off and on.
The machine worked.'
See the rest at http://www.catb.org/jargon/html/koans.html
There were some interesting reads in the AT&T Technical Journals for the 5ESS switches (e.g. Kintala, Bernstein , Wang “Components for Software Fault Tolerance and Rejuvenation”) a long time ago. The journals are not free accessible, but I have found another text from one of the authors about the same topic http://www.crosstalkonline.org/storage/issue-archives/2004/2... (defense related web site) or http://srejuv.ee.duke.edu/shaman02_secure.pdf (other author)
As per your comment, you are right: if there is a kernel bug and you'd rather restart your computer to quickly get back to what you were doing and avoid distraction, all good.
If you are part of a kernel engineering team and you quickly restart your computer and pretend it never happened then maybe you didn't give your best shot.
Not really. If it happens once, it can / will happen again. In my opinion a bug is a bug and should be fixed. But I do agree this attitude does seem prevalent.
If the problem only ever occurs once and power cycling fixes it then that is effectively a permanent fix.
If it comes back next year and power cycling fixes it again, you know there's something wrong but it probably won't be an actual problem for any practical purposes unless it gets worse.
If you find yourself to need to power cycle every month or every week or even worse, increasingly more often, chances are you're going realize that the root cause must be investigated and fixed in a matter or months or weeks or even days.
"Have you tried turning it off and on again ?"
- from Graham Lineham's excellent I.T Crowd , the phrase is spoken by Roy from Technical Support, the I.T. dept.
https://www.youtube.com/watch?v=p85xwZ_OLX0 - Irish brogue ?
https://www.youtube.com/watch?v=tPxrACTPxEk - Minessotan US ?
It results in the software engineer basically saying, "lets try it again and see if we can reproduce it".
:~$ uname -r && uptime
3.2.0-4-amd64
01:33:09 up 461 days, 3:50, 13 users, # uptime
14:34:38 up 587 days, 21:12, 1 user, load average: 0.00, 0.00, 0.00Seriously: I'm assuming that these are both servers running a single application or small group of related applications. OA was talking mainly about client style devices.
The bottom ice maker stopped working and no matter what setting I had it on would make it create ice.
My current theory is that I changed the ice settings while the freezer was open (a few days) earlier. I noticed this because the water in the door shuts off if you open the other door.
Sure enough, after leaving the fridge unplugged for 15 minutes I now have ice.
"one error per month per 256 MiB of ram was expected"
http://stackoverflow.com/questions/2580933/cosmic-rays-what-...
Using it to solve a daily problem is a terrible idea, if you have any control over the system.
TL;DR: Most software bugs that make it past testing are transient "heisenbugs". That is, they're the kind of bug that goes away when if you restart the program.
Related: This is actually a core tenet of the Erlang ecosystem -- spend any length of time around Erlangers and you're bound to hear the phrase "let it crash". Erlang actually has support for this built into the system: Supervisor processes exist to automatically "power cycle" your code if an unhandled error occurs.
Not disappointed, but not what I expected.
If not, why not? That'd catch any "mad user bugs" and all kinds of accumulated collateral damages from the sea of complexities which might be overlooked if routines are only tested individually.
It seems so common its mandatory in some places / fields, and totally unknown in others unfortunately.
You must disconnect its power by turning off the PSU or yanking the cable. I suspect a majority of tech people know this but your PSU passively supplies power to a few things on the board, this can be enough for controller errors to persist across a reboot.
The worst time for power cycling a machine is during an emergency.