When power cycling your (x86) server isn't enough to recover it
utcc.utoronto.ca
utcc.utoronto.ca
It was not uncommon to return from lunch to find than an embedded computer board that had been working when you left wasn't any more. One way to debug them was to put them in the refrigerator for a while. If they then worked, you knew you had a bad solder joint or an IC that was on the verge of failing.
First one is dell latitude laptop with fingerprint reader, randomly after few days of operation, fingerprint reader stops responding and login screens freeze for a minute until it timeouts few times. reboot does not solve it, nor suspending machine. it needs to be powered off and on again (hibernation to disk also works).
second case is my pc with ASRock creator x570, after long time if keeping it suspended, WiFi card stopped to function and just throwed some errors in dmesg on driver initialization. here even power off and on did not help, but flipping switch on power supply for few second resolved the issue
For obvious reasons AMD boards don’t tend to ship with Intel wifi, but in my experience anything else sucks. The intel 6e cards are amazing and dirt cheap.
I also dual boot, in addition to being an incurable distro hopper, and these AX210 cards worked out of the box in basically everything.
Cause realtek checks the box for has wifi and costs probably $3 less? If you care, you can swap it, and if you don't, you don't.
Funnily enough, the threadripper (at least WRX90, and at least asrock) come with an Intel dual 10Gb LAN card. Probably because none of the alternatives are good enough for a pro board.
I had to do this secretly because company warranty/service deal would require it to send to dell/request technician
Some Dells have a "feature" where something, somewhere, in their mess of a UEFI/iDRAC stack will get corrupted and will stay wrong through power cycles until you physically unplug the servers from power and hold down the power button to discharge a capacitor and clear out the NVRAM where the corrupted value is.
Most recently this impacted a PowerEdge R7525 server we have where the iDRAC was enforcing a power cap of ~300 watts leaving the system to be less than 1/10th as performant as it should have been. Manually setting a new power cap did nothing except update the values displayed in the UI. Multiple six minute (because of their mess of a UEFI/iDRAC stack) reboots of both the server and the iDRAC did nothing.
Dell was less than useful except for the fact that they hosted the answer. After raging against their CSA script/LLM auto-reply bullshit for days an aggrieved user with the same issue looking for help in their forums finally posted that he did the cap drain trick and it worked.
Saved me tons of wasted time. Thanks, anonymous fellow frustrated dell customer!
[1]: https://github.com/torvalds/linux/blob/9b2ffa6148b1e4468d08f...
I unplugged it and left it overnight, and the next day, the ME was gone.
This was the ARC version, but it can remain operational for some time after power is removed.
That being said, there are some versions of BIOS that do allow turning the ME off, but most motherboard and laptop manufacturers will not allow general consumers to install that version of the firmware. There are some groups that have figured out how to sign a patched fully feature-unlocked BIOS on a per machine basis (disabling ME is a simple Y/N flag), but YMMV given these tools are nearly impossible to get working.
AMD should end the clown show of RATs, and eat the remaining Intel market. =3
Some motherboards allow you to disable it, and it doesn't do as much as ME in the first place (no network modules or built-in remote access purpose like ME)
[0] https://en.m.wikipedia.org/wiki/AMD_Platform_Security_Proces...
It's still there, but unlike most consumer BIOS can apparently be turned off (whatever that means to Intel.)
Personally, I don't hold a lot of hope outdated on-chip minix OS can't be exploited/activated anyway. =3
Yes, if ME detects a problem when initializing it grants you a 20 minute window as a grace period, presumably to allow users to attempt to fix it.
> There are some groups that have figured out how to sign a patched fully feature-unlocked BIOS on a per machine basis (disabling ME is a simple Y/N flag), but YMMV given these tools are nearly impossible to get working.
You can also just flip the HAP bit[0], I'd assume that's what those advanced (usually leaked dev build) BIOS firmwares do anyway.
> AMD should end the clown show of RATs, and eat the remaining Intel market. =3
AMD has PSP[1], which is functionally equivalent (though with a significantly smaller attack surface, when left enabled)
I personally am of the belief that both technologies are likely backdoored. There's so much pointing against them[2], that the simplest explanation is they're more likely than not a mandated backdoor that chipmakers eventually expanded for other purposes (such as recent versions of ME handling suspend-related power management)
[0] https://github.com/corna/me_cleaner/wiki/HAP-AltMeDisable-bi...
[1] https://en.m.wikipedia.org/wiki/AMD_Platform_Security_Proces...
[2] https://en.m.wikipedia.org/wiki/Intel_Management_Engine#Asse...
This is why we can't have nice things. =3
Yes, power-cycling is more unambiguous, but afaikt, the example here is purely that power cycling really needs a noticable off-period so that all devices can fully come down. Otherwise, there's no real standard on what should happen - this or that component might stay up or retain state.
The other reason I like 'reset' is that lots of devices (fans, disks, probably all power systems - definitely including PSUs) have lifetime limits in power cycles. Mostly this is minor, unless you do something like reboot cluster nodes after a job (concievably a paranoid security requirement), or some automation gets in a loop and continually zaps a server.
!!!! X64 Exception Type - 12(#MC - Machine-Check) CPU Apic ID - 00000000 !!!!
My story: I had an Intel NUC running Linux back in the day, which would get stuck in standby such that I had to remove and replace the CMOS battery to get it to boot again! I never figured that one out...It too, 3 reboots to clear up the errors. Generally on the linux system one extra reboot was necessary about half of the time.
You want your cold boot to truly be a cold boot.
The subject was a cheap little black & white TV set that my folks had. Dad was an amateur radio operator, who mostly built his own equipment. He could have dissembled it, traced circuits, and calculated the wait time if he'd cared to.
Like, at all. Would just hang when you tried. Couldn’t exit from BIOS after changing settings, couldn’t suspend to RAM. Had to yoink the cord whenever I needed to restart. Wild stuff.
Perhaps like Frankenstein the lightning was a breath of life, and with its new sentience my PC was trying to preserve its existence. At any rate I reflashed the BIOS after a few months and it never happened again.
Standard troubleshooting before getting the vendor involved or replacing parts is to action a power drain.
BMCs with high uptime can be especially prone to this, often forgetting how to talk to the system they’re attached to.
Given the problem occurred only once, I didn't do any more investigation on why.
Also, a topic which can spur some interesting comments.
You unplug/plug it, cold boot it, and then it works again.