Reverse engineering Dell iDRAC to get rid of GPU throttling
github.com
github.com
I like the idea, it is a small computer that is used to monitor and control your big computer. But hate the implementation. Why are they all super secret special firmware blobs? Why can't I just install my linux of choice and run the manufacturers software? This would still suck but not as bad as the full stack nonsense they foist on you at this point.
Definitely an area where a more open ecosystem would improve the pace of innovation.
GPL can force manufacturers to cooperate with users. Of course, they can still use closed source binary modules and userland programs ...
The BSD "freedom" is not for the user of the software, it's for corporation to take.
The original BSD release will stay open and free for everyone to utilize.
If it's GPL you can at least interrogate the code release and make an informed decision
https://airbus-seclab.github.io/ilo/BHUSA2021-Slides-hpe_ilo...
I agree with you that you should be able to run whatever since in the end it’s just another computer, but the manufacturers believe otherwise since there’s “valuable IP” or whatever nonsense (insert rollseyes emoji here).
There are open specs like redfish but still doesn’t get to the heart of the matter.
AMI sells a bmc software stack, https://www.ami.com/megarac/ Intel and small manufacturer were unhappy, about always paying the ami tax. So intel created openbmc, as a hedge against ami's monopoly for small manufacturers. I have heard Openbmc has user from facebook, google, ibm, bytdance, and ali.
Dell owns their own stack in idrac, I have heard most of their systems are nuvoton based. I am suspect dell pays some big bucks to keep their systems at feature parity with the other options, and they view it as a an investment.
There are also silicon devices on the motherboard, that have drivers that are not able to be shared. So it not surprising that companies don't share source in a way that would be useful.
If you wanted a system that as a bmc that could be tested try the asrock-e3c246d4c, it looks like there are hobbyist, that have it running coreboot, and openbmc. (impressively)
https://github.com/AspeedTech-BMC/openbmc (or see the upstream OpenBMC tree)
HPE are adding some OpenBMC support https://www.hpe.com/us/en/compute/openbmc-proliant-servers.h...
This work from Supermicro looked promising https://lore.kernel.org/openbmc/CACPK8XdE0sRmt4x54YJVJO2wDT5...
BMCs lack the security fundamentals and often behave like cheap IoT knock-offs. They often use outdated kernels, libraries, and security mechanism.
Woe betide you if you run into one of the BMC implementations that shares a host network interface; no separate cable! These things are terrible from a security standpoint.
I mostly work in the cloud now but when I last had to manage a bunch of physical machines we had a physically separate network accessed via its own VPN to get onto the BMCs. Because yeah, the security situation was a joke.
Id put money on there being preauth vulnerabilities in those things, judging by the engineering quality.
On a much less-important note, it might explain some weirdness I'm seeing lately with one of my home Supermicro servers. (The docs say the BMC should only listen on one port, but the switch still sees some degree of responsiveness on the normal non-management port when "off".)
Try:
`sudo lspci -vv | grep -P "[0-9a-f]{2}:[0-9a-f]{2}\.[0-9a-f]|downgrad" |grep -B1 downgrad`
Besides the speed, you can have another problem with lanes limitations.
For example, AMD CPUs have a lot of lanes, but unless you have an EPYC, most of them are not exposed, so the PCH tries to spread its meager set among the devices connected to your PCI bus, and if you have a x16 GPU, but also a WIFI adapter, a WWAN card and a few identical NVMe, you may find only of the NVMe benchmarks at the throughput you expect.
Most AM4 boards put an x16 slot direct to the CPU, and an x4 direct linked NVMe slot. That's 20 of the 24 lanes; the other 4 lanes go to the chipset, which all the rest of the peripherals are behind. (There's some USB and other I/O from the cpu, too). AM5 CPUs added another 4 lanes, which is usually a second cpu x4 slot.
Early AM4 boards might not have a cpu x4 NVMe slot, and those 4 cpu lanes might not be exposed, and the a300/x300 chipsetless boards don't tend to expose everything, but where else are you seeing AMD boards where all the CPU lanes aren't exposed?
I'm sorry, I oversimplified, and said "most of them" while I should have said "not all of them" as 20/24 is more correct for B550 chipsets (the most common for AM4) instead of trying to generalize.
Your explanation is more correct that mine.
For anyone who might want extra details about the number of lanes per CPU, https://pcguide101.com/motherboard/how-many-pcie-lanes-does-... is a good read that shows the difference for APUs.
Lanes behind the chipset are multiplexed, and you can't get more than x4 throughput through the chipset (and the link speed between the cpu and the chipset varies depending on the chipset and cpu). But that's not a problem of the CPU lanes not being exposed, it's a problem of "not enough lanes" or more likely, lanes not arranged how you'd like. On AM4, if your GPU uses x16, and one NVMe uses x4, then everything else is going to be squeezed through the chipset. On AM5, you usually get two x4 NVMe slots, but again everything else is squeezed through the chipset; x670 is particularly constrained because it just puts a second chipset downstream of the first chipset, so you're just adding more stuff to squeeze through the same x4 link to the CPU.
Personally, I found that link to be more confusing than just reading through the descriptions on wikipedia for a particular Zen version. For example https://en.wikipedia.org/wiki/Zen_3 ... just text search in the page for "lanes" and it explains for all the flavors of chips how many lanes, and how many go to the chipset. Similarly the page for AMD chipsets is pretty succinct https://en.wikipedia.org/wiki/List_of_AMD_chipsets#AM5_chips...
Mine just go to second NVMe weirdly enough.
example from my X670E board
* first NVME = 4x gen 5
* second= 4x gen 4
* 2 USB ports connected to CPU (10/5 Gbit)
and EVERYTHING ELSE goes thru 4x gen 4 PCIE bus, including additional 3x nvme, 7 SATA ports, a bunch of USBs, few 1x PCIE ports, network, etc.
You can check if it's active using `nvidia-smi -q | grep Slowdown` as shown in the post
This article doesn't mention at all what the max TDP of each gpu is, which makes me suspicious. Or things like max tdp of cpus (such as when running a prime number calculatio multi core stress benchmark to load them to 100%) combined with total wattage of GPUs.
If you have never built an x86-64 1U dual socket server from discrete whitebox components (chassis, power supply, 12x13 size motherboard, etc) this is harder to intuitively understand.
I would recommend that people who want four powerful GPUs in something they own themselves to look at more conventional sized server chassis, 3U to 4U in height, or tower format if it doesn't need to be in a datacenter cabinet somewhere.
This is for example present on thinkpads, and while you could patch the bios before, Intel bootguard now prevents you do that "for your own protection" :)
I hope the MSI leak contains actual bootguard keys for intel 11th gen+, and can be used to allow "unauthorized" PCI modules on modern thinkpads!
Could you please explain which serial? Can you do dmidecode and tell me which Handle/ UUID is all zeroes?
Even if it may not directly apply to current thinkpads, it implies the UEFI module might have other conditionals before going on checking the positive list - something that should be easy to check by reversing the LenovoWmaPolicyDxe.sct PE32.
Would love to learn how to figure out what the changes are because with my limited knowledge I didn’t figure anything out.
If you want to learn, https://erfur.github.io/2019/03/28/down_the_rabbit_hole_pt3.... is a good guide!
It's for an X230, I wanted 802.11ac wifi in it.
> The automatic system cooling response for third-party cards provisions airflow based on common industry PCIe requirements, regulating the inlet air to the card to a maximum of 55°C. The algorithm also approximates airflow in linear foot per minute (LFM) for the card based on card power delivery expectations for the slot (not actual card power consumption) and sets fan speeds to meet that LFM expectation. Since the airflow delivery is based on limited information from the third-party card, it is possible that this estimated airflow delivery may result in overcooling or undercooling of the card. Therefore, Dell EMC provides airflow customization for third-party PCIe adapters installed in PowerEdge platforms.
You need to use their RACADM interface to update the minimum LFM for your card.
No upgrading my wifi, thats a nono!
Of course, I can use a USB wifi card, no problem. Just the convenient PCI card inside the system that is specifically designed to enable a wide variety of functions aside from wifi that is locked in the BIOS just in case I wanted to use a 3rd party card of a specification that you don't happen to sell.
Lenovo did that to reduce their customer's ability to buy PC expansion devices from other manufacturers, in hopes of making more aftermarket sales in the future.
It was a pure greed tactic.
Are you sure this isn't a 1u rack will fry itself if you put a spaceheater type of thing inside of it?
I see the Z6 is a workstation unit, they're going to be more flexible there.
From the article
I still tell people about capacitor plaque and how we should have class actioned Dell out of existence over it
Sure, they "put up $300 million for repairs" according to https://www.theguardian.com/technology/blog/2010/jun/29/dell..., but I bet the lions share of that went to their largest purchasers, so the people who blew their school's IT budget on Dell computers were just SOL.