ASRock motherboard destroys Linux software RAID
forum.asrock.com
forum.asrock.com
Unfortunately (and I genuinely hate to be "language lawyering" when someone has lost all their data), the ASRock firmware is doing the right thing according to the letter of the UEFI spec -- see section 5.3.2 of http://www.uefi.org/sites/default/files/resources/UEFI%20Spe... :
> If the primary GPT is invalid, the backup GPT is used instead and it is located on the last logical block on the disk. If the backup GPT is valid it must be used to restore the primary GPT. If the primary GPT is valid and the backup GPT is invalid software must restore the backup GPT.
Good time to promote my videos visualizing how GPT and Linux RAID work :-) I did some cool live visualizations of mdadm, mkfs.ext4, gdisk, injecting errors etc. https://rwmj.wordpress.com/2018/11/26/nbdkit-fosdem-test-pre... https://rwmj.wordpress.com/2018/11/06/nbd-graphical-viewer-r...
So unlike with MBR disks, you can't just nuke the first few sectors and call it a day, you have to wipe the whole disk (or use a GPT-aware tool like wipefs).
It's much easier to use a tool like sgdisk
sgdisk -Z
sgdisk --zap-all
(note the capital Z)I would still recommend having a GPT, though. There's nothing useful to be gained by avoiding it, and we've just seen what sort of failure modes can come up.
Every Linux installer that I've used sets up disk partitions for software RAID volumes. I don't think an installer for a major distro will let you use a bare device, even if you wanted to, and all of the instructions/tutorials I've seen on setting up mdadm include the configuration of partitions.
Seems like the OP went out of their way to configure software RAID in a non-standard manner, without understanding the consequences.
I hate to kick the OP while they're down, but it doesn't look like this is an issue that someone would encounter with a default/recommended configuration.
(Also note I managed to fully recover from this without data loss, and I have backups.)
This RAID setup is already quite old. For newer RAID setups I routinely put the RAID on partitions as you said, after more advice has appeared on the Internet that this is recommended.
However, I didn't find hard arguments for this beyond "if you make the partitions a bit smaller, you can accomodate disks that aren't exactly euqually large" -- certainly I didn't encounter "if you don't do this, UEFI may trash your setup".
It would be inaccurate to say I went "out of my way to set it up in a non-standard manner" though. I didn't use a distro installer, simply because the machine was already installed when I added this RAID to it. I just used straightforward `mdadm --create --create /dev/md0 --level=1 --raid-devices=2 /dev/sdc /dev/sdd` on bare devices, which is a supported use case by mdadm, and much documentation and tutorial material still mention that you can choose bare devices vs partition (probably hard to blame them for that, given that this is the first and only mainboard I have encountered that does this).
I use mdadm RAID1 on full disks on my desktop with the metadata at the end of the disks (superblock version 1.0). Partitioning tools will complain that my backup GPT partition table is invalid and offer to fix it (I always politely decline) but the primary one is fine. My Dell desktop doesn't mind this situation and also does not ever write any data into the EFI partition so it works well for me.
Unaligned partitions is a common source of bad performance, and it's not a thing that (used to) "just work".
Therefore I've seen "don't use partitions" not be recommended, since it's otherwise a source of potential unalignment.
Just wish I could remember where I read that. I remember taking it on board as it made sense and the source seemed very credible.
Is the part you quoted the one that's relevant here though? As per my post, gdisk reports "corrupt GPT" (due to CRC mismatch). In section 5.3.2 you posted, it says:
> If the primary GPT is corrupt, software must check the last LBA of the device to see if it has a valid GPT Header and point to a valid GPT Partition Entry Array.
> If it points to a valid GPT Partition Entry Array, then software should restore the primary GPT if allowed by platform policy settings (e.g. a platform may require a user to provide confirmation before restoring the table, or may allow the table to be restored automatically).
> Software must report whenever it restores a GPT.
So wouldn't the spec-compliant behaviour be to "provide confirmation before restoring", or at the very least, "must report whenever it restores a GPT"?
After cycling to the shop and buying a couple TB disks to make extra backups before messing with headers, I created new superblocks (you can tell `mdadm --create` to pick up existing data).
I then repeated the repairing and subsequent re-deletion by the mainboard on boot around 10 times, in order to ensure it's reproducible before I post on their forums.
I did not spot it outputting any kind of message about it.
interesting comment though thanks a lot. makes much more sense now than reading just the article.
This mainboard's UEFI even has a built-in GUI for filing tech support tickets, for the case things are broken to the extent that you cannot boot.
(This feature is highlighted in the manual, "if you are having trouble with your PC".)
Loaded backup partition table instead of main partition table!
And there's no "backup partition table CRC mismatch".
This tool can be very cryptic to decode the messages, though.
> If the primary GPT is corrupt, software must check the last LBA of the device to see if it has a valid GPT Header and point to a valid GPT Partition Entry Array. If it points to a valid GPT Partition Entry Array, then software should restore the primary GPT if allowed by platform policy settings (e.g. a platform may require a user to provide confirmation before restoring the table, or may allow the table to be restored automatically). Software must report whenever it restores a GPT.
Emphasis on "if allowed by platform policy settings". Booting up the system should require read-only access to the disks and writing a new GPT should be prevented in that case. At the very least after asking fo ruser confirmation, as the following section suggests :
> Software should ask a user for confirmation before restoring the primary GPT and must report whenever it does modify the media to restore a GPT.
In that case, the motherboard uefi neither asked nor reported that the primary GPT was restored.
I had a `dd if=/dev/zero of=/dev/sda bs=1k count=1000` or similar to wipe the disk just enough that the provisioning would be happy... Or, at least, it was happy most of the time.
I bet this was my root cause!
Of course there is also the "nuclear" option of running blkdiscard.
This really discourages me from buying new hardware.
Meanwhile the software embedded in these boards becomes more and more complex with basically a full blown embedded OS, graphical interface etc...
Software shops have a hard enough time writing robust software these days, ASRock probably has a team of a dozen or so glorified interns copy-pasting from stack overflow (or worse, contracting third parties do to it). I'm exaggerating a bit of course but not by a massive amount from my personal experience working with hardware shops.
And perhaps even other people having had a look what the firmware actually does, or fix the behaviour myself (like adding a "Do you really want to?" popup before it wipes the superblocks).
This mainboard was released in Q2 2014. It still has DDR3. Its manufacturer warranty expired last year!
This is hardly representative of new hardware.
My favorite one was when the integrated EFI shell wouldn't scroll, so every new line would just print over the previous one.
The most obnoxious and common is hardcoding the boot path to be "\EFI\BOOT\BOOTX64.efi", ignoring what the boot entry says on that matter.
Tell us what manufacturer does that, so I can avoid ever buying from them.
I mean really, all that was needed was to read the content of the disk, and if you couldn't use it. just print that it was unusable. but somehow someone put in write functions to destroy peoples data because they assumed they knew all cases that would come by. instead of assuming they don't know everything and building a more careful product. the unknown unknown is what hit this developer and user of the product. it's a shame this happens especially with such a framework who is supposed to replace 'legacy' things. if this is a sign of how modern frameworks work for these kind of interactions , i'll take my legacy rubbish over it any day of the week.
ASRock, as you see from this HN posting, has a reputation for not being quite as polished. Probably best to skip for now if you don't want to be messing with potential weirdness.
If you're looking for higher end things, Supermicro boards are good. Recent weirdness - "Chinese installed backdoors on Supermicro boards" - aside that is. ;)
Great detective work. When reporting a bug like this, it's this extra mile investigation that gives a lot of weight/credibility.
I can almost forgive it for OS designed for novices but even then they tend to ask permission - however, putting these assumptions into hardware and not even asking permission - that's a broken piece hardware IMO.
The UEFI bios may be following the spec to the letter, here. If so, the takeaway is to always use `wipefs -a`.
`wipefs -a` is a good recommendation, but perhaps it can delete too much? In my case, I used `sgdisk --zap` to delete specifically the GPT bytes without deleting other data.
My `man mdadm` shows:
RAID devices are virtual devices created from two or more real block devices.
This allows multiple devices (typically disk drives or partitions thereof) to
be combined ...
"Disk drives or partitions thereof".I couldn't spot the mention you're referring to.
0: http://event.asrock.com/tsd.asp | http://forum.asrock.com/forum_posts.asp?TID=5265&title=howto...
If asrock wanted people to contact them about bugs like this they should create a public bug tracker, otherwise they deserve the delayed response time from not getting this via the preferred channels.
I have an X399 Taichi for my TR2. I'm using 2 older U.2 NVMe drives that have a legacy boot option rom that causes the board to hang 9/10 times at POST. The only option is to disable CSM, so it will only try to use the UEFI option rom. This works fine, EXCEPT that the CSM disable is lost every time the board looses power. Eg, that configuration setting is not properly saved to non-volatile storage.
When I reported this to ASrock (via email), I was told in very broken english to "re install windows". This is great advice, considering I'm running FreeBSD, not to mention that the entire issue happens at POST. Sigh.
I'm also using X399 Taichi, it had flashed probably all released non-beta firmware versions except 3.00 and 3.10, with permanently disabled CSM and it had never lost that setting.
[1] https://www.asrock.com/MB/AMD/X399%20Taichi/index.asp#BIOS
FWIW, I'm typing this on chrome running in a bhyve VM, and I pass a USB controller through via PCI-passthru (using IOMMU) for webcam and U2F dongle use by chrome.. The FreeBSD IOMMU driver whines a bit, but I just assumed that was buggy FreeBSD IOMMU support. Eg: "ivhd0: Error: completion failed tail:0x7e0, head:0x0."
> It is important to know that a disk configured to be part of an mdadm RAID array can look like a broken EFI disk.
I hate to say this, but it is obvious that a software RAID like mdadm cannot grab raw disks as a whole. If the disk doesn't have a valid GPT and EFI partition at the beginning how it is possible to boot from GRUB, load initramfs, and then load mdadm?
I think something must be misconfigured at the first place.
These are data disks containing no operating system, so there's no booting from them and GRUB doesn't come into play.
The GPT is only useful when you want to use part of the disk for RAID - partitioning disk into multiple parts where only some of it will be raided. Some people prefer opposite. First raid raw disks and then partition that md storage.
The other reason to put GPT (or old style partition table) is so that no system or tool will mistakenly be helpful with detrcting "uninitiated drive" and then asking to perform quick format.
For example, they have BIOS updates still in 2018, while the vendor of my previous mainboard from the same generation stopped pushing updates in 2016.
Edit: Before you ask, yes the RAM was listed in the compatibility list. Same RAM, same CPU, same chipset on a Gigabyte mobo works just fine.
See other sections in this thread: mdadm itself recommends using a partition to prevent hardware from tampering with the data.
Therefore I don't see any real problem with anything around here.
GPT with 1MB aligned partitions will probably work for the foreseeable future; if for no other reason than (popularity across OSes) inertia and it being a nice binary multiple.
As a note; I refuse to use the silly 'MiB' style syntax. 1M bytes is base 10, 1MB is base 2 'near' to that, since bytes implies the native to the device base 2 alignment rather than the common among humans base 10 default.
why does this random authors of code needs to inform people of the fact their hardware has shitty bugs which are apparently never fixed because 'everyone will read best practices'. That seems problematic to me....
I would expect it sooner from the mainboard manufacturer in their information, - with perhaps a warning and a choice (uefi as gui so why not eh!) before destroying disk content.... (the OPs system clearly just does it without even prompting or giving an indication it will do it.) Before this firmwares modify the disk they could easily prompt to check if it's ok with the data's owner to modify this data.
manufacturers shouldn't rely on assumptions in their code for processes which can destroy someones personal data in any case. there being a warning from some people who dive into disk drives a lot that this kind of problem exists doesn't make that situation any better.
Usually each NVMe drive has own pcie lanes, so having two drives raided gains performance over having one drive. AMD is now really pushing into having bonus lanes, so for threadripper they published reports on 6x raid0 of nvme ssds, where they got 21gigabytes/s reads out of storage.
https://community.amd.com/community/gaming/blog/2017/10/02/n...
Please note, it is now important to check motherboard information regarding pci lanes availability, and ensure that between GPU and M2 there are plenty pcie lanes. Some early motherboards had 4 m2 slots, but only first 3 slots guarantees pcie lanes, while utilizing 4th would force sharing of pcie lanes and be slower than first 3.