Of the 17 servers, only three are now remaining where I haven't had to do this. Not all of these were actual SD card failures, some where done preventively. Still, there have been several SD card failures requiring emergency repair work at inconvenient times. There have been 0 RAID failures requiring similar emergency work on the systems where /boot has been migrated to HDD based RAID.
Of course there is no replacement SD card handy with working /boot filesystem. Actually there is no such thing as "handy" in my case. I can actually reconstruct /boot partition on the HDD faster than anybody could go to the datacenter and replace the SD card. And if there were a replacement SD card, I would need to keep it up to date manually every time the kernel or initrd are updated.
I never want to see this kind of setup again, and frankly it felt insane the first time I saw it.
Of course I should have migrated all of the machines off the SD cards by now, but my excuse is that maintaining Linux on these machines is not really my responsibility (it is nobody's responsibility apparently, although many people are interested in keeping these systems online).
When I die I want to have a SD-card shaped tombstone.
I think the parent views on storage is still valid. Do not boot from the RAID volume where you store many terabytes.
I usually have a small mirror SSD for the OS (like 128G or max 256G). Those drives are super cheap, easy to replace.
For the /data I usually use ZFS or mdadm + xfs. Depending on the use case (for example do you want to add more capacity later? ) you can decide.
Using SDcards for server filesystem is negligent at the very least.
1.HA
2.Had an OS that would get fully loaded into memory (think ESX).
The idea was that SD-cards take less power and where cheap enough you could have a whole gang of spares with images ready to just drop in place. I'm not saying it was a good idea, just that I've encountered this more than once.RAID is used for two things:
1. improving the performance of slow disks (at least the read performance) 2. having 1/2 disk fail with your system remaining completely usable, till you replace the faulty disk as soon as possible.
The second point is fundamental to me: there shouldn't be any disruption of service whatsoever, meaning that only the sysadmin should notice the fault (beside maybe a reduced performance of the system since you have 1 less drive). Database transaction that were in progress when the disk did break shouldn't fail, writes/reads on the FS shouldn't fail, the only thing that should happen is an alarm triggered in the monitoring system to inform that a disk needs to be changed as soon as possible.
Having a RAID with manual recovery... it means that you could end up a Saturday evening in front of a computer to bring a system back online, and still some data corruption may have happened.
If it's your own personal box, and you're okay with that fiddling, fine. But if people are supposed to get work done while the admin gets a new disk, it's not great.
Would it randomly pick one of the bootable disks? That way you have an n-1/n chance of avoiding the bad disk?
If a disk with a boot vol failed, it would have a 50/50 chance of still booting depending on which one failed. You'd create a backup grub entry to boot off the other one manually.
Not as transparent as hardware RAID, and maybe now grub is aware enough to auto handle mdadm? I don't know - I haven't done bare metal mdadm for a decade or so.
The same way it has worked for the past 30+ years: the firmware boots from first designated boot device; if that device isn't bootable, it moves on to the next one. The next one boots since it's part of a mirror and has all the necessary data to do so, and in the rare case where the device is "half bootable" one would simply intervene and select the next good bootable device manually.
I have 2 NixOS-based NASes that run ZFS. each one has 3 equal-sized 256gb SSDs in addition to the pile o' spinning rust.
each SSD has a small UEFI boot partition, then the rest of the space is cut in half.
the root filesystem is a 3-way ZFS mirror of the first half of the SSDs. the second half of each SSD is another 3-way mirror, this time as the "special" / metadata device for the main hard-drive-backed zpool.
any of the 3 SSDs can fail, and it will boot up and mount the storage perfectly fine. I could also easily upgrade to larger SSDs, in-place and with minimal downtime (zero downtime, if my case had hot-swappable SSD bays)
it would work just as well with only 2 SSDs, but the incremental cost of a 3rd is small enough relative to the whole that I went for it.
Machines booting from SAN had only one local disk, others would have 3 local disks (small and cheap ones because application data was on SAN anyway).
Main raid would do the job for availability in term of disk failure, 3rd bootdisk did the job in case of user erroror data corruption on the main boot env. It saved our asses a few times, bringing apps back online quickly and saving us from reinstalling or fixing stuff from a tenporary live environment.
That lets you then even hotswap them if you need to while keeping the root filesystem workable.
Servers go boom, refuse to boot, or even get trigger a few alarms. Disk is replaced, server is booted, and provisioning paves the entire thing faster than you'd be able to debug.
I can't knock it, but I imagine I'd be yelling for at least RAID1 if I had to wait for a replacement.