Some of the problems:
* RAID cards usually have under powered CPUs, and can easily be a bottleneck. Often the best performance with a RAID card is with the RAID disabled.
* Even battery backed RAM is often pretty slow and on the wrong end of a high latency connection between CPU/RAM and disks. Generally investments in ram or intentlog/writelog/slog is a better investment.
* Metadata is often undocumented, often needs backed up via obscure device dependent methods, and is required for a RAID adapter failure
* The firmware is often buggy, especially in the handling of the numerous failure modes.
* Recovery often requires the same card with the same firmware and a copy of the backed up metadata. Not all cards can recover all metadata from just the drives.
* Often are "too" smart, and won't export raw drives for use with more advanced filesystems like btrfs or ZFS. Some require ugly work arounds like exporting each disks as a single disk RAID0.
* Often RAID cards hide the SMART info from the OS, often crippling the ability to predict drive failures. Or hides the functionality behind a weird software stack that assumes a SMTP server and integrated poorly into whatever monitoring/alerting system you use for the operating system.
* Some RAID cards require a network connection and run a buggy and insecure out of data web stack on some undocumented and rarely (if ever) patched CPU that was obsolete the day it shipped.
I have been running production servers with LSI hardware RAID cards for the past 15 years and have not experienced the items you note. The only thing I did experience was a bad RAID card (solved by replacing it with a similar LSI model).
To contrast your "problem" list:
* Compared with MDADM (mirrors, RAID-5, RAID-6), the LSI RAID cards tend to be on par as far as performance (esp with BBU)
* A failed drive won't prevent the array from coming on line (in my experience)
* All my servers have battery backed RAM; never had an issue with slowness or high-latency. Can't say the same for ZFS.
* Never, ever had a problem with metadata on the disk. Not sure why this seems to be an issue
* LSI cards have pass-thru mode and can easily be used with BTRFS/ZFS. I have done lots of tests; never an issue
* LSI cards have their own "patrol read" mechansim to scan drives for bad sectors and send alerts when something seems wrong
On a positive note: * Hard drive failures are truly plug-n-play. Send a tech out to the cabinet, spot the RED light, replace drive. Done. No OS work, no console, no crash cart needed.
* You can logically divide the drives using the RAID manager tools just like Linux MDADM, ZFS, BTRFS, etc. Easy.
* You don't need the specific card to replace a failed card. You just need a card that can read the metadata on the array to bring it back to life.
* Lots of monitoring scripts available in the wild to get RAID stats, rebuild times, etc.
I am not saying hardware cards are indestructible. But, all in all, hardware RAID cards are solid devices that have been in many, many production servers all over the world. I certainly would not dismiss them in favor of ZFS RAID (which tends to be slow for many operations). Pick the best tool for the job.Generally I (and colleagues) consider a hardware RAID failure a nightmare and a SAS HBA failure an annoyance. Doubly so if your storage design included cross connected servers and you just mount the storage from the other server. I've never tried similar with a hardware RAID. Can you cross mount and easily import/export RAID sets between controllers?
The main hardware RAID performance issues I see is with higher drive counts or NVME drives when combined with RAID6. Seen any hardware RAIDs and can manage a few GB/sec with 2 disks of redundancy? 3 disks? Even spending a few $k on hardware RAID seems to lose to a random 5 year old server using ZFS or software RAID by a large factor. Seems like even pretty old x86-64 servers manage a few GB/sec per core.
Careful on the passthru, I had many generations of LSI cards that worked, and 3ware and Areca before that. However I've been hearing that the new LSI RAID cards lack passthru/JBOD mode. I've seen reviews and complaints about the lack of it. One case had involved a LSI hardware RAID connected to a Dell JBOD array, but maybe it was specifically crippled by Dell.
I avoided ZFS for years, generally slower at the relevant workloads (used IO logs collected with systemtap and representative loads created with FIO). However once I added cache and the benchmark included multiple streams of writes, ZFS was a huge win. In particular 64 or more sequential write streams to 120 disks. In fact ZFS did better with 3 disks of redundancy that other filesystems did with 2.
In that past I had more luck with ZFS or MDADM for doing things like /dev/sd[ab]1 for a 32gb RAID1 for boot, then /dev/sd[abcd]2 for RAID5 or RAIDz2. Sounds like the hardware RAID tools are get better.
I've heard of some hardware RAID setups managing to recover hardware RAID metadata with a new card, and some not. Even cases of successes and failures from the same company. Even cases where support claimed it would work, but then decided the card and/or firmware wasn't close enough. Sure decent support can overnight a card, but nowhere near as nice as just being able to mount the drives on any Linux box.
Thus far, we don't have any high-density NVMe servers so I can't really test high-end servers. My experience is limited to 8x SSDs or 16x spinning drives in our arrays.
Finally, I avoided ZFS for a long time as well. Early on (OpenZFS 0.6, 0.7, etc), I spent an enormous amount of time trying to tweak/tune the system just to get on par with our HW RAID devices. No matter what I did, I simply could not get the server (typical Linux NFS NAS) to get any decent performance. In fact, I clearly remember HW RAID hitting 1.8GB/sec on our SSDs while OpenZFS could only get around 400MB/sec. Only until OpenZFS 2.1 mark did I see any real performance gains.
...Speaking of ZFS... I recently worked on a project to replace XFS with ZFS for Postgresql servers. I learned a lot about OpenZFS - specifically the memory latency when using compression. I had a lengthy discussion with the OpenZFS "gang" (https://zfsonlinux.topicbox.com/groups/zfs-discuss/T5122ffd3...) to identify some large latency numbers. Turns out, you need to disable ARC compression and "ADB Scatter Enabled" otherwise you will definitely hit some performance issues. Take a look at that thread - especially the very end where I publish some tuning suggestions. It was a fun learning exercise :-)
I find it useful to implement some sort of tag or link to wiki be added to alerts information [like on failed drive in your case] - makes much easier to every team member to fix the issue without guessing/invent own way during the outage.
Do you have any experience with hardware RAID provided by motherboard manufacturers like Asus? I have a spare dual xeon system I may utilize for my next NAS build and I'll pair it with two more RAID cards that can be put in HBA.