OP described going from 4x disks, add 3x, end up with a 7x raidz2 pool. The steps are
4/4 (steal disk, set up new pool) -> 3/4 + 4/5 (copy and destroy old pool)-> 4/5 (add missing disk) -> 5/5 -> 7/7.
Your suggestion is to add 4x disks, so you can do
4/4 (set up new pool) -> 4/4 + 4/5 (copy and destroy old pool) -> 4/5 (add missing disk from old pool) -> 5/5 -> 7/7 + spare (hot or cold?).
is that a suggestion or do you meam you never used it? because lvm basically uses mdadm under the hood.
Example: I have 5 XFS volumes across 45 drives on a 4U NAS box with 1 cold spare. Granted, I probably should've chopped them up with a single-box hypervisor RAIDless Ceph cluster of VMs in retrospect.
With your extra drive solution, I still have to recreate all my datasets and shares, whereas in my solution, they migrated intact, and I still had backups in case of pool failure. I could zfs send into a giant file on the 18 TB drive, but I'd be reticent to do that because it's just an opaque file that I can't verify will successfully restore. Whereas with my solution, I had the two pools running side by side and could verify everything restored successfully onto the new pool before blowing away the old pool.
Why did you pull a drive from your RAIDZ1 to purposefully degrade it?
I am not sure why you keep saying 18TB. Your drives are 8TB. I am suggesting that you should have simply bought another 8TB disk so you wouldn't have to degrade your RAIDZ1.
Yes, I agree I could have reduced the risk of pool failure if I bought an extra 8 TB disk and not degraded the pool.
So, it came down to do I for sure spend an extra $120 on an extra drive or do I just take the <1% chance that one of my three other disks fails in a 6-hour window while I'm migrating data. I took the chance, but I also had my data backed up in cloud storage at the file level in case there was a pool failure.
In other words, it wasn't worth $120 to me to avoid a <1% risk that I'd have 8ish hours of hassle of recovering from cloud backup after a pool failure.
I have to point out that the surprise risk of the $300 bill from Wasabi dwarfs the cost of the 8TB extra disk.
In retrospect, I would have paid the money for the lowered risk, but everybody has a different tolerance for that.
Again, great work and very detailed.
Raidz3 (survives 3 lost drives) is pure insanity in terms of computational demands. Better off splitting up pools, using multiple systems, or using multiple copy systems up the stack like with Ceph on multiple, simple RAID1 or RAID10 storage nodes.