It can be configured with a hot spare so the "resilver" starts immediately on failure detection, not "when you happen to fix".
It can be configured with a hot spare so the "resilver" starts immediately on failure detection, not "when you happen to fix".
We lost everything on a FreeNAS system once because we were alerted to a failed-disk error, but decided to wait until the next week to replace it. But then a second disk failed, and we lost the storage pool. Lesson learned!
But this was an operator error, not a system problem. FreeNAS has been tremendously stable and reliable for us.
We buy a batch of 20 drives all at the same time, and they are all the same manufacturer, model, size etc... Possibly even from the same batch or date of manufacture. Then we put them in continuous use in the same room, at the same temperature, in the same chassis. Finally they have an almost identical amount of reads/writes.
Then we act shocked that two drives fail within a short interval of each other :)
Now its 18 disks (3x6 raidz2) from 3 different manufacturers and every vdev has 2 of each. And the vdevs are physically evenly spread throughout the case.
I sleep so much better. It was kind of a miracle the first setup survived the 4.5 years it did.
The nice zfs built-in autoreplace functionality doesn't work on Linux or FreeNAS. You need some scripting/external tooling to do the equivalent. See
https://github.com/zfsonlinux/zfs/issues/2449
https://forums.servethehome.com/index.php?threads/zol-hotspa...
I've had a few drives fail on my ZFS for Linux fileserver and wondered why my hot spares weren't automatically kicking in, and this is why.
On Linux, if you don't use the zed script that's referenced in that Github issue above and just replace a failing drive manually, a hot spare is worse than useless, because you need to remove the hot spare from the array before you can use it with a manual replace operation.