Seagate drives by chance?
Seagate drives by chance?
Anybody running 100 TB+ ZFS arrays should really have been aware of this class of problem before deploying (and doubly so if using consumer drives). ZFS protects from bitrot and hardware failures... but it's not magic. If data is written once and not read again for a long time, you won't know it's been sitting on bad sectors until it's too late, and you may end up with multiple failing drives at once if you lose a disk and try to resilver. Far better to check periodically so you can throw bad disks early.
Hopefully they had a good backup strategy!
All that to say yes they lost all that data, no it wasn't backed up. Not critical to the business.
I'm yet to see anyone make a really good set of resources to either watch or read that are not much too technically deep.
A lot of us have "learned the hard way," as it sounds like Linus himself eventually did. I think this highlights an issue with the "learn through youtube video" approach. An internet celebrity may acquire enough knowledge to do accessible demonstrations, presenting totally valid, useful, and correct information in them, and still miss crucial "unknown unknowns" that they simply hadn't encountered in their own research.
It's hard to know what to recommend for a class of issue you aren't even aware of!
[0] https://blogs.oracle.com/oracle-systems/post/disk-scrub-why-...
Edit: wait, I think I get it... "Scrubbing simply needs to read every file from the disk so the RAID layer notices and repairs a URE". So it's to avoid bit rot? I have to say, as someone who has a few TB of personal data on a NAS, bit rot is a bit scary. Data backup is mainly why I bought the lifetime pCloud+encryption package during the last Black Friday. I wonder how (if?) they avoid bit rot?
If you don't know about ZFS scrubbing, but are using ZFS, you may wish to spend some time researching ZFS some more.
> I find it a bit scary that your have to do monthly manual work on your NAS or you'll lose.
Most distros / ZFS packages set up a cron job and e-mail out any errors. You do receive error reports from your NAS, right?
Yes,
it's also trivially to automatize in more or less all situations.
The gotcha is that you have to do it/know about it...
LinusTechTips recently did a video about how they installed ZFS on Linux and didn't have a scrub cron. They started to lose data before noticing.
The problem with by default installing a cron job when installing ZFS is that for a general purpose OS there a good default for when and how often to run it. And running it on the wrong time might even be a major problem.
Through then tbh. having a bad default is probably still better then no default in this case.
> lose data before noticing
Is a bit of an overstatement as they didn't look for quite a while, they also did not only fail to do scrubbing, they also failed to setup automated health checks and reporting.
Turning a non NAS focused Linux distribution into a well working and tuned NAS isn't easy (compared to using a good NAS OS/distribution), but making it somewhat work is easy. Which makes this a pretty common mistake for non-specialized people (i.e. like in their case).
Quite easy for me to setup, even though its my first NAS that I built myself and first time using ZFS. Very surprised that LTT effed that up to be honest.
And yes, Seagate SMR drives (obviously not ideal but I'm cheap)