HGST gets closer to shipping 10TB HDD
zdnet.com
zdnet.com
We're nearing the point at which the throughput relative to the capacity and the risk of failure or corruption makes further capacity less and less useful, unless you're doing some pretty sophisticated replication, validation, & caching.
Naively I would think that it would simply have to stream large chunks of each drive in parallel, apply the parity coding and write out the stream of calculated blocks to the new disk.
Of course, that's in the context of an actual RAID array (or ZFS pool, etc.) - if you've got a replication strategy that spreads multiple entire copies over multiple systems, you're less vulnerable and these kinds of disks are a good choice.
Of course having no redundancy at all during a rebuild that takes more than a negligible amount of time is still worrisome, but that's why we have multiple-redundancy raid schemes.
But you still have to untangle your concerns. A single point (or cause) of failure has to be handled differently than increasing probabilities of encountering independent UREs on separate drives during rebuild.
These drives (like the Western Digital Purple) are used predominantly for archival data. You should never have to randomly write to, for instance, surveillance video, or similar streamed media.
This isn't for your SQL database server.
Two disk failures in the same array within 15 hours is very rare.
And if it can't recover it? You can prove it by failed checksum check. This is why, in my opinion, ZFS is so much better than all other file systems: if it fails, you can prove it instead of just sorta vaguely questioning everything because you can't tell but you suspect it, and it slowly drives you insane.
Once you try ZFS, you never go back.
(We didn't lose any data - everything is replicated elsewhere in addition to the raid)
Now, what does this mean in, exactly? In that four drive scenario, RAID10 has a 1/3rd chance of failure (both failed drives being both the A or B parts of the inner RAID 0s is two out of six possible scenarios), RAID5 will fail no matter which two drives, RAID6 will survive no matter which two drives, ZFS's RAID-Z3 (three parity RAID5/6) will survive no matter which three drives, and RAID1 will survive no matter which three drives.
RAID10 is the most performant of the possible outcomes, and the usual makeup of a 4 drive array unless you absolutely need the storage, then its usually RAID5; unless you seriously do not care about sanity at all, then its RAID0 or just independent disks.
Now, my suggestion for 15 hour rebuild times? Whatever you do, have a hotspare so it can begin immediately rebuilding and do not use any RAID variant that can't handle more than 2 failures in 15 hours. This means no RAID10 or RAID5, only RAID6 or RAID-Z3.
Ceph mitigates this 15 hour time because it can rebuild a lost drive by just allocating new blocks on every other drive simultaneously and maxing out your network and/or storage IO (whichever is slowest) so the window of potential doom is much smaller (depending on Ceph cluster size, obviously, bigger the better in this case).
They're just so rare that the difference between ~6h rebuilds with 4T spindles versus 15h rebuilds with 10T spindles simply doesn't matter. Any homeopathic advantage of slightly faster rebuilds is dwarfed by the benefits of higher density (less hardware per capacity).
Natural age failures are pretty rare.
Also, their drive failure is 10 disks/day out of a population of 44,100 drives, about 0.02%.
So, between the lack of correlation of failure (their Nodes, or, "Pods" are in different racks), their ability to lose an entire POD regardless (they have 3 parity pods), and the relatively low disk failure rate - the large drives aren't a problem. A rebuild in 2-3 days is more than fine.
Simple structures, with a large safety margin and simple control flow vs complex structures, very light weight but at the limits of capacity.
Like an ethernet switch doesn't need complex flow routing algorithms when the raw backplane bandwidth is some multiplier higher than will ever experience contention.
It always seems to me, that things built near their design limits will eventually expose a catastrophic flaw.
Human engineered items that have stood the test of time, Roman Aqueducts, Brooklyn Bridge, DC3, Dodge Dart, Toyota Corolla, AK47, PDP8 all have commonalities in their design philosophies.
Well there's one issue that another whole drive will fail and you're screwed.
The other issue is that modern disks have an unrecoverable read error rate compared to their size such that a total cover-to-cover read -- necessary on every remaining disk to rebuild a RAID5 -- is kinda unreliable, even with a supposedly healthy disk.
http://www.enterprisestorageforum.com/storage-hardware/selec...
http://www.techtravels.org/amiga/amigablog/?p=280
This is probably because Amiga didnt have real hardware floppy-disk controller, just a general IO (CIA) chips, and read raw serial datastream into ram, all the decoding was done in software. Similar to Apple II 2x 8bit XOR checksum.
You do know that in a RAID6 two drives may fail without causing data loss? In any way you need one or more hotspares ready plugged in in order to keep the time window short.
RAID6 is not perfect but the probability of 3 drives failing during the rebuilt window is much lower than the probability of 2 drives failing. (One also has to consider errors such as memory corruption, chip failures or catastrophic failure to the power supply where no RAID level will protect you from.)
A RAID1 built of two RAID6s may be necessary to avoid performance drops during rebuild. In the case of multiple failures a RAID6+1 setup will protect you from at least 4 hard drives failing in the rebuild time window.
> The other issue is that modern disks have an unrecoverable read error rate compared to their size such that a total cover-to-cover read -- necessary on every remaining disk to rebuild a RAID5 -- is kinda unreliable, even with a supposedly healthy disk.
This is another reason for a RAID6. Not only does it recover when one or two disks fail. It also recognizes and recovers broken blocks when one disk returns the wrong data. You scrub the disk weekly and remap broken sectors or swap out (soon to break) harddrives.
Also, does Ceph distribute it's objects so that two drives don't contain the same set of objects? I.E. it's probabilistically impossible that a number of drives going down in a large array can wipe out all the copies of the object?
Yes, you create a CRUSH map which lets you define a hierarchical list of bucket types that reflect your failure domains (host, chassis, rack, room, datacenter, etc).
http://ceph.com/docs/master/rados/operations/crush-map/#crus...
And yes, in the clusters I've built I've always calculated what the chances are that X simultaneous drive failures will take out any data, and it's always been astronomically low.
I've mostly seen people use triplicate pools. For semi-warm storage I've been testing erasure coded pools, with a triplicate cache tier on top, and had good experiences on my test cluster.
I've never used Ceph in production, but it doesn't seem unreasonable to expect the sysadmin to keep clock skew in check.
Ceph object-store known as RADOS is production ready and supported by Red Hat [1], in addition the block-device service known as RBD which is an access method to RADOS is also production ready and supported. RBD provides many SAN like features like snapshots, cloning, and soon mirroring. The CephFS (POSIX filesystem access method) is not yet a supported product but is used by several sites in production. Ceph also provides S3/Swift object-store gateway (radosgw) so existing applications can access the object store with a compatible API.
Ceph provides unified storage: providing object, block and filesystem access from a single cluster, so Ceph is the more general purpose technology. Ceph MDS provides metadata service for the CephFS POSIX-filesystem but is not needed if you're not using that feature. For development and testing Ceph will run on a single VM, or more realistically on 3 small VMs so the cluster characteristics can be explored.
[1] alternatively you can buy a production grade fully supported appliance product from Fujitsu based on Ceph known as ETERNUS CD10000. For most deployments Ceph runs well on commodity hardware and the Ceph community provides excellent support too.
DreamHost would beg to differ. They've been using it for their DreamObject[1] for few years. Granted that they have have better visibility earlier because Sage Weil, but I've seen Ceph deployment for production in the wild since mid-2014.
EDIT: And I'm not suggestion it isn't badass, just that the idea Ceph as a first suggestion for general purpose RAID replacement is probably not a good idea just yet, and that there's an alternative.
What John is not talking about is the Object Store (S3 compatible) or RBD (rados block device, EBS like block storage). There's countless people using both Object Store and RBD in production today. Both of these are the most deployed solutions for OpenStack (for storage or object storage). For example the Bloomberg folks use in production.
We use CephFS, and contributed to it earlier, and helped find a lot of bugs in the kernel and the MDS. Last time we uncounted a bug was over a year ago. We currently only use one MDS with a hot spare (backup node actively replaying the log).
What's really missing today from CephFS is the fsck and restore utilities and I believe people are working on them as we speak.
[1] http://www.macrumors.com/2015/03/11/13-inch-macbook-air-ssd-...
http://www.tomshardware.com/reviews/samsung-sm951-m.2-pcie-s...
The way it works is by realizing there are only a few bad sectors on the failed drive, so rebuild that area first, the used the "bad" drive to help rebuild the rest of the array at full speed.
And SSD's only complain if you listen. Does Windows come with a smart checking tool by default? Don't you have to manually install one and check?
I suspect most people with an SSD have no idea where it's holding.
But it doesn't change the "defective by design" nature. It's one thing if you can't do anything about it - but the bricking design is not good. It should go read only on max writes, not brick.
HDDs don't have a set end of life. Instead you use them till the die naturally. SSD you need to be notified when the end of life is.
There is nothing inherently wrong with either way, it's the action AFTER the death (of either) that is the problem: HDDs do they best they can even after a failure. SSDs just die because they are programed to, not because they have to.
Saying: But you are warned does not excuse that.
The plan is to move to JBOD soon and just use software to store each file on 3 random disks across 3 separate servers.
You have to write 256MB to change a single bit?
And you have various zones, and you have to keep track of where data is written, because it's not written in order, and can be on any zone?
You would need some sort of battery backed scratch/cache area to pull this off, so that you don't have to write very often, otherwise I can't imagine performance will be very good.
Or a copy-on-write filesystem. With some tweaking they can essentially fill one zone by linearly writing new copies (and metadata) to it while garbage-collecting (copying GC, with compactation) old zones in the background