A Look at Backblaze’s Toshiba Hard Drives
backblaze.com
backblaze.com
The first NetApp Filer I used had 4 GB drives, total capacity a few hundred GB, don't recall cost (not cheap), in 1997. It was the size of a small closet.
The first EMC I used had drives of size I don't recall, total capacity in the TB, for unimaginable prices, in 2000. It was the size of small room.
We're up to 8 TB drives. In 3.5". For under $300. It's mind-boggling.
Is simply printing and adding a label to each hard drive not enough, is it too error-prone, or what?
So why is this not just part of the "take all drives from large box/crate/etc" process for the drives?
IE whoever is taking them out of boxes puts them down, one by one, in a simple little labeling machine that slaps label on them (and records into a stupid database the label).
If you want something more advanced, labeling machine has small camera that takes picture of top of drive for further identification, you can process and store all the barcoded info that exist in the image (unlike OCR, the barcodes should be 100% accurate)
A very large consumer of drives at a previous employer :-) did this pretty efficiently. When we expanded our cluster for Blekko we did this for the 5000 drives we got from Western digital (well the scanning, we didn't really need an asset tag) and it goes really quickly with a code scanner in hand and a python script recording the values.
But in your article you claim that the unlabeled Toshiba drives lengthen the maintenance time by "a few minutes" every time they fail.
Since all drives eventually fail, wouldn't it make sense to trade those "few minutes" at the end of the lifecycle for a constant 30 seconds at the beginning?
Given that you have multiple vendors and you also seem to have a risk even in handling any drive, I would suggest that asset tagging the disks and using a bluetooth scanner, or android data entry app, would make sense for all your disk assets. You can then automatically document and track the entire lifecycle of a disk live, as it is inserted or removed.
Your refusal to generate your own labels seems strange.
I did a bit of searching (there's a noticeable shortage of photos of Toshiba HDD ends on the Internet...) and figure that it's probably some sort of batch code - the 3 I could find and read were POU34250025173, POU34250027527, and POU37250019573. The one in Backblaze's photo is POU34350016620.
At least the human error part would be mitigated. With some clever connector design, I guess a drive could be labeled in 5 seconds or less. If you do it for all drives before mounting, you'd get a single standard label for every unit.
Please don't. I wouldn't like to be responsible for such a disaster. Those are nice folks.
Performance-wise, I have nothing but good things to say about the WD Black series and recent 7200rpm drives from HGST (formerly Hitachi, now owned by WD). Of course neither is any match to a decent SSD, but the 1TB models are 1/10 the price of a similarly sized SSD :)
For energy efficiency, on the other hand, any 5400rpm drive from WD or HGST will do. They are quiet and reliable.
As for Seagate, I've had at least two of their drives fail on me in the last few years, not to mention they feel significantly slower than similar drives from WD and HGST in typical laptop usage. Even a 5400rpm WD Blue can run circles around a 7200rpm Seagate drive.
<evidence type="anecdotal" />
I thought it'd be interesting how your "1/10th the price" sounded ok but was actually more like 16% - 20% the cost. I'm constantly impressed by how the SSD prices keep falling.
That does introduce a human error based point of failure though - a few from a batch could have their labels mixed up.
Given how blackblaze's setup is described someone powering off the wrong drive or entire node by mistake due to a misidentified non-failed drive will not affect service at all, the built-in resilience to hardware failure will easily cover this, but there will still be some impact if only in the wasted man-time (for the original error and any resulting investigation & relabelling effort) and any light performance degradation as the affected node is brought back into service.
So for the number of drives they use, perhaps the manufacturers labelling the drives in a consistent manner is valuable enough to complain about not being the case.
I wish they fed that to a sound synthesizer with a 'Borg' voice. Coolest datacenter ever.
say -v Trinoids "Failure: Disk 0491/sdag doesn’t contain a valid partition table. Pod0491: Replace sdag (Z252A34AS) with a new 3TB Toshiba DT01ACA300. Reboot Pod0491 and re-add new sdag to sync. Your biological and technological distinctiveness will be added to our own"