Western Digital HDD boss mentions archive disk drive idea
blocksandfiles.com
blocksandfiles.com
Gets me thinking tho: Why stop at 5.25in drives? Let's go back to washing machine sized units or even bigger. Let's stack platters so dense and heavy that we can use the spindle as a flywheel UPS system too.
Or how about making a truly gigantic tape spool? Perhaps kevlar threads impregnated with iron could serve as "tape" and we could spin kilometers of that onto a spool.
They argued that a 3600 RPM drive was as fast as 5400 RPM 3.5" drive because you're spinning a larger diameter platter. Technically true.
They also claimed that since you could fit more data on a track, since it was larger, you would do less seeking, and that's what makes mechanical hard drives slow.
This didn't stop BigFoot drives from being slower than 3.5" drives that were 3 years older and 1/4 of the capacity.
It's useful if you get a good deal or if you like the unique advantages of tape (smaller medium with decent shelf life), but otherwise the price advantage seems dubious.
A lot of this doesn't apply as you go smaller, but also the drive cost starts to dominate (can't be amortized over many cheap tapes). There's a pretty narrow band where tape is likely to have a real advantage over super-high-capacity disks.
If anything, we're seeing hyperscalers becoming more heterogeneous with CPUs. Google has training and inference TPUs, custom silicon for encoding media, and many more custom CPUs. It makes sense for storage if the benefit is there.
I've watched colleagues in a team adjacent to my own work through these issues. Have you? Slapping a commodity card into a box and loading some commodity software on it seems like a cakewalk by comparison. Obviously people thought it was worth it anyway, but I think you're seriously misunderstanding where the difficulties lie and what kind of resources are needed to overcome them. 200+ engineers would be overkill (unless FB engineers are ~4x more productive than Google engineers) but you'll need other kinds of specialists as well and the installation/operational costs will still be high. It's unlikely to be worth it for an organization much smaller or different.
Aside from the fact that most people using tape for archival storage don't want to pay extra for the read/write heads, SATA interface, etc., there is no reason why you couldn't package all these things into a self-contained tape unit with a small flash disk acting as a small cache and directory listing.
You could definitely package such a thing for consumers, for example, but most workloads there aren't a great fit for the medium. Basically the only thing that makes sense is using it for archival and backups.
Because that gets rid of the main advantage of tape. Which is that the tape-media has no read/write head and is therefore much much much cheaper to mass produce.
---------
In practice, people buy tape-libraries entirely. Like a 3U unit with 50-tape slots + a few drives to read/write to those tapes, and then hook them up to the network.
https://www.quantum.com/en/products/tape-storage/
From this perspective, you buy as many tapes as you want storage (aiming for 500TB? Buy like 40 LTO7 tapes for your library. Aiming for 1000TB? Buy 80 LTO7 tapes for your library, assuming compression of course).
From there, you just read/write to the library, and have the underlying software handle the details, like any other NAS.
I don't know why this seems the standard practice in the industry, but it really annoyed me when I realized a “15TB” LTO-7 tape has actually only 6TB real, “native” storage coz it assumes some average compression ratio.
Why is this acceptable? What if I use the tape to store incompressible data like video and images? Feels like intentional cheating.
When a company is spending >$20k on a tape system, the people in charge of buying it will talk to the sales people, tell them the use case, and get a more accurate estimate.
The fact that buyers may talk to salespeople is really not an excuse for the deceptive behavior.
"well, in our restaurant medium rare means well done and rare means medium rare, but that's okay our customers are well-paid professionals, they'll talk to waiters to get a more accurate picture"
Clothing manufactures also lie about the waist measurements of pants. (Go measure yours and see.)
But I guess work the data density required these days and the extreme closeness of the head to the platters this won't work anymore as dust would get in.
But this is how the first hard drives worked, we had some with our pdp-11 at the computer museum.
So I just buy external $200 USB hard drives.
There is a market for something lower latency, higher capacity, but pay as you go (no upfront cost).
Frankly I wouldn't mind some sort of giant disk drive, vinyl record size, if that makes sense and is lower cost/TB.
The modern tape drives are miracles of precision and the tapes are these cool little cartridges but why? The meat and machines handling tape carts liked the of 3/4in format just fine. With modern electronics for the write heads and slightly less ruinously expensive loading and running hardware we could still get decent data density, even if we have to hunt for the tracks each tape load.
Basically you have a nextcloud instance that you fill up with your data, then you ask for an LTO tape to be written with it. Repeat ad infinitum.
The tapes are yours; when you buy a tape it comes with 3 free operations (read, write, checksum control). If you want your tapes back, we'll mail them to you, else we keep them in storage for you.
As the tapes are written using LTFS, you can easily read them back anywhere with the proper drive.
You only pay a fee for the cloud storage; it also gives you access to your archive's database (what file is on which tape, etc).
We also implemented the ability to turn off entirely (cut the power) disk chassis when unused, to save power (using 60 or 102 drives chassis). So if you have enough data (a few hundred terabytes to petabytes) it can make sense too.
Do you? Virtually every filesystem has some kind of error detection, maybe even error-correction built in.
That seems like a solved problem to me. CD-R solved this by writing Reed-Solomon codes every couple of bytes so that if any error occurred, you could just fix them on the fly. (As such: you could have scratches erase all sorts of data, but still read the data back just fine)
I have to imagine that tapes have a similar kind of error-correction going on, using whatever is popular these days (LDPC?). Once you have error correction and error detection, you just read/write as usual.
-------
If Tapes don't have that sort of correction/detection built in, you can build it out at the software level (like Backblaze used to do)... or maybe like Parchive (https://en.wikipedia.org/wiki/Parchive).
par2 even has options for specifying level of redundancy. I've had good experience in recovering large corrupted files from an external drive - since then, I've incorporated it into the automated backups of my personal infrastructure.
Nitpick: you mean physical encoding, not filesystem.
Not nitpick: it’s nowhere near enough. Blu-Ray bitrot is a huge issue and if you don’t either write out your data twice (to the same disc or to distinct discs) or use PAR2 or similar, your backup isn’t worth the money you paid for those shiny coasters.
Not sure if they really do. We won't know for another 90 years or so :)
Because modern 20TB hard drives already take like a full week to read from beginning to end (aka: time for a RAID6 rebuild), which is too long as it is.
The problem with hard drives is that they need "more read heads per byte" than the current technology. You can solve this by making more hard drives (ex: RAIDing together smaller drives, say 4x5TBs), so you have more read/write heads going over your data faster.
There's multi-actuator hard drives coming up (2 independent heads per drive), which will report themselves to the OS as basically 2-hard drives in one case. I think that's where the storage market needs to go.
------------
Tape is the king of density: having the most bytes of storage on the fewest number of read/write heads possible. But Hard Drives encroach upon that space, especially with like 20TBs on one read/write head.
> Or how about making a truly gigantic tape spool? Perhaps kevlar threads impregnated with iron could serve as "tape" and we could spin kilometers of that onto a spool.
There's no advantage to that. A stack of 300 LTO-tapes will practically use up the same amount of space as one-tape that's 300x longer. (Besides, LTO-8 tapes are 960 meters long, roughly 1km. The idea of pulling on a thing that's 300km long and hoping for it to not rip itself apart is... pretty crazy. There's definitely some physical constraints / practicality with regards to just material science: shear/stress kind of calculations)
The "jukebox" concept... really a tape-libraries (robot that picks tapes out of storage compartments and shoves them into the drive) is the ultimate solution to density.
You'd have (for example) a four-platter, eight head drive, and instead of storing a given byte serially at cylinder 123, head 0, sector 5, bytes 0-7, you'd store it at cylinder 123, sector 5, byte 0, and light up heads 0-7 all at once.
Now, this might have been hard with early logical drive designs that tracked the physical geometry and expected a more serial format, but that was long ago hidden beneath LBA translation.
Maybe there's a power or crosstalk issue with activating 8 heads at once.
Need to 300 TB of capacity?
50 tapes at $50 totals $2,500 + a few thousand dollars for the tape drive.
Or you could buy about 12x 18TB external drives for about $450 each totaling $5,400. About the same.
Current HDD tracks are around 50-60nm wide, and their limitation is not magnetic grain density (those are around 10-12nm for non-HAMR/MAMR substrates and even smaller for HAMR/MAMR substrates.) The limitation is the long (relatively speaking) actuator arm trying to precisely stay on a very narrow data track. There's actuator flexure, disk flutter, and aerodynamic noise, all of which increase with increasing platter radius. This increases your minimum track width, and minimum disk thickness/minimum disk spacing and end up defeating the purpose of bigger disks in the first place.
Also it's important to keep in mind that in order to manufacture a drive, the entire drive is written to and read back multiple times, which takes an increasingly long time, creating a huge push for some steps to be combined.
That all being said, I'm not entirely sure where the sweet spot in terms of data density per unit volume. It might be slightly more efficient with 5.25 drives. But I can assure you that the record sized or washing machine sized HDDs of olden days are not practical anymore ...though they would be dope if they did exist!
I want a 1TB optical drive with ~100MB write speeds which supports incremental writes. I know I'm asking for a lot but if you could give me the ability to buy a 1TB disc for $1 or less I'm all over it. Archival discs that last for 50+ years would be a huge bonus as well. Perfect cold storage solution.
I'd buy a BLu-ray drive RIGHT NOW if the media didn't cost $66 for a 5 pack of 100GB discs or $90 for a 50 pack of 50GB. That's more per GB than SSD or spinning rust. If the cost of the discs was 1/10th I'd already own one.
> Wikibon argues the cross-over timing between SSDs and HDDs can be determined using Wright’s Law. This axiom derives its name from the author of a seminal 1936 paper, entitled ‘Factors Affecting the Costs of Airplanes, in which Theodore Wright, an American aeronautical engineer, noted that airplane production costs decreased at a constant 10 to 15 per cent rate for every doubling of production numbers. His insight is also called the Experience Curve because manufacturing shops learn through experience and become more efficient.
[…]
> “Wikibon projects that flash consumer SSDs become cheaper than HDDs on a dollar per terabyte basis by 2026, in only about 5 years (2021),” he writes. “Innovative storage and processor architectures will accelerate the migration from HDD to NAND flash and tape using consumer-grade flash. …
* https://blocksandfiles.com/2021/01/25/wikibon-ssds-vs-hard-d...
Interesting prediction.
Additionally, lightly-used flash has error rates that are orders of magnitude smaller than flash that has reached or is approaching its write endurance limit. Which is why an archive-only SSD can very reasonably be expected to provide unpowered data retention far in excess of one year.
[1] For enterprise drives, the specified duration is shorter but the storage temperature is higher. The warrantied write endurance is also typically higher, so those drives are willing to take their flash to a more thoroughly worn-out state.
That's not how flash storage works. You can unplug a flash drive and put it on a shelf and it will keep its data. Here's a review of portable, external SSDs that don't lose their data when not plugged in:
* https://www.tomshardware.com/reviews/best-external-hard-driv...
Are you thinking of DRAM perhaps?
Could not the same be said of the magnetism of the bits on spinning rust? What's the shelf life of data on an HDD?
Tapes also have magnetic charge, but are designed to be "unrefreshed" for longer periods of time.
Flash has a better insulator than DRAM, but it wears out with use and leaks more with higher temperature, densities, less margin with more bits per cell...
Please see [1] for one source.
[1] https://www.ibm.com/support/pages/flash-data-retention-0
The fact that the SSD controller has to do anything at all to read data that was stored only a year ago is a hint that data retention on a SSD requires some (powered) effort.
- a switch from storing charge in a conductive floating gate structure to storing charge in a non-conductive charge trap layer (Intel's flash is the one holdout here)
- a switch from planar to 3D fabrication, allowing a huge one-time increase in memory cell volume and thus the number of electrons used to represent each bit, and also opening up avenues of scaling capacity that don't require reducing cell volume
- dedicating far more transistors to error correction in SSD controllers, greatly reducing the performance impact of correctable bit errors but also enabling the use of more robust error correction codes
WD develops both NAND ( via SanDisk / Toshiba ) and HDD. They know the roadmap of both HDD and NAND. There is nothing on the current NAND roadmap which suggest we get another significant cost reduction. As much as I want to see 2TB SSD below $99. I would be surprised if we could even get to that point by 2024. Today a portable 5TB HDD cost $129, ( or $109 with discount )
This is similar to DRAM, we might get faster, higher efficiency DRAM. But we are not getting any cheaper DRAM. The price of DRAM / GB in the past 10 years has had the same price floor.
HDD is in similar case, it is near the end of the S curve.
That said Windows and Linux could certainly get some polish when it comes to accessing high latency storage. Opening a high-latency or even missing network share on Windows (still?) causes explorer.exe to hang.
Linux might handle multi-second IO OK but I have had cases on extremely bad flash USB drives (hacked a bunch of write-once drives) where the IO is blocked for minutes after the write appears to have finished.
In any case it would be really could to throw bcachefs on a HDD drive like this with a cache SSD.
It's not like Windows doesn't have good APIs that allow you to do better. explorer.exe (or at least the file explorer part) is just an objectively bad product. It wasn't that long ago that it didn't allow you to create files starting with a dot, despite that being an entirely legal path name. And it still doesn't support long paths or alternate data streams, making it pretty easy for software to create files that you can't view in explorer.
I'm kind of surprised that it doesn't seem to have any popular alternatives, considering how consistently bad it is (outside of small subsets, like compression or file copying)
"two major problems of traditional storage devices: data density and durability. One of the densest forms of storage is tape cartridges, which house about 10GB/cm3. Iridia is on a path to having a storage device that could store 1PB/cm3 and reduce latency by 10–100 times compared to magnetic tapes. The other problem that Iridia is solving is durability. Rotating disks tend to work for 3–5 years, while magnetic tapes for 8–10 years. Because DNA is extremely durable, Iridia’s technology on the other hand has an estimated half-life of more than 500 years."
[0] https://outline.com/a97xFX
[1] https://www.microsoft.com/en-us/research/project/dna-storage...
EDIT: according to some random website, the human genome contains 3.2GB, and it takes a cell 24 hours to divide. That works out to 37KB/s, which is not very promising.
But even if we define a new standard of (say) a hundred 12" platters in 12" high cases, the fixed infrastructure (r/w heads, arms, case, board) does not scale as well as tape.
Looks like interesting technology but I don't see how that helps here.
* 5.25'' drive which is taller (more platters) and wider (more sectors per platter)
* Slow rotational speed to reduce consumption and vibration
* SMR again, but label the products accordingly
* Small (64-128GB) SSD embedded, acting as transparent cache especially for quick response to write commands.
* Possibility to disable this caching layer with a SATA command.https://devnull-as-a-service.com/ there you go, robust support for the write once read never use case.
Might be an excellent choice for the Iron Mountains of the world, especially for long-form media storage, though I think that the majority of personal long-term storage is actually shrinking, in terms of growth rate.
https://www.backblaze.com/blog/open-source-data-storage-serv...
https://www.backblaze.com/blog/wp-content/uploads/2016/04/bl...
Edit: My mistake! I was confusing 5.25 form factor with 3.5 :/ much shame.
Background: I used to work on a very large storage system at Facebook, though the one most relevant to this discussion belonged to our sibling group. I've also missed any developments in the year-plus since I retired.
Probably including taping, which most non-enterprise folks are often surprised still exists.
There's an upfront cost for the infrastructure (drives, usually robotic libraries), but once you get to certain volumes they're quite handy because of the automation that can occur.
Tapes are suitable for tape-oriented async-retrieval products (not sure if any Clouds have one?), or for putting _some_ replicas of data on as an implementation detail if the TCO is lower than achieving replication/durability guaranteed from HDD alone. But that still puts a floor on the non-tape cold bytes, where this sort of drive might help.
I'd be willing to use that over huge HDDs. Give me 1TB platters at 5-15€/platter (consumer HDDs come close to €18/TB for large capacity). Actually, I wouldn't mind having them more expensive than HDDs per TB, as I wouldn't have to pay €250 at a time for bulk capacity upfront.
I nearly calculate the cost of using tape every year and it's just not end-user useful :-(
The drives are too damn expensive and probably not that nice to use.