Seagate's Roadmap: The Path to 120 TB Hard Drives
anandtech.com
anandtech.com
I guess an even higher level of redundancy will be required at these capacities.
What it interesting about this spec is that it is statistical, meaning that the more times you read the bits, the more likely you are to hit a URE. HDD manufacturers do all sorts of things to avoid them (re-trys, track mirroring, etc) but ultimately it is a mechanical system and there are a LOT of bits.
When I was at NetApp, the company would get reports of UREs that were fixed by RAID reconstruction (so you read a disk, you get an error and you fix it by the ECC in the other drives). And we had determined that by 4TB it would be "stupid" to using mirroring for protection since if a drive failed you couldn't count on being able to re-silver the mirror by getting a clean read of the other half. (something like 1 in 20 attempts would fail and you'd have data loss). And of course reconstructing a RAID-5 group for a failed drive you need to read the other drives, for larger groups of disks (which people did to maximize data storage) you started becoming at risk of being able to reconstruct all of the stripes. That was the genesis of doing dual parity (which Netapp called 'diagonal parity' because of the design).
The ZFS folks have also considered this and designed both dual and TRIPLE parity which makes sense for these larger drives!
It already takes a long time to reconstruct any "wide" vDev in ZFS that is zraid2. So one wonders if you could solve for device capacity where 'reconstruction takes 3 years' (the depreciated life span of a drive at Google).
This kind of holistic systems' awareness is a sign of software that's designed well.
Not sure if they do actually implement such algorithm already.
The solution isn’t more parity disks, it’s to move away from that RAID set up entirely.
I fully agree with this part. De-clustered RAID really should have been mainstream a long time ago. Thankfully dRAID in OpenZFS will bring this to the "masses" in the sense of it being an open source implementation, whereas this has traditionally been a proprietary feature.
> The solution isn’t more parity disks, it’s to move away from that RAID set up entirely.
But I strongly disagree with RAID being the problem. Rather the increasing of capacity with bandwidth not really improving is the issue. Of course this is a inherent limitation of disk drives.
I mean, yes, you could have ssds in raid and not have to worry about the drive failing before it's rebuilt, but we're talking about spinning disks here.
What's the current recommendations around raid capacity before you have to seriously start worrying about drive failure before it can be rebuilt? This is a genuine question I don't know.
That depends on the size of your drives and the number in the array. If you are running a small NAS with only one parity drive, you don't want drives that take more than a day to onboard. That limits you to 4/6TB per drive. If you have two parity drives in the array you can probably risk the 8/10/12 TB size. If you are running commercial-scale arrays of dozen of drives, arrays that can handle multiple failures at once, then the sky is probably the limit.
But the story gets more complicated. Are all of your drives the same age? Are all the drives of the same model/manufacturer? Such things increase the risk ofone failure evolving into a multi-drive failure during rebuild. If/when I build a new array I want drives of different ages. So If I had 10 drives, I would start the array on five or six, holding some for later so that the entire array isn't the same age.
If you have two drives with the same number of platters, the same diameter, and the same rotation speed, the one with the highest storage can only do that by increasing storage density - making the bits smaller. Smaller bits means more tracks, but also more bits per track. So in a single rotation (1/7200 of a second, for instance), the read head sweeps across more data. If it reads it as fast as it sees it (which they do), then that means a higher transfer rate.
1) the rebuild takes time comparable to characteristic lifetime of the drive, which on average is >5years. Rebuild times are more in the realm of days, maybe coming to weeks with the proposed 100TB models. Still far from substantial probability of failure.
2) the drives are very old or have bad SMART data. Then the probability of drive failure and (array loss) shoots up.
What people often talk about regarding big drives and reliability concerns is probability of rebuild failure for RAID5, RAID6, i.e. some bit is read or written wrong and the array becomes inconsistent. This is much more probable but isn't really a big problem in practice because this can be detected and rebuild can be repeated.
As with the shift from 5.25" to 3.5" to 2.5" and even 1.8" drives, the only way out of this bind is really to take advantage of the improved areal density to make drives that are the same capacity but smaller and pack more of those into the same volume or power/heat envelope. Drive manufacturers could help this along a little bit e.g. by sharing a motor and some environmental bits between what are otherwise completely separate drives (including separate external interfaces) within a single package, but mostly we'd all better get used to higher drive counts. Dual actuators - and this is far from the first time they've been tried BTW - are mostly a red herring.
Background: I worked on exactly these problems for the latter half of a thirty-year career, most relevantly at my last job working on an exabyte-scale storage system at a FAANG.
Retailers in my area (e.g. Best Buy) don't stock hard drives larger than 4 TB; I was going to tell somebody who lived in the Valley that he's lucky to be able to go to Fry's and then Fry's closed down.
Those same retailers stock both budget and quality SSD's up to 2TB in size. Most upgraders and the system builders are happy.
For the rest of us there is Amazon where the Seagate Exos "enterprise" drive costs half of what similar "consumer" drives cost, has a great reputation and does not seem hard to live with at home.
I would not take it for granted at all that a backup, RAID rebuild, restore, metadata scan or any full scan would work on such a disk if I hadn't tested it -- it is just that kind of technology. Synology is not crazy at all when they make you buy branded large drives to go in the enclosure.
FWIW, I filled an 8T hard drive and then a 4T follow-up hard drive in just one summer of torrenting. Full-size Blu-ray rips are large, and the canon of great cinema is vast. Sure, it'll take me many months afterwards to watch everything that I have downloaded, but these large hard-drive sizes are well within the bounds of what a cinephile with a home theater building a collection would need, and they definitely just aren’t for large businesses. Of course, the hard drive discussed in the linked article is something else entirely.
Video production & storage.
My company works on VR content. We currently have over 10PB stored remotely and have no less then 40TB of ACTIVE files in use every day.
These numbers grow by 175% monthly.
We are running out of rack space for our file-servers, just trying to KEEP UP with demand.
I wouldn't consider myself an expert on market numbers, but I'd say you generally shouldn't be using the larger drives. As others have said, disk is the new tape. Unless you have the kind of system where you might once have used tape - i.e. one with a substantial ice-cold-data component - larger disks are likely to be an ill fit. Even where they are a good choice, that's mostly going to be in a tiered architecture with flash etc. to suck all the "heat" out of the data going into them.
It’s nice that 2.5” portable USB hard drives are cheap enough now (and USB ports providing enough power) that you don’t need the bulky power supply any longer. That means if you want to do any kind of bulk video storage, you can afford to do it. You can’t with SSD, which costs 10x as much per TB. A factor of 10 still matters to most people, ESPECIALLY if they’re not FAANG.
And I think relying on FAANG infrastructure is over-rated. Still a very good case for local storage. USB now provides enough throughput and power that it’s easier than ever to have a significant amount of local storage for cheap. And carry it with you if you travel.
I can get a used server on eBay with 96 TB of storage, usually in 12 x 8TB configs for around $2000 - $2500.
Flash drives? I'll spend 10 times that... hell, probably more. I love having my movies, music, television, etc. on my home server, but I don't love it at the level of 25 large. Now $2500? That's a lot more reasonable.
Spinning disk is still really useful for video. Which is not shrinking any time soon.
If you want to do video stuff on a laptop you have room for maybe one extra hard drive. And if you are willing to have an external drive (which isn’t too bad), it’s not going to work to plug several hard drives in and expect to RAID them together.
Secondarily, there are only so many drives you can stuff in a workstation. A lot of compact ones only have a couple slots in them, and not everyone wants a RAID or JBOD controller with a bunch of ports on it just to do a video workflow. And if you have a video surveillance small server or appliance, you might only have a few slots in it (4 is fairly common).
So again, even one or two hard drives is fine for most uses. Not everyone is gonna put a 24 slot JBOD/RAID chassis in to just provide enough storage space for their video surveillance system or whatever.
The scenario you describe hardly sounds like "most uses" to me.
I answered.
The home prosumer market. And if anyone ever bothered to develop a multi-terabyte solution for "normal people", whereby they could easily have their DVD / Blu-ray / Compact Disc collection converted from disc to a media server, probably a lot more people.
100 TB hard drives will end up in consumer PCs by 2035, I'm certain of it, even if we won't really need it. 8K is just plain stupid for home use, I honestly don't know why its being pushed, since you'd need a 120" screen or more, but I'm sure there'll be new media types beyond 4K... 4K60 FPS for instance, or high-resolution VR movies. I could easily see a world where 4K60 FPS movies take up 200 GB space each, and high-res VR movies are that large or even larger.
No, I didn't. PaulHoule did.
> 100 TB hard drives will end up in consumer PCs by 2035, I'm certain of it, even if we won't really need it.
I think you're confusing the need for more capacity with the need for more capacity within a single drive. This is the same distinction you'll find in the original RAID papers. With a big single drive ("SLED" in the papers) your performance per megabyte stored and your reliability per megabyte stored keep going down. At some point this will make your system unusable. I contend that we're already on the edge and these newer technologies that only increase capacity will push a lot of people over. Increasing video resolutions only make the problem worse, not better. How many 4K60 streams are you going to get over a single interface?
Now consider an alternative: same amount of storage, across multiple drives. In the same physical space, because that higher areal density can be used to make drives physically smaller just as well as it helps make them logically bigger. In a similar power/heat envelope, because smaller also means lighter. But with better performance across more heads and more external interfaces. And with better reliability because with multiple drives you can add some redundancy (and even if you don't at least a single failure will only cost you some of your data).
Why would you want a single big disk instead? Complexity? It's not actually that big a deal. Non-specialists build such systems every day. The only thing the drive vendors need to do is use those improvements in areal density to make drives physically smaller instead of logically larger. Stop making the capacity/bandwidth gap bigger. Let people build balanced systems that actually work, instead of on systems with a striking resemblance to those used for "virtual tape" cold storage in the bad old days.
How many does a prosumer need? One, maybe two. Maybe 5 in a NAS. A single drive today can already handle that in 8K.
When it comes to media files, we hit storage limits all the time but we're nowhere near the bandwidth limits of a hard drive. We're not on the edge when it comes to performance per megabyte. Program files and game files are miles away on one side, and videos and photos are miles away on the other side.
There's some point where increasing the density of hard drive platters is too slow for a prosumer media library, but I'm confident it's far out there, past a petabyte.
According to my calculations, SATA-3 could support almost four 4K60 streams, but only if the data for those videos was very carefully interleaved (never happens). Also literally nothing else happening on the drive, no bad-block relocation, no bottlenecks elsewhere in the system, etc. Closer to reality, those files would be laid out on different parts of the disk and you'd lucky to get even two concurrent streams without seeks between them ruining your throughput. So yes, you are at the real-world bandwidth limits of a hard drive.
By contrast, two physically smaller drives adding up to the same capacity in the same space could reliably deliver one stream each, plus probably a third with data stored on both if you had decent buffering (because now you have enough MB/s headroom to buffer) to cover the seeks that remain. Just as when I was building storage systems for video professionals in 1994-95, if I had to deliver such a system and it had to work before I got paid I know which way I'd go.
If you have raw 4K footage, and you're editing with it, and you need multiple streams and the ability to scrub around, I would simply say not to use any hard drive.
But the media library would be photos you take, videos you take, DVDs you rip, etc.
I have no idea what fraction of the server market.
But there's also an important thing to note about product families. You talk about using density improvements to make drives smaller, and then install more of them to keep performance up. I think that's reasonable, but I also think that one of the best ways to do that is to reduce the platter count and drive height. In that world, where the main product is thin drives, it takes only a small amount of engineering effort to keep making an XL model that has lower performance but is significantly cheaper per TB.
Yes, size does matter a lot for those markets. If you want performance use flash. However, again, size per drive is a red herring. What matters is the capacity you can fit into a system, whether it's a laptop or a server. Having that much capacity present as a single volume through a single interface is simply not ideal either for performance (which might not be the same goal but still has a lower limit) or for reliability. That's all I've been saying. You're better off combining multiple lower-capacity drives, even if you can get by (at least for a while) with a single larger drive. Serious video folks and even gamers have known the advantages of dual drive RAID-0 or RAID-1 for years.
> where the main product is thin drives, it takes only a small amount of engineering effort
What you're now suggesting is no more than what I suggested nearly a day and several posts ago (look for "drive manufacturers could help"), which you and "others" took issue with. Yes, drive manufacturers can and should make those thin drives, and then sell multiples packed into a single enclosure like we already have today. The fact that it's multiple physical drives could be more or less transparent. The transparent version would be cheaper and offer the system designer more flexibility. The non-transparent version, akin to existing HW RAID or even multiple platters today, would be a bit easier to conceptualize for people not used to thinking of enclosures and spindles and platters and heads as separate things, but it would be a bit more expensive (controller plus memory as part of the package) and not necessarily better.
In short, using "drive" to mean both the package with connectors on the side and the piece(s) of oxide-coated metal inside it is sloppy, and leads to wrong conclusions. Once you realize that higher density creates more options than "every limit the same except for higher capacity" then it quickly becomes clear that 120TB on a single spindle isn't the best use of that technology.
I agree, but the argument I'm making is about cost per terabyte. I'm not inappropriately clinging to terabytes per drive.
> What you're now suggesting is no more than what I suggested nearly a day and several posts ago (look for "drive manufacturers could help"), which you and "others" took issue with.
I didn't take issue with smaller or multi-component drives existing, I just don't think they are necessary for all use cases. I was referring to what you said before on purpose, but disagreeing with the conclusion that "mostly we'd all better get used to higher drive counts". I got the impression you were treating it as a temporary transition measure.
But can be countered by more careful engineering of the device.
And if wishes were horses...
No, you can not simply state some arbitrary level of reliability and design to that level. At any given level of reliability more parts = less reliability and more time or resources spent on design are not by going to remediate that unless you are willing to accept much higher costs.
Engineering is a trade-off and physics determine the sweet spot for the balance of that trade off. Once that sweet spot has been determined all other things being equal more parts will cause your reliability to go down unless you will also accept that your costs will go up and/or other parameters will be affected in a negative way.
Software people in general have a very hard time to understand this because to them 'parts' are free, but even in software, assuming zero costs for parts (adding a library, a function, a line of code) has an effect on reliability. That's why computers used to be much more reliable than they are today, and that's before we get into details such as cognitive load while trying to understand complex systems.
HDD are much faster when they're only half full.
1. Without RAID, drive fails and you're stuck waiting for restore from backup.
2. With RAID, risk double drive fails < risk of #1 and are stuck waiting for restore from backup.
3. With RAID, risk of single drive fails > 0 and < #2 and continue working while waiting on drive clone while still having the backup in #1 and #2 to fall back on.
Maybe the above is wrong because UREs are just the unrecoverable medium errors. I can't find the official definition of URE rate.
If I have a bunch of 8TB drives, I don't care what my odds are of doing a single-drive operation without a speed bump. I'm going to back them all up at the same time, and I only care about speed bumps over the entire operation and how well they're handled.
120TB.. Man, to those in the future laughing at this comment with your 120TB games and your 60TB super-high-definition VR feature length movies.. I'm jealous.
Either you used a very pessimistic estimate for drive speed and forgot to factor in the increase in linear density, or you used a relatively accurate calculation and mixed up megabytes per second with megabits per second.
> If the drives are used heavily while rebuilding, this can go much longer.
True but that's a separate factor and depends on how the array is laid out.
Seagate should sell such a machine capable of running arbitrary Linux distributions as a kind of mini-NAS.
[1] https://www.jeffgeerling.com/blog/2020/building-fastest-rasp...
[1] https://en.wikipedia.org/wiki/File:Conner_Peripherals_%22Chi...
(Decreasing IOPS/TB is also why SATA SSDs haven't gone beyond 8TB while NVMe and SAS SSDs are pushing past 30TB.)
I use HDD for bulk storage and SSD for high-speed storage. I don't mind less speed, except insofar as it impacts reliability (e.g. a drive in a RAID fails, and the redundant drive fails while rebuilding the array).
It seems like competitors to HDD have become increasingly noncompetitive over time. Tape drives are insanely expensive at reasonable capacity. DVDs hold 0.1% of an HDD and are more expensive. Various newfangled optical drives (e.g. M DISC) appear more reliable than HDDs, but also cost much more per TB. Plus, you need to keep a big pile of coasters.
I don't think that will be true. By having the heads only read half the platters, the bytes per track are <areal density increase>/2. That's more seek operations for linear reads and small offsets. For completely random access that's probably the same number of seek operations as before, but queuing latency is lower.
Where it might reduce seeks is where multiple processes are trying to stream data at the same time, which sounds like is happening more and more.
Their goal is to retain that TCO edge by ensuring that HD performance doesn't get too much worse as HD capacity increases.
I could imagine certain resonances (eg. the height of the head oscillating) might differ between the two heads, effectively meaning data written with one set of resonances cannot be read back with another.
IIRC the next step is to put the logic for that in the head itself with a tiny processor so the feedback cycle can be nearly instant. I suspect that would eliminate issues with one head reading something written by another.
I suppose as things like 8k video editing come up and file sizes explode there will be use cases for this kind of density, but without read/write and throughput increasing it seems like it won't be super useful for a little bit.
S3, BackBlaze etc. all focus on cramming as many hard disks in to a single machine as they can do, without running in to other bottlenecks on the machine level (CPU, memory, NIC bandwidth, controller etc).
You very much want to get out of the RAID business in those environments too. Backblaze mention their use of Reed-Solomon which is fairly common on large scale storage, and moves you much closer to resiliency on an individual object basis, rather than thinking in terms of the entire drive.
Consumer-grade HDDs barely managed 100MB/s 10 years ago. Now they can often do 200MB/s, and enterprise disks are even faster. With these much larger Seagate drives I guess SATA will be the bottleneck, not the drive's sequential read/write speed.
SAS is 12Gbps or 1200MBps (I don't know if it uses 8/10b encoding, but I'll assume that for simplicity)
------------
Hard Drives are no where close to breaking the SATA3 barrier, let alone the enterprise SAS 12Gbit barrier.
So we might be headed for what used to be a single write head (and its single throughput stream) to double, triple, or more.
Especially since the additional read heads enable the datacenters to scale shared object storage more effectively with more dense drives, which seems to be the main customer/application for HDDs at this point.
On an unrelated note, are you the same Dragontamer that I've met and played with at PDXLAN? Or do you just happen to use the same alias?
Unlikely me. We must have the same online alias (it is a very common alias in my experience...)
To answer the OP's question, it seems to me that after around 12TB or so, it makes more sense to move away from implementations that require rebuilds such as raid 1, no raid, or jbod solutions.
7200 RPM / 60 == 120 rotations per second. A "half-rotation" to move the typical data on the disk to the head (half the data is within the first half-rotation, the other half of the data is within the 2nd half rotation).
If you want to reach the data faster, you need to physically rotate the disk faster: such as a 10,000 RPM drive, 15k, or 20k drive. To allow for faster rotations, you shrink the drive to 2.5" or even 1.8". Alas, SSDs have taken over this niche entirely, so we only really have 3.5" and 7200 RPM drives anymore.
[1]: https://www.youtube.com/watch?v=yHgSU6iqrlE (presentation)
[2]: https://www.snia.org/sites/default/files/SDC/2019/presentati... (slides, page 41)
Not sure if there will be consumer drives with this eventually or if the cost is too prohibitive.
A multi-actuator drive isn't really "one hard drive" anymore, its really just two hard drives ganged together. While more physically convenient, it doesn't seem to really offer the true 2x increase we're looking for.
Actuator#1 cannot give more IOPS over the data that Actuator#1 is assigned over. You only get more IOPS if you can split the work between the two actuators. Same problem as RAID0 or RAID1 multi-read hard drives (you gotta figure out a way to "split the work" to get RAID0 truly 2x the IOPS).
RAID1 can give you a 2x increase in reads, but suffers even more than RAID0 when it comes to writes.
Dual actuators, implemented in a straightforward way, can both access the entire drive surface which means they can give you a true 2x increase. Sometimes even better than 2x, because each arm can focus on one side of the disk. For read/write workloads it completely outclasses RAID.
That constraint means nothing here. You can issue two parallel reads to two drives in RAID-0 just as easily in RAID-1. The only case where this doesn't work is where you're reading more than 2x the interleave size and you're issuing separate requests for each interleaved chunk. With command queuing, a smart storage system should even recognize the pattern and buffer to reduce the damage, but you'll still pay a cost in extra interrupts and request handling though so it's better to learn about scatter/gather lists.
> they can give you a true 2x increase
I already explained why this isn't actually the case, and have observed it not to be the case with multiple generations of dual-actuator drives. Stop presenting theories based on misconceptions of how disks and storage stacks work as though they were fact.
Under RAID 0, the odds are 50% that two independent reads are on the same drive. It's impossible to get a speed advantage in that case.
> I already explained why this isn't actually the case
You said they "improve parallelism, not media transfer rate or latency", and I'm arguing about parallelism. Plus large transfers can be rearranged into parallelism (fact, not theory).
And you said that they can face internal contention "elsewhere" but implied that could be fixed.
So that doesn't sound like what you said disagrees with what I said.
If you have a single sequential stream, then no. You'll either have parallel reads across the two drives, or you'll have alternating reads that the aforementioned semi-smart storage system can turn into parallel reads with buffering. If you have multiple sequential streams, then it's practically going to be like random access, which you already put out of scope. So there's no relevant case where RAID-0 is worse than RAID-1 for reads.
But you know what will be worse? Dual actuator drives. Why? Because of what dragontamer (who was right) mentioned, which you overlooked: the two actuators serve disjoint sets of blocks. They even present as separate SAS LUNs[1] just like separate disks would, so you would literally still need RAID on top to make them look like one device to most of the OS and above. But here's the kicker: they still share some resources that are subject to contention - most notably the external interface. Truly separate drives duplicate those resources, enabling both better performance and better fault isolation. Doubled performance is an absolute best case which is never achieved in practice, and I say that because I've seen it. If Seagate could cite something more realistic than IOMeter they would have, but they can't because the results weren't that good.
The only way dual actuators can really compete with separate drives is to duplicate all of the resources that change behavior based on the request stream - interfaces, controllers, etc. Basically everything but the spindle motor and some environmentals, as I already suggested now two days ago. You'd give up fault isolation, but at least you'd get the same performance. That's not what Seagate is offering, though.
[1] https://www.seagate.com/files/www-content/solutions/mach-2-m...
> But you know what will be worse? Dual actuator drives. Why? Because of what dragontamer (who was right) mentioned, which you overlooked: the two actuators serve disjoint sets of blocks.
They don't have to do that.
I was talking about what you can do with dual actuators, not product lines that already exist.
I didn't realize how mach.2 was designed, though. That's a shame.
> But here's the kicker: they still share some resources that are subject to contention - most notably the external interface.
Each head, even at peak transfer rate, uses less than half the bandwidth of the external interface.
So even if both of them are hitting peak rates at the same time, and the drive alternates transfers between them, things are fine. For example, let's say 128KB chunks, alternating back and forth. Those take .2 milliseconds to transfer. That makes basically no difference on a hard drive.
> Doubled performance is an absolute best case which is never achieved in practice, and I say that because I've seen it.
I completely believe you, about drives where each arm can only access half the data.
> The only way dual actuators can really compete with separate drives is to duplicate all of the resources that change behavior based on the request stream - interfaces, controllers, etc.
Or upgrade them to 1200Mbps, which isn't a very hard thing to do.
Since you didn't know they're different until a moment ago, you were talking about both. Don't gaslight.
> Each head, even at peak transfer rate, uses less than half the bandwidth of the external interface.
So two will come damn close ... today. With an expectation that internal transfer rates will increase faster than standards-bound external rates. And the fact that no interface ever meets its nominal bps for a million reasons. Requests have overhead, interface chips have their own limits, signal-quality issues cause losses and retries (or step down down lower rates), etc. Lastly, request streams are never perfectly balanced except for trivial (mostly synthetic-benchmark) cases, and the drive can't do better than the request stream allows. There are so many potential bottlenecks here that any given use case is sure to hit one ... as actually seems to be the case empirically. Your theory remains theory, but facts remain facts.
Rebuild time per TB will actually slightly improve because of better throughput due to higher density and higher number of disks inside the drive, so the recovery time for small arrays will actually get better.
True, rebuild time for a whole drive will get very long which is not great, but if the array is designed with good enough redundancy, this won't be a problem, less alone a blocking issue. The very point of RAID is that the system is functional even in the state of rebuilding. If enough drives are used, it does not matter that the rebuild takes 1 month.
Higher areal density won't improve rebuild times unless internal transfer time is the bottleneck (it's not), and it very much does matter if rebuilds take a month. If that additional capacity isn't accompanied by proportional amounts of external-interface bandwidth and CPU/memory somewhere, then bigger disks will mean more risk of data loss. The math is unforgiving.
Of course rare failures and loss of data do happen. There is no storage strategy that prevents these with certainty.
Data loss and performance degradation should be expected and designed for. Maybe RAID6 isn't cutting it for petabyte projects, but it is fine for vast majority of RAID users (small businesses, <12TB arrays).
I've noticed that special hardware and design requirements of the few largest operators are somehow proselytized as a standard that everybody should adopt. People just like to talk about how they understand the biggest deployments in the worlds and how that is the best practice for everybody. But for most users of RAID, these bigboy strategies are irrelevant. Arrays below 12TB are very common and work acceptably well with RAID5 / RAID6, and occasional stripe failure very often isn't a big deal for home users or small businesses.
> Higher areal density won't improve rebuild times unless internal transfer time is the bottleneck (it's not), and it very much does matter if rebuilds take a month.
Why? It matters only if running in degraded state poses performance/reliability problems to users. Which means the array wasn't designed with proper redundancy and performance in the first place. That is the problem, whether rebuild takes a day or a month. Large drives 100TB will be fine if enough of them is used in the array so it works well in degraded state. Also, most probably URE rate will go down due to better ECC measures with 100TB drives.
So one one hand you say that "big boy stuff" doesn't matter to anyone else, but on the other you say that "proper redundancy" requires higher scale. Seems a bit Goldilocks-ish to me, or perhaps even a bit slippery. There's a pretty well established trend, especially in storage, of things that happen in large systems becoming very relevant to smaller ones over time. RAID itself was considered a super-high-end niche once. And don't assume that my knowing about the high end means I don't know the low end as well, or make appeals to authority on that basis. Rebuild times have always been an issue worth addressing, from 1994-95 when I was working on the then-highest-density disk array (IBM 7135/110) to now, from high-end HPC to SOHO. Don't act like you occupy some magical space where what's true everywhere else is not true as well.
I agree with you that in time, the high-end tech becomes the standard tech. But that takes some time. There is quite a non-magical space of small providers who do not care for super reliable storage or super fast rebuilds and this will be the case for a long time. Yes the faster the rebuild the better, and "it is a concern" is fine. One week or month rebuild can be lived with. There is nothing magical about one day, one week or one month. They are all very short compared to typical drive lifespan.
At the same time, yes I believe 100TB drives, if they come, will be used in those extremely reliable big deployments, simply because of better TCO and expansion of data. Even if rebuild times will be longer than today, I believe it can be made to work reliably.
NAND flash data retention is related to how worn-out the flash is, in terms of program/erase cycles. A drive that's at the end of its rated write endurance is still expected to be able to retain data for one year (consumer) or three months (enterprise). Flash that isn't significantly worn out has much longer data retention.
Did anybody calculate the point where it will not be economical anymore to build hard drives? 1 year? 5 years? 10 years?
You can break down drive costs somewhat into the fixed costs (SSD controller, or hard drive spindle motor and actuators) and the costs that vary with capacity (NAND or platters+heads). The fixed costs tend to be lower for SSDs (or at least SATA SSDs), and the variable costs are higher for SSDs because adding NAND is more expensive than another platter.
Hard drives smaller than several TB are no longer getting new technology (eg. you won't find a 2-platter helium drive), so whenever NAND gets cheaper the threshold capacity below which hard drives don't make sense moves upward.
For the near future, hard drive manufacturers have a clear path to outrun the capacities available from cheap SSDs. NAND flash gets you more bits per mm^2 than a hard drive platter, but platters are far cheaper per mm^2—enough to also be cheaper per bit. I think that relationship will still be true by the time hard drives are using bit-patterned media and 3D NAND is at several hundred layers.
Ultimately, it might make more sense to ask when we will see SSDs having taken over the former hard drive market with hard drives having moved entirely into the market traditionally occupied by tape.
Also, higher RPM hard drives are still a thing and are holding their own to some degree. The higher RPM can help with rebuild times and to reduce the impact of the lower random IOPS. Also, conventional hard drives do not have the write limitations that (especially cheaper) SSDs have, although that has improved over time.
So I think we’ll just see a continuation of the three-tiered system of storage for many years to come, but with hard drives increasingly disappearing into the cloud and away from consumer devices. SSDs for most things, hard drives for bulk server/cloud storage, and tape still for cold, long-term-stable storage.
I think we’ve already seen a plateau in storage cost reduction as SSDs are not cheaper per TB than HDs. I think we’ll put more effort into being efficient with storage management in the future as we can no longer simply rely on doubling storage capacity every couple years.
Depends, 15k RPM are indeed gone. 10k RPM drives are sold and used.
As far as I can tell, WD's 10k RPM drives are discontinued and no longer listed on their site. Seagate lists 10k RPM drives up to 2.4TB and 266MB/s, with a 16GB flash cache. Looking on CDW, it's more expensive than a 3.84TB QLC drive. It uses more power at idle than a QLC SATA SSD under load. I can only imagine a few workloads where the 10k RPM drive would be preferable to the QLC SSD, and I'm not sure the 10k RPM drive would have better TCO than 7200 RPM drives for such uses.
Are there any situations that you think still call for 10k RPM drives to be selected, rather than merely kept around due to inertia?
1) A high-throughput MX/message broker server that is set up and works next 10 years without a need of drive replacement.
2) ZFS SLOG (or any similar scratchpas/transaction log) for 24/7 intensive writes.
In both cases similarly reliable SSD (enterprise ssds) would cost more and require more frequent replacements.
It'll be curious if the hyperscale clouds decide to self-manage more HDD functions, and dumb down the devices ($), or leave that to HDD manufacturers ($$).
I imagine there's some savings to be had stripping memory & controllers out of drives, when you're deploying in large groups anyway. Similar to what was done with networking kit.
Or maybe this already happens? Moreso than "RAID-edition" drives.
For drives using Shingled Magnetic Recording (SMR), the storage protocols have already been extended to present a zoned storage model, so that drives don't have to be responsible for the huge read-modify-write operations necessary to make SMR behave like a traditional block storage device. I suspect these Host-Managed SMR drives are not equipped with the larger caches found on consumer drive-managed SMR hard drives.
I've also been told that we have the technology to get data OFF of hard drives in cases of catastrophic failure. we don't have that capability with SSDs.
So for archiving, I think hard drives should stay a long time.
I think hard drives will be the "tape drives" of the future, relying on capacity more than random i/o speeds.
If you want to archive at rest, bluray/dvd or tape in good climate controlled storage is the only way to do it. Your HDD will spontaneously die simply because it is a moving part well before a comparative SSD reaches the write-limit.
I'm not sure it would be economical for bulk/cloud storage. It might be cheaper to just have geo-diverse redundant storage such that a failed drive can be rebuilt anew from redundant copies of the same data.
I'd be interested to see if optical technologies find a niche, they seem most stable for long-term storage, e.g. M-DISK
see also: https://arstechnica.com/gadgets/2019/11/microsofts-project-s...
How many years it will take to write 120TB on it?
A 120TB flash array can be written in minutes with enough upstream bandwidth, and takes just one rack.
Tape on other hand was always speed limited, is just fine in its nice with that limitation.
The PM1643 16TB SSD is about 7 times as expensive as 16TB disks. (2600 vs 360, euro)
The PM1643 actually uses more than twice as much power in use (r or w) than a spinning disk. Idling they use the same 5W.
The most power-efficient consumer SSDs are an order of magnitude more efficient for sequential transfers than the Samsung PM1643. Eg: https://www.anandtech.com/bench/SSD18/2460
Here are some off-the-cuff possibilities:
"Photos" that are interactive "spaces" composed of combinations of many high-resolution images and (LI/RAY)DAR point clouds.
"Videos" that are the same or far beyond 8k
Individuals using constant-capture systems to archive their day-to-day, providing the ability to "never forget." This could be supplemented with local capture and storage of extensive biometric data for personal-health analysis.
Amazing graphics / texture sets to locally provide standard libraries for very high resolution AR / VR experiences.
Local archives of reference and web content so that it can be browsed privately to ensure privacy.
A general push-back against cloud storage in favor of local privacy. This might be powered by some federated system that allows individuals to give up some local storage in exchange for storage on other's systems.
120TB may still be more than you need for now, but 1-2TB in a gaming PC is now "cozy".
[1] https://www.theverge.com/2021/2/25/22302057/call-of-duty-war...
At the moment, I'm content with 14T drives for personal use, which is about 4x bigger than what I was using just a few years ago.
But if I had 120T, I would just start hoarding every bit of data I could get my hands on. I'm not quite https://www.reddit.com/r/DataHoarder/ level yet, but I could totally see myself heading in that direction.
And at $DAYJOB we need to store data in the order of magnitude of a petabyte. This also seemed huge just a few years ago, but now feels restrictive. 30x 120T drives to shove into a rack in the data centre would totally be considered if it were possible.
I remember in the 90s I heard someone ask how you could ever fill a 1GB hard drive ...
If you do any 3D work (engineering or artistic), the assets for these projects can be huge. If you're doing ML work training sets are massive. If you do more conventional development/content creation, software and assets needed keep getting larger with a never ending list of things to pull down from github etc. If you're just working with data/databases, the amount your working with never seems to get smaller and if it's time-based there's always more history to deal with as the years tick by.
Even if one does everything in 'the cloud', all you're doing is shifting the location of the storage needed rather than the quantity. So it's pretty much the same old story of 'more' when it comes to how much storage is needed.
Data density is great and all, but this line got me wondering if we wouldn't be supporting more than just the 2.5" and 3.5" form factors. I'm suspect there would be a market for storage devices with twice the volume and 50% or more storage. Or does the 3.5" form factor have something to do with the limitations of spinning rust tech?
There are issues. As the platters get bigger they get wobbly, more probe to vibrations. The difference in platter speed between the inner and outer 'rings' of data also increases. There is no hard reason for 3.5 over 4.5 or 4.8 inches, but the advantages of going bigger do not outweigh the practicalities of trying to push a new form factor for a commodity product.
>> spinning rust
This is a common refrain, but seagate and the others are proving that HDD tech is not 'rusting'. With lower per-TB costs, and phenomenal reliability numbers (<1% chance of failure per year) HDDs are going to be around for a while. Outside of the datacenter, most users are more limited by more by network speeds than storage speeds.
> This is a common refrain, but seagate and the others are proving that HDD tech is not 'rusting'.
It is a common refrain because disk platters were originally coated with iron oxide, and iron oxide is rust. It isn't a value judgement.
I think the original comparison to 'spinning rust' was tongue-in-cheek and not meant to be taken so-very-seriously, but it's kind of taken on a life of its own and, personally, I feel that it's gone a bit far and undeservedly undervalues hard drives.
"Logically, if you utilize a 'memory' disk you are ridding yourself of the physical limitations of disk drive technology — basically trading in the spinning rust. In a disk drive, you are battling physics to squeeze more performance out of your array; there are certain physical properties that just can’t be altered on a whim." - 2004
https://web.archive.org/web/20041018152106/http://www.dba-or...
As to the question of its present use as a method to hurt the feelings of hard drives and hard drive cognoscenti... I dunno.
I think you're reading too much into it.
- more power usage
- more heat
- more noise
Datacenters in particular are not interested in increasing heat or power usage. 5.25" hard drives used to be common, but were phased out sometime around Y2K.
There used to be larger form factors for HDs, for instance full-height 5.25 inch (think two DVD drives stacked on top of each other). This Wikipedia page has pictures: https://en.wikipedia.org/wiki/List_of_disk_drive_form_factor...
This was apparently a product that took the strategy you described in the 90s: https://en.wikipedia.org/wiki/Quantum_Bigfoot
There was a variant of the WD Raptor that took the opposite approach: it put a 2.5" 10k RPM drive in a 3.5" enclosure, but this drive played the role SSDs do today; they weren't for mass storage.
More the opposite, really - same capacity in lower volume. The people who buy most disks are a lot less concerned with capacity per disk than capacity per rack. Sometimes capacity per kW or capacity per BTU, but mostly capacity per m^3 in my direct experience. I also suspect that a lot of consumers would prefer the same capacity in a smaller package, since they mostly boot off flash and their disk is basically a cache for much larger amounts of data elsewhere.
With your extreme example, I doubt it. Already you can get two drives of 75% of the capacity for just a bit more money, and that would perform far better than one drive of twice the volume.
The current form factors are baked in, how would we know what the free variables are?
You can see them in this magnetic image - the whole 25% to 75% of the image is the servo signal area. The tracks are seen on the bottom 20%, top 80%:
https://commons.wikimedia.org/wiki/File:Aufnahme_einzelner_M...
Bottom to top:
- sync (the regular lines)
- positional track number, gray encoding (irregular lines)
- track fine tune (the block trio)
- fine track ID
After fine track ID is read correctly, the position is deemed correct, the RX head is turned off, and the TX head is turned on. This is the rx-tx gap. Then track data starts.2^64 bytes would be 16 exibytes, or about 18,446,744 terabytes.
I'm not even sure YouTube has that much data.
So in 10 years we will have 100x120TB = 12PB per disk. Put 8 of these disks in a 2-U unit and we'll have a 100PB box. Stack 10 of these units together in a rack and we'll get an 1EB storage rack.
Remember some file systems use couple bits off the 64-bit addressing space for other purposes. Shaving 4 bits off and we get 1.15EB. That's within the 10-year storage capacity estimate.
While present day Youtude might not have that much data, I'm sure when the 16K, 32K, and 64K videos come along the space will be filled up. Let's not forget the voluminous 3-D Lidar scan data generated by every iPhone when people start recording 3-D VR-Holodeck from their phones. Besides wouldn't it be cool to host an entire copy of the Youtude in your closet? ;)
I can't imagine how long it will be until we see 16K, considering 4K caught on in around 2014/15. 8K video exists, but is still exceptionally rare, at least for consumers.
> Besides wouldn't it be cool to host an entire copy of the Youtude in your closet? ;)
As long as you're not still hamstrung by Xfinity's 1 TB data cap...
How about some napkin math on the cost of that data overage? hah
This was a few years ago, so it might be possible that Youtube has more data than what can be addressed with a 64-bit value.
Of course, a failure once in a while is expected. With Seagate I had a string of failures over a few years. I tend to think that randomness was not alone in explaining it (or I was really unlucky).
It seems that shingled recording (SMR) could still be the future, although the original article does not seem to mention it?
At what point do we stop because reliability has gotten unacceptably low?
I have 3 1TB WD Blacks that reach 300 mb/s, when I was shopping for my drives, similar SSDs were either crazy expensive (10 times are much for example) or they were tiny, or the ones with "use-able" size and reasonable price, were SLOWER than the WD Blacks in sequential writes and reads (they were still faster in random access, for obvious physics reasons).
To be honest I am perfectly happy with my WD Blacks and never felt the need to switch to SSD, usually my speed bottlenecks are somewhere else (often network speed! it is surprinsingly hard to find good network hardware, all I find is random cables, adapters and routers, that shops don't even know what their speed limits are, asking if a cable supports gigabit ethernet just get me confused looks)
EDIT: This post score is fluctuating, suggesting to me people are sometimes downvoting it. I wonder why... I mean, I posted some anedacta, but why the downvote? There is not even anything to disagree with on my post, unless I wrote something very wrong, but nobody replied to say what is wrong...
The WD SN850 NVMe flash drive does 7000mb/s reads. (Close to maxing out a PCIe 4.0 4x slot) That's 23 times faster than the hard drive speed you quoted. SATA SSDs have been obsolete for a decade now.
Edit: Just checked some prices on Newegg. Looks like the XPG GAMMIX S50 goes for $139, and does 3900 MB/s reads, while SAMSUNG 860 Pro does 560 MB/s reads... and costs $179.
It is unfortunate that the invisible hand does not seem to care about drive speeds. Absent price signals, online retailers do not do a terribly good job signposting the enormous gulf in performance between SATA and PCIe. Newegg lists 999+ 2.5" SATA SSDs, and there is no reason at all to buy any of them when building a new computer.
Also some workloads perform poorly on SSDs (due to their limited read/write cycles). But mostly dollars per terabyte.