So exactly 14 years passed, does someone have 2^64 bytes for a single ZFS filesystem (or anything close to that)? I don't really feel like storage capacity (or 1/price) doubles every year.
So exactly 14 years passed, does someone have 2^64 bytes for a single ZFS filesystem (or anything close to that)? I don't really feel like storage capacity (or 1/price) doubles every year.
So with 5 doublings, if we saw datasets on the order of 2^50 back then, we should see datasets on the order of 2^55 today (30 PB). And sure enough, here is a computer with a 30 PB global filesystem: https://www.fujitsu.com/downloads/TC/sc11/k-computer-system-...
So another 9 doublings to go to 2^64 bytes. With one every 34 months, this should happen in 25 years (2043). But again, if you are a filesystem developer you should be on the conservative side and assume it might double every ~12 months.
If ZFS lives as long as the FATs have (three decades) even at a doubling every 34 months, that means six remaining doublings before 2035. Not quite 2^64, but that's on the liberal side. Half that (doubling every 17 months) is twelve doublings, which is easily over 2^64.
That has to be one of the worst typos I have ever seen.
Of course, the real reason for the mistake is that phone spelling correction will convert endian to Indian.
;)
That’s 7 PB to 19.2 PB. If they had replaced all 3,500 drives it would be 4X.
I think that's the point from some of the other comments (below). Instead of having every larger and larger single filesystems, once you get to a certain point (~1-2PB is my guess), it simply isn't worth it to organize your data under a single filesystem.
That's not to say ZFS was wrong in choosing 128 bits... at the time ZFS was designed, the level of horizontal partitioning that we are now accustomed to just wasn't a thing. Compared to HDD times, networks were far too slow to support such a thing. Now, however, we talk about HDDs being too slow compared to our networks. At the time, 128 bits was a solid choice, even if it was overkill.
Plus, Sun was definitely a scale up kind of company... if you want to get bigger, by a bigger ($$$$) box. So having a large capacity FS was good for their bottom line. Now, we think in terms of scaling out with multiple (cheaper) smaller boxes... which is probably the only way we've managed with current levels of data.
No, that is not true. A vault acts as a single unit, effectively making it a single filesystem as much as a ZFS pool is a single filesystem.
Internally, a vault operates on a set of "tomes", which are groups of 1 harddrive from each storage pod (20). These "tomes" handle the redundancy aspect as well. However, the "tome" is not the filesystem.
The concepts map cleanly to the ZFS world: A vault is equivalent to a ZFS pool, and a tome is equivalent to a ZFS vdev in redundant configuration. Neither vdevs nor tomes are filesystems on their own—they only become a filesystem when combined into pools and vaults, respectively.
Now, your argument does hold if interpreted differently: One does not use a normal filesystem for those absurd sizes. However, it is a single filesystem.
I guess I was assuming it worked like Lustre or BeeGFS where there were smaller FSs (ext4, XFS) at play that handled storing file (chunks) and then a higher level interface that managed it all.
Either way -- the original comment was referring to Backblaze swapping 7PBs for 19PBs. Which we agree is not a normal size for a "normal" file system! ZFS is a strange beast... it was designed to support a world that just never showed up. It is a "normal" FS that works from laptops to servers, but was designed to support a big iron world that never really appeared (much). Instead, here we are comparing it not to XFS, but to cloud FS's that were designed for a completely different "cloud" world.
I would assume that the tome, beneath network protocols, interacts directly with the individual block devices. They have implemented things like Reed-Solomon error correction themselves to handle redundancy across the tome (which to recap is their cross-server vdev equivalent), so they have indeed implemented tasks of a filesystem and RAID manager on top of their lower-level primitives.
I also believe that what their low-level primitives are is actually a rather irrelevant implementation detail. A filesystem presents a way to store and organize files, and is a filesystem regardless of whether it is implemented on a block device or an egg-engraver. It could be using FAT16 as object store, sharding across files and handling metadata elsewhere. I would still consider such setup to be a true filesystem, as it is not simply a 1:1 network-to-disk protocol.
And yes, ZFS won't see the large systems it is designed to handle anytime soon. What Sun did not see coming was the death of large, high-capacity servers, and the rise of distributed solutions. But then again, no one saw that coming in 2001 when ZFS development started.
However, considering that storage sizes will continue to increase, who knows what will happen in the future... 128 bit filesystems might come in handy at some point.
(Note that my information on backblaze's systems is based on public information—I am not an insider of any sort.)
Is that a fundamental thing or is it just because there aren't many filesystems that can handle it? It seems to me like a single filesystem would be conceptually easier, but I'm far from knowledgeable in this area.
But really, I do think it's physics at play... or rather logistics.
Let's say you're trying to put together a 10PB storage system in a datacenter. If you're using 60 disk JBODs with 10TB drives, thats 600TB of raw space per 4U (or a max of 6PB per rack, or 12PB in 2 racks). Now realistically due to power, network gear, etc... instead of maxing out 2 racks, you'd probably split that into at least 4 (or more) racks. If you're Backblaze (see above), you split those 20 pods across 20 separate racks (and add more volumes, but that's another story).
But once you're at that level, for redundancy, you'd not only managing which disk data is written/replicated to, but also the JBOD/pod, and the rack (and the data center). So, at that point, you're not going to working with a "normal"/traditional file system. You could with ZFS define redundancies that would work like that, but realistically, because of the extra overhead, you're better off changing how you organize your data.
(Plus: ZFS isn't very good at expanding a filesystem with new storage. You can do it, but your pool ends up unbalanced and performance suffers. When you decide to have a 1+PB system, you normally want to have the ability to add more storage space, which means a different type of FS)
So basically we moved away from singular large unified file systems and built swarms of little file systems.
I can't certainly blame their choice though, and 16-million terabyte hard drive arrays (the limit of 64-bit addressing) on a large mainframe are really just a few quantum leaps away. Heck, I still vividly recall buying a MASSIVE 40 GB HDD in 2000 - it seemed excessive beyond wildest dreams at the time (and was promptly filled to brim at a LAN party.)
Since, advances in mass storage tech, most importantly the successful commercialization of perpendicular recording in 2005, mean that nowadays 12 TB drives are commercially viable - roughly 1000x the size of a 40 GB drive.
In the next 20 years, petabyte-sized drives might well be available (though almost certainly not based on any form of magnetic recording). Once you hook hundred 50-petabyte drives to a single fileserver, you start closing in on limitations of 64-bit addressing.
When you're looking at the same price for two systems, and one is 100x-1000x faster at the cost of being 'merely' hundreds of TB per gallon instead of a few PB per gallon, nobody is going to bother engineering the latter just to save a few bucks on shipping.
Everybody uses it to mean "big". Oh, language.
I assumed it implied something far 'larger' such as a leap through spacetime...
The interesting thing there was that the electrons existed only at discrete energy levels, and the transition between them appeared instantaneous.
It typically implies a revolutionary change over evolutionary improvement.
The point is, that the increase in frequency is so large, that using MHz is no longer sensible.
There has been a "quantum leap" in frequency.
No cloud vendor has a billion harddisks, that's for sure.
(The filesystem might only support a maximum file size of 2^64 bytes, but that's nothing to do with the capacity of the storage array)
ZFS aside, the 512 byte sector isn't even a hard limit at the storage medium level. Disks with 528 byte sectors — and other odd sizes — have been available in the enterprise world for a long time now: https://en.wikipedia.org/wiki/Hard_disk_drive#Market_segment...
HDDs and SSDs with 4 kB sectors dominate the consumer market now: https://en.wikipedia.org/wiki/Advanced_Format
https://cloud.google.com/files/storage_architecture_and_chal...
No, they're not a single filesystem, as the papers make clear here and there.
Rates were quicker back around 2000-2005, but that pace has fallen dramatically.
https://en.wikipedia.org/wiki/Mark_Kryder#Kryder's_law_proje...
May I remind you that flash, DRAM and SRAM already are 3D structures.
I calculated it at some point and it was several years of the entire earths manufacturing capacity for disks in a single filesystem.
[added the missed "several years of"]
Or, exactly 2 Zettabytes.
—Wikipedia
ZFS being an acronym for Zettabyte File System was neither retrofitted nor all that brief in use it seems.
Former employees of Sun Microsystems, I know some of you browse HN. Care to shed some light on this?
There is also EOS[2] which is based on XRootD and is planned to replace AFS[3] as a general purpose networked filesystem.
[1] http://xrootd.org/ [2] http://eos.web.cern.ch/ [3] http://information-technology.web.cern.ch/services/afs-servi...
Moore's law only ever really held for transistor density on integrated circuits, not magnetic storage. Magnetic storage capacity always grew more slowly.
Flash storage capacity today has more to do with die stacking than transistor density.
In any case, Moore's Law is definitely dead now. Like, for real this time.
https://www.backblaze.com/blog/hard-drive-cost-per-gigabyte/
They also link to this great (logarithmic) chart:
http://www.mkomo.com/cost-per-gigabyte-update
TL;DR: price hasn't been dropping by half annually, but prices have come down from ~$0.70/GB to < $0.03/GB now. So roughly a factor of 20 over 13-14 years, somewhere in the neighborhood of 25% improvements compounded annually over that time.
Its unlikely. There are few opportunities to compile a data set of that size. One might be the NSA collections center in Utah, one might be Google or Microsoft's web cache, and one might be something like the Internet Archive's cache.
However the length choice gets 'weird' if you want something larger than 64 bits and less than 128. Your next intermediate choice is 96 bits. That is 64 + 32 (or one additional 32 bit long word. Growing by just 8 means you now have a 9 byte pointer, by 16 gives you a 10 byte pointer. Some architectures are penalized when doing off word alignment, and so on those architectures the structure is padded out to the next word size anyway.
I've always found Jef's reasoning in that post amusing but I continue to thing 64 bits would have been fine (just like I think it would have been fine for the Internet v6 work :-). Fortunately time has given us a crap ton of memory for our boxes so the storage penalty isn't too great.
The on-disk format could even remain the same; it wouldn't really matter except for lots of small files and those are already so horrid that ZFS might have some other way of handling them. (Like treating /all/ small file names and data as part of a larger directory file or something.)
Even if we ignore that factor of 800 or so (400PB spread across 500 nodes), they still would need another factor of 128,000 before they would need ZFS's 65th bit.
Keep in mind it's 2^128 blocks, not 2^128 bytes.
Backblaze would need to consume the entire planets supply of disks for a very long time AND put all those disks connected to a single linux, freebsd, or solaris box. Only then would they need the 65th bit that ZFS has for addressing blocks.
https://www.anandtech.com/show/10098/market-views-2015-hard-...
I don't know about that exact disc/format, but if we and as we are finally able to stably write at such capacities... (whatever technologies end up being used to do so).
We're going to run into another "640KB should be enough for anyone", unless we are forward thinking with regard to potential storage format capacities.
(And if you think no one will ever be able to consume that much data and detail, just think of all the modeling that will be done and the data behind that. Think of the data being put out by CERN and not just how it will expand but also how people will want to explore lower significance hints, corner cases, and who knows what. Etc.)
https://www.emc.com/collateral/hardware/white-papers/h10719-...