Bitrot and atomic COWs: Inside “next-gen” filesystems
arstechnica.com
arstechnica.com
People sneer at me for being a stickler for it, but between 50GB of memory and 26TB of ZFS-protected storage, I see ECC corrections about as often as I see disk checksum errors - maybe half a dozen of each in the past year or two. Frankly I think it's idiotic it's not more common and better supported.
What doesn't seem easy is to get ECC in laptops. I don't think even traditionally business-focused lines have it (e.g., Thinkpads).
It would be nice if the intel ultrabook standard started including ECC as well. One can dream...
Prices for these start about 3x higher than other boards, their choice is extremely limited, and availability even more so :/
Granted I spent nearly 10 years immersed in storage systems (5 at NetApp, 4 at Google) but still there is some easily checked stuff missing from this article and so its point (which is hardening in the filesystems is good) is lost.
Nope. On a non-degraded array with parity (R5, 6, ...) it will only read the block it needs to from the drive it is on. For an array that mirrors without parity (R1, ...) it will read the block from one of the drives. There is no need to bother another drive with the read operation unless it needs to check parity (because it has been told to on each ready as some controllers can be so told), you actually reduce the performance benefit of the striping if you do as you may be moving the heads of the other drive(s) "unnecessarily" and potentially away from another block they were about to be asked to read.
With an un-degraded array unless explicitly told to check parity on read the controller will not touch the parity blocks until a write happens at which point it will read the other relevant data blocks in order to regenerate the parity block. This is why the RAID 5 write penalty exists but there is no read penalty (in fact there is a read bonus due to striping over multiple devices).
Parity blocks will, unless checking parity on every read, only ever get read if the array is in a degraded state, in which case you can only derive some data blocks by reading the other blocks in that stripe (data and parity) and working out from them what the missing block should be.
For R1 you can in theory use an elevtor algorithm to chose which or the two mirrors serves a given read depending on how close the block may be to the most recently accessed blocks on each device, reducing average head movement distance. For bulk reads or an I/O pattern including much writing this will make little or no difference, but for significant random access read patterns you might see a measurable (perhaps even noticable to a human user) latency improvement.
Last time I tested this myself I found that the Linux software RAID arrangement did not perform any such optimisation for RAID1. Results like https://raid.wiki.kernel.org/index.php/Performance agree with my findings, though they do suggest that with certain layout options the newer RAID10 single-level driver provides such optimisation for reads (at the expense of slower write performance). Linux's md-raid "raid10" module is not your traditional RAID10 (a nested arrangement of RAID 1 and RAID0, a stripset of mirrors needing at least four drives) - it is a single level driver that offers greater layout flexibility with regard to the copies of data on each device and even supports operation on three devices (this operates similarly to what IBM RAID controllers call RAID1E and some others call RAID0+1) or even just two devices (in which case it behaves as RAID1 but with the extra layout options). See http://neil.brown.name/blog/20040827225440 amongst other references. IIRC you should avoid the "far" layour for your boot devices as boot loaders don't tend to understand it.
I'm fairly sure some current systems (MD on Linux for example) have this sort of optimization. And that is because the designers assumed that the disk would be fine or just fail, and not be in some state in-between.
This assumption used to be more true that it is now. The MTBF for a read error has not been increasing as fast as hard drive capacity has been. In the old days, it was highly unlikely to get a silent read error. However, these days it is more common due to the massive increase in capacity.
Interesting, if you make that assumption you will get burned. I have personally experienced disks that through firmware errors returned what was essentially a freed memory block from their cache as the sector data, returned success status on a write that never actually happened, and flipped bits in the data they actually returned (without error). The whole point of RAID for data reliability is to catch these things. (RAID for performance is difference and people tolerate errors in exchange for faster performance).
We used to start new hire training at NetApp with the question, "How many people here think a disk drive is a storage system?" and then proceed to demolish that idea with real world data that showed just how crappy disks actually were. You don't see these things when you look at one drive, but when you have a few thousand to a couple of million out there spinning and reporting on their situation you see these things happen every day.
The issue of course is that margins in drives are razor thin and they are always trying to find ways to squeeze another penny out here and there. Generally the expresses itself as dips in drive reliability on a manufacturing cohort basis by a few basis points. You can make a more reliable drive, you just have a hard time justifying it to someone who is essentially going to get only one in a million 'incorrect' operations.
This is the same reason ECC RAM is so hard to find on 'regular' PCs, the chance of you being screwed is low enough that you'll assume it was something else or not be willing to pay a premium for the extra chip you need on your DIMMs for ECC.
>> "And that is because the designers assumed that the disk would be fine or just fail, and not be in some state in-between."
> Interesting, if you make that assumption you will get burned.
Probably why the designers of RAID have moved on to other storage systems. Yet, RAID is currently the only viable option for consumers to try to get "reliable" storage. Which is why the article was written.
> The whole point of RAID for data reliability is to catch these things. (RAID for performance is difference and people tolerate errors in exchange for faster performance).
No, RAID is for withstanding disk failures, where the disk fails in a predicted manner. This is a common misunderstanding. Nothing in the RAID specifications say that they should detect silent bit flipping. If you find a system that does this, it goes beyond the RAID specification (which is good). But you seem to be under the impression that most systems do that, but that is not true. Most RAID systems that consumers can get their hands on won't detect a single bit flip. The author demonstrated it on software raid. I am fairly certain that the same thing will happen if you try it on a hardware RAID controller from any of LSI, 3ware, Areca, Promise etc.
For instance Google has multiple layers of checksumming to guard themselves from bitrot, at filesystem and application level. If RAID did protect against it, why don't they trust it?
If you use a 15/16ths scheme you only "lose" a bit more than 7% of your drive to check data, and you gain the ability to avoid really nasty silent corruption. I'm not sure why even a "consumer" RAID solution wouldn't do that (even the RAID-1 folks (mirroring) can use this for a bit more protection, although the mean-bit-error spec still bites them in trying to do a re-silver)
As for Google, If you'd like to understand the choice they made (and may un-make) you have to look at the cost of adding an available disk. The 'magic' thing they figured out was they were adding lots and lots of machines, and most of those machines had a small kernel, a few apps, and some memory, and an unused IDE (later SATA) port or ports. So adding another disk to the pizza box was "free" (and if you look at the server in the Computer History Museum you will see they just wrapped a piece of velcro around the disk and stuck it down next to the motherboard on the 'pizza pan'.) In that model an R3 system which takes no computation (its just copying, no ECC computation) with data "chunks" (in the GFS sense) which had built in check bits) and voila "free" storage. And that really is very cost effective if you have stuff for the CPUs to do, it breaks down when the amount of storage you need exceeds what you can acquire by either adding a drive to a machine with a spare port, or replacing all the drives with their denser next generation model. Steve Kleiman the former CTO of NetApp used to model storage with what he called a 'slot tax' which was the marginal cost of adding a drive to the network (so fraction of a power supply, chassis, cabling, I/O card, and carrier (if there was one)). Installing in pre-existing machines at Google the slot tax was as close to zero as you can reasonably make it.
That said, as Googles storage requirements exceeded their CPU requirements it became clear that the cost was going to be an issue. I left right about that time, but since that time there has been some interesting work at Amazon and Facebook with so called "cold" storage, which are ways to have the drive live in a datacenter but powered off most of the time.
I can't agree with this statement, "No, RAID is for withstanding disk failures, where the disk fails in a predicted manner." mostly because disks have never failed in a "predicted manner", that is what Garth Gibson invented RAID in the first place, he noted all the money DEC and IBM were spending on trying to make an inherently unreliable system reliable, and observed if you were willing to give up some of the disk capacity (through redundancy, check bits) you could take inexpensive, unreliable, drives, and turn them into something that was as reliable as the expensive drives from these manufacturers. It is a very powerful concept, and changed how storage was delivered. The paper is a great read, even today.
Anyone that is actually available for "normal" people (which is relevant for the discussion of the article), i.e. not enterprise SAN?
If you have an "iffy" sector that is on th eway out you might get a decent read after a few attempts. You can then make sure that the blocks on the other drives are OK and (assuming all the failures are in the same device) drop the problem device when done.
In reality most drives these days do this for you: a certain number of sectors are reserved for reallocation in the case of small surface failures so the drive itself will retry a few times to get a reliable read (each sector on a disk has checksums that allow error detection) then the controller will remap that sector. This happening once or twice is considered normal wear and tear, this heppening beyond a certain threshold or a certain rate is a sign of imminent failure and is measureed by SMART indicators - so running the scan will not let md raid do much directly but will give the drive chance to remap data that is in danger (and if you have mdadm setup right, email you a warning if you shoudl consider dropping the drive immediately and replace it).
for raid in /sys/block/md*/md/sync_action; do
echo "check" >> ${raid}
done
which will check and fix errors (or fail). i run this weekly as a cron job.of course, it doesn't help with any data that are corrupted and read before it runs.
Um, isn't this the EXACT point of the article you are so happily criticizing? That filesystems don't do this, and they should?
Now in reality in big enterprise SAN arrays, the vendors usually add checksum blocks onto the underlying disk (aka larger sectors in the Clariion, or checksums into the filesystem on NetApp) OR use special certified firmware on the drives that make sure the drives never lie themselves! So largely if you are on a big enterprise SAN, you probably don't seriously need to worry about bit rot.
But outside of that space, most PCI card RAID controllers, and I suspect at least a few "enterprise lite" arrays, probably don't do checksum calculations. So BTRFS and ZFS provide notable value. (and definitely on raw drives)
Intel could easily offer ECC support in its consumer line of desktop (none) and laptop processors (only offers it on 3 of them: http://ark.intel.com/search/advanced/?s=t&FamilyText=4th%20G...), but doesn't.
To me, ECC support should be like SSL/TLS for modern applications; don't do it without it!
It's the only FS I've ever had silent corruption on, and it used to happen (10.4? 10.5? 10.6?) all the time.
I'd kill to have Apple just buy a license to NTFS.
HFS beats that quite handily (and gets obliterated by ExtFS2 performance.)
Would I want to store a lot of important write-once data on HFS? Reluctantly, if at all. High-throughput database? Yuck, no.
Would I trust it with a source-tree that is stored in a DVCS and immediately restorable, while also giving me much faster compiles? Yep.
Would I choose either FS for a server? Nope. That was the point - there is no single "best" file system, or the world would've long ago settled on it. Look at what you need it for, and make your decision accordingly.
I am legitimately asking because I have a good 2TB of family photos on hard drives and spooky stories about random bits flipping freak me out.
But you don't need anything fancy; 2TB is small enough though that you can just buy another HDD and back it up manually.
Scrubbing (periodically reading the data and comparing it to checksums) is one way to get around this. It's very effective against small numbers of sector errors in backups (you should have more than one), and in detecting marginal data on drives. It's less effective against some other kinds of corruption. Another option is to store multiple copies (replication), or additional ECC information that can be used to recover the data if one copy is lost.
How much effort you put into this really depends on how much you need or want the data.
[1] http://www.cs.wisc.edu/adsl/Publications/latent-sigmetrics07... [2] http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.64....
Seagates at least report their error correction rates via SMART, and make very obvious curves that show them scanning the heads across the disk surface gradually when idle.
InternetFS(Encrypted, distributed, always in sync, cheap snapshots, p2p)
Want to share Photos or movies with your friends and family? No problem just right click on the file, select the friend from a list, they see the file and they can view it on their computer.
This next gen file system will be incredibly easy to use, cross platform. As simple to use as Facebook, Skype, Email.
> You need to run special gateway servers
You don't _need_ to, each peer can run its own gateway on localhost. Gateway servers are useful if you're on a computer where you can't install anything.
> I don't think it uses transparent encryption
I'm not sure what you mean by that, but if it means you need to manually encrypt things before storing it in tahoe-lafs, you're wrong. The whole point of Tahoe-LAFS is to push your files to it, and it will automatically encrypt, erasure-encode and distribute it to all nodes in the swarm you're in. All you need to care about are URIs returned by Tahoe-LAFS, and which contain all you need to use files (no need for extra password or keys, in fact they're contained in the URIs)
> there isn't a global namespace
Tahoe-LAFS swarms are defined by their introducers, the central piece that ties all storage nodes and potential clients together, kind of like a bittorrent tracker. There are no namespaces because swarms are not globally defined (although a node can take part in multiple swarms, this is completely invisible from Tahoe-LAFS users. Exactly like bittorrent). My proposition was to build a world-level swarm, where anybody could register and participate in the swarm. If you wonder about the reliability of a central introducer, there is work in the pipes for creating decentralized introducers.
1: http://ori.scs.stanford.edu/ 2: http://dl.acm.org/ft_gateway.cfm?id=2522721&ftid=1403940&dwn...
Needs encryption, access controls and some kind of public namespace for publishing.
Space Monkey looks really interesting but I agree that this isn't something that is going to perform well running on residential internet connections. It also doesn't look like there is anything like non-explicit sharing (i.e. public files) which I think is an essential part of a distributed data store because it would be a boon for small-time publishing (no more web servers!).
I would love the net to be more decentralized but everyone living behind NATs means we need central servers to connect the users between them.
Do you know if your NAS4Free solution a viable solution for a household full of Macs that need a shared Time Machine destination? I'm tempted: it's cheaper, and NAS4Free will keep evolving.
SMB, SSH/SFTP (even set up key-based auth), regular FTP all work great. Transmission works great and has a decent web interface. UPS support seems to be okay, haven't really tested it.
DLNA/UPnP (fuppes) crashes all the time and I only tried it with the very buggy VLC for iOS, so I'm not sure how stable it actually is.
Also, from what I remember, due to the fact ZFS keeps previous versions of files around, the effective data capacity you get is something like 1/2 or even 1/3 of the capacity of the HDD, no?
You can create snapshots of FS to keep old versions, but how many, and therefore how much space is used, is up to you.
The whole point of ZFS is that if a bit gets flipped, it's corrected the next time the file is accessed. There's a scrub command that manually hits every file to check for integrity and you can schedule it through cron.
And I assume you can control the size of the snapshots, I haven't looked into it but it would be pretty silly if you couldn't. I have four 3TB drives in RAID-Z and I have 9TB to play with, same as RAID-5.
For snapshots - they only get made, if you make them (or setup an automated script that makes them). You can list the snapshots and see how much space each uses, as well.
Also, I'm not sure where you got the "ZFS keeps previous versions of files around" idea from, it only does that if you make a snapshot, and only modifications of the file cause the file to be copied; this is waht copy on write is. It means you can have the whole history of your file system with almost no space overhead except for changes. Apple's time Machine does the same thing, but in a different way: each directory that hasn't been changed since the last backup is hard linked to the previous backup's version (which might itself be a hardlink). this makes is quite space efficient (and a pretty easy to understand hack to get versioning on a file system that doesn't natively support snapshots)
[1]: http://5by5.tv/hypercritical/56 [2]: http://5by5.tv/hypercritical/57
[0] http://blogs.msdn.com/b/b8/archive/2012/01/05/virtualizing-s...
> By default, when the /i switch is not specified, the behavior that the system chooses depends on whether the volume resides on a mirrored space. On a mirrored space, integrity is enabled because we expect the benefits to significantly outweigh the costs.
Plus:
> When this option, known as “integrity streams,” is enabled, ReFS always writes the file changes to a location different from the original one. This allocate-on-write technique ensures that pre-existing data is not lost due to the new write
OS/2 had installable filesystems. Does Windows have something like it?
Windows has ReFS - http://en.wikipedia.org/wiki/ReFS
ReFS is beta-quality and lack several features ZFS has since its first production-grade release. It's not really an apples to apples comparison.
I use it for my backups, both automated and manual. For my photos, I write a set of 10-15 DVD-sized tar archives containing files that haven't been backed up yet, using a tool I wrote. I then run parchive to generate two or three DVD-sized ECC files, so if any one DVD gets trashed I can recover it. Then it's just a matter of burning the DVDs and stacking them somewhere off-site.
I'd like a file system that duplicates sensitive data on the same drive. The file system data should be duplicated too, and marked in some way as to be able to reconstruct a disk after failure.
I don't need to save on space, I only need safety. And keep in mind that it takes many hours to dump a single time the contents of a disk - so, it's practically inaccessible as a whole over short periods of time.
For now, I get by with DropBox and TimeMachine but it's far from perfect. My photo collection alone is 1TB, so, no luck in backing it up in the cloud.
But still use mirrors or raidz to protect against failure of a drive.
Amazon glacier would store that for $10/month, though the retrieval costs if you needed to restore the whole thing are a bit more complicated.
Nowadays, we use something more like a hash function, where we take the block of 4kB or so and produce a 64-bit hash of it. In order for this check to fail, the failure has a 1 in 2^64 chance of guessing a correct checksum for the corrupted data, which is very unlikely to happen.
Also, we have the possibility of using Reed-Solomon error correcting codes in some applications, which can not only detect errors but also correct them.
The problem would be if the checksum AND the data _both_ became corrupted in a way that the checksum were still valid. After all, if the checksum were to change, you'd get a failure, and then you could just verify that the data was not actually changed and/or lost. Not to mention that if you stored the checksum in two places, it'd be pretty easy to see that it was the checksum that itself changed.
In ZFS, a storage pool is a merkle tree[0], the checksums themselves are checksummed (as part of their ancestor blocks). The one risk is that the uberblock itself (the root of the tree holding a checksum for the whole thing) becomes corrupted, which is why IIRC zfs stores the last 128 revisions of the uberblock in 4 different physical locations[1].
[0] http://en.wikipedia.org/wiki/Merkle_tree
[1] sadly does not mean you can rollback to any of them, when a new uberblock is created (which is common) the previous's set of metadata (meta-object set MOS) becomes reclaimable, and once the MOS has been reclaimed/reused the corresponding uberblock is useless
Generally when you design such a system you can often say what would have to be true for you to "miss" that something was corrupted. In our parity example, an even number of bits would have to change state. In the CRC example bit changes would need to be correlated across a longer string of bits. Once you have ways that you know you would not be able to detect errors, then you start breaking the system apart to change up detection and correction. So for example at NetApp a block (which was 4K bytes at the time I was there) on disk was 8 sectors, then there was an additional sector that included information about both a CRC calculation for the available bytes, as well as information about which block it was and what 'generation' it was (monotonically increasing number indicating file generation). The host bus adapter (HBA) would do its own CRC check on the data that came from the drive, passed through it, and landed in memory. That would detect most bit flips that occurred on the channel (SATA or Fibre Channel port) as data went through it. ECC on memory would detect if memory written had its bits flipped. Software would recompute the block parameters and compare them to the data in the check sector.
So if data on the disk was bad, that check sector would not work, if the data had been written correctly initially and gone bad, the RAID parity check would catch it, if the data was corrupted crossing the disk/memory channel the HBA would catch it, if the data got to memory but memory corrupted it, the ECC would catch it, if the memory some how didn't see the corruption the check against the block check sector would catch it. All layers of interlocking checks and re-checks in order to decrease the likelyhood that something corrupted yoru data without knowing it.
(Also I agree--yes MIME types in filesystem please)
Take a snapshot before you make changes in a script, then run the script, if the script fails, just rollback!
On Linux for example you can set the user.mime_type xattr and it's possible to have Apache use that for it's mime-type.
Versioning makes sense for performance but adding metadata sounds like a huge can of worms with no obvious benefits for a low level implementation IMHO.
Besides your metadata wouldn't carry well to other filesystems (most USB drives still use FAT...) so it would be a bit of a headache to get right. Look at the mess that are file permissions and ACL already.
Because you haven't added it. Extended attributes can store arbitrary data. Propose something, experiment, see how it works, and maybe you'll come to something we can standardize on.
Because it's a bad idea. Apple tried it and it didn't work.
See: http://en.wikipedia.org/wiki/Creator_code and http://en.wikipedia.org/wiki/Type_code
COW: Copy-On-Write
Ref:
* http://en.wikipedia.org/wiki/CD-R#Lifespan
* Library of Congress "CD-R and DVD-R RW Longevity Research" http://www.loc.gov/preservation/scientists/projects/cd-r_dvd...
The actual study: http://www.loc.gov/preservation/resources/rt/NIST_LC_Optical...
* http://www.thexlab.com/faqs/opticalmedialongevity.html (references the NIST study above)
2) I would never trust my data to a medium that can get destroyed if touched with your finger the wrong way and you're required to handle it manually to read the data on it
With what? A time machine?