Is there any docs that explain why COW make Btrfs so slow compared to QCOW2?
So qcow has a read only base image that gets updated when we change things. The image format just had the changes from the original image. So you update a package, it adds some metadata to point at the new stuff and adds the data in and you are done.
So with btrfs you have this image on top of btrfs, so you update a file and its metadata inside the image. Say you start with a pristine image that's in nice big extents. You update a package which changes small chunks all over the file. Let's say you update 12 4K extents. So now instead of one extent you now have 36 extents. This affects everything, fsyncs take longer because there's more extents we have to write out, the space is more fragmented so cold cache reads are more expensive, the csums are no longer contiguous so they also take up a larger more fragmented area. It has this really terrible cascading effect.
I think Btrfs for a guest F's is best pointed to an LV, rather than qcow2. It's been awhile since I benchmarked that compared to 'qemu-img create -f qcow2 -o nocow=on' which will set xattr +C on the file making it nocow. The nocow xattr helps a lot with this problem.
That said, I do not think the tests involved nesting CoW file systems.
I don't remember the state of it when it was first introduced into Solaris, maybe someones memory is better then mine. Was ZFS better of in 06-07 Then BTRFS is now?
By the way, ZFS was deemed production ready after 4 years of development.
http://m.youtube.com/watch?v=dcV2PaMTAJ4
I am the guy who asked the LZJB question. To summarize my recollection, formal design work on ZFS started in 2001 when Matthew Ahrens started working at Sun. Jeff Bonwick had promised Matthew Ahrens a job at Sun making a filesystem a 6 months to a year before then when Matt was still in college. I am sure that both Jeff and Matt had some thoughts on it during that time, but there was no formal effort until Matt's employment started.
By the way, I was under the impression that ZFS development started in 2001 while btrfs development started in 2007. That would be a 6 year difference.
ZFS does suffer from read-modify-write on partial record writes. The effect of that is apparent in the benchmarks. However, the benchmarks are being done on mechanical disks, which have low IOPS. The IOPS of a mechanical disk are roughly the same on a given sequence of IOs at different positions regardless of whether they are 4KB or 128KB in size, so it only has to pay a penalty once. If the record size were changed to 4KB, this penalty would disappear and ZFS performance should increase, provided that the VM internals are properly aligned.
Also, read-modify-write overhead reduces IOPS by at most 2 and bandwidth to the smaller of the link bandwidth and the IOPS times the record size. A CoW filesystem should be able to perform roughly at that level when it does read-modify-write on records/extents. Unless btrfs' internal extents are huge, there is an issue somewhere. Of course, having huge extents by default on which read-modify-write is done could also be considered a design issue.
I ran into issues with maintaining a huge ~1M file Maildir, it seemed to do very badly with huge directories. Some kernel thread would be at 99% CPU while I was trying to populate a Maildir, and the entire system ground to a halt.
More importantly I would run into issues like running out of space on / with 20G left, but it was the "wrong kind of space". I.e. I had run out of metadata space but had plenty of file space left and had to run a rebalancing operation on the filesystem.
I didn't need any of the COW etc. advanced features that btrfs provides, and didn't test them, but as just a normal user needing a general purpose filesystem it was a bit too much hassle, especially with needing to administer the filesystem in a manner similar to how I might administer a RDBMS.
I ended up just going back to ext4.
If btrfs could modify that behavior automatically using heuristics it would truly be general-purpose. However there are some inherent tradeoffs of that approach that users need to consider, and I consider that downside a basic constraint of the Unix filesystem abstraction.
I walked away with the impression that btrfs was a powerful tool for certain use-cases, but definitely not something I'd call a "general purpose" filesystem given the need to be fairly knowledgeable about its internals to use it for common desktop use-cases.
To me a general purpose filesystem is something like ext4, it doesn't have amazing performance, it's not bad either, but generally nothing unexpected happens with it and I can just leave it there and don't have to worry about it.
The btrfs filesystem seemed like the opposite of that. A very powerful tool whose power and flexibility made it less general purpose by virtue of needing to keep close tabs on how you were using it.
I am only superficially familiar with btrfs internals, but I do not see any way to implement snapshots without either doing CoW or duplicating the data in its entirety. If you are checking to see if the data is part of a snapshot, then you should be doing CoW.
Seems like the expense of checking could be largely removed with clever enough metadata caching, which probably noone has had time to implement.
The other aspect that I haven't talked about is our fsync performance is kind of shit compared to other fs'es. Now this does get better in the nocow case but it's still pretty heavy and needs optimization.