TFS: A file system built for performance, space efficiency, and scalability
github.com
github.com
I understand this filesystem is still nascent, but shouldn't data integrity at least be one of the design goals?
What makes you think it isn't? It definitely is. In fact, it borrows several ideas from ZFS wrt/ integrity.
For example, it uses parent block checksums like ZFS.
The first section in the README is called "Design goals", with 13 items. None of them is "data integrity", and none of them even talks about validating the data or handling any failures aside from power loss.
By contrast, in the canonical slide deck on ZFS[1], the first slide talks about "provable end-to-end data integrity". In the paper[2], "design principles" section 2.6 is "error detection and correction".
I'm glad to hear that's also a focus for TFS. With ZFS, the emphasis on data integrity resulted in significant architectural choices -- I'm not sure it's something that can just be bolted on later. As a reader, I wouldn't have assumed TFS had the same emphasis. I think it's pretty valuable to spell this out early and clearly, with details, because it's actually quite a differentiator compared with most other systems.
[1] https://wiki.illumos.org/download/attachments/1146951/zfs_la...
[2] http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.184...
Fair enough.
Both ZFS and Btrfs were initially developed by really high caliber people and experts with good track record. ZFS had five years for full time development until release, next five years to get close to the features and stability that ZFS has now. Btrfs started 10 years ago and it's still trying to catch up.
For good reason in my opinion. An unstable kernel will cause application glitches, unnecessary slowness or OS crashes. They are annoying but not persistent (ie a reboot and you can carry on for a bit). However an unstable file system are persistent and thus could destroy all of your data forcing you to recover from backups (assuming you're diligent enough to keep tested backups).
Plus file systems still have to deal with buggy consumer hardware and other similar edge cases (eg storage devices, power failures, etc) just like a kernel would.
Also, I'll point out that Apple just dropped a new FS on millions of devices, with no issues... that was developed in 3 or 4 years. I'm still blown away that they pulled that off.
It's an impressive feat, regardless of the differences in target devices. Even with the hardware configurations well known, the fact that it was done at such a large scale successfully means that even unusual edge conditions didn't crop up.
This shouldn't be downplayed, it actually speaks to why it's important to have incremental stages of software delivery. First target highly constrained environments (iOS, Watch OS, tvOS), then work on the more difficult and less constrained general computing environment.
Meanwhile commodity hardware will accept a flush to disk command, and return true when it hasn't flushed it to disk. This is done for performance reasons.
I also note that Apple deployed APFS by default only on the most locked down of the two OS's they produce: it's not the default file system on macOS where there's a much wider assortment of user intervention and hardware availability.
And HFS+ had enough of its own issues, hence why it was being replaced, that it's entirely plausible the problems are fewer expressly because of the change in file systems. I imagine their intent and hope was exactly that or why bother?
I wouldn't compare the two.
I went to the OpenZFS website (http://open-zfs.org/wiki/Main_Page) to try and get a feel for the amount of work currently going on. They have some videos and slides, they have a mailing list. The mailing list is mostly full of messages via GitHub but at least that lead me to https://github.com/openzfs/openzfs which I was previously unaware of and which I had not found when trying to look for an OpenZFS specific repository.
As far as I have come to understand, the goal of the OpenZFS project is to merge changes made to ZFS by Illumos, FreeBSD, ZFS-on-Linux and other projects. Their videos might answer this but I wish their website had a clear and simple overview of what has been done. Most of what I can find on their wiki is various ideas for things that they want to do, but I can't know if that means that the OpenZFS project is mostly talk and not so much action, or if it's just that they have done a lot of those things but because they are busy doing they don't have time to write about it. Could be that they have some pages on the wiki I haven't seen also but in that case I think it should be organized better.
I know that Joyent picked up several highly skilled software engineers and programmers that used to work for Sun, and that their SmartOS operating system builds on code descendant from OpenSolaris and that ZFS is as integral on SmartOS as it was on Solaris, perhaps even more so on SmartOS. I have not seen mention of SmartOS on the OpenZFS wiki though.
All in all I feel that ZFS is still being actively developed and maintained by many people. But to what extent they are able to cooperate as much as a lot of them seem to want to I would like to know. It would be a shame if btrfs overtook ZFS simply because the development of ZFS was too fragmented :/
My strong impression is that the ZFS code base is a sufficiently big and tangled thing that no one wants it to fragment. With bug-fixes and improvements happening in Illumos, ZoL and FreeBSD both want to be able to incorporate them on a regular basis and to push their own changes upstream to reduce the maintenance burden of carrying those changes.
ZFS was started internally at Sun with a bunch of resources dedicate from people who had been involved in and thinking about storage devices and file systems and their myriad problems for a long time well before they started ZFS. And the cat wasn't out of the bag to end users for about 4 years, during which time there wasn't the contributor pile on effect.
Btrfs had to contend with some balancing act of not saying "no thanks" to patches, meanwhile those contributions very likely did clutter up the code every bit as much as allowing the contribution aided in hyping the project rather than turning people away from contributing at all.
As for raid56, ZFS emerged before cluster file systems like Ceph and Gluster were even conceived. The companies that need to store tons of data, and fund various storage related projects don't care nearly as much as they once did about raid56. They can just replicate the data elsewhere using Gluster and if a whole brick, be it XFS or Btrfs implodes, they can just make a new brick and the data gets replicated again.
Anyway, ZFS and Btrfs are really not comparable even though on the surface they seem to do really similar things.
I don't know if the authors are here, but if they are - would you comment on fragmentation and the dangers of growing a filesystem past 95-98% full ?
In the world of ZFS, performance can become significantly degraded with as low as 90% space filled. Further, our experience has been that you can permanently degrade filesystem performance by churning the usage above 95% for any significant amount of time. Which is to say, even if you reduce usage back down to 80%, the zpool maintains poor performance until it is destroyed and recreated.
This is exactly what you would expect to see with a fragmenting filesystem that has no defrag tool.
Unfortunately, creating a defrag tool for ZFS is a very daunting technical hurdle and it appears that nobody is interested in pursuing it.
How does TFS behave ? Does it have, or do you plan for it to have, a defrag utility ?
> I don't know if the authors are here, but if they are - would you comment on fragmentation and the dangers of growing a filesystem past 95-98% full ?
Fragmentation isn't an issue in TFS, at all. Because it is a cluster-based file system. Essentially that means that files aren't stored contagiously, but instead in small chunks. The allocation is done entirely on the basis of unrolled freelists.
This does cause a slight space overhead (only slight, coming from the fact that metadata of the file is stored in the full form), but it completely eliminates any fragmentation.
Rotating disks are surely affected by not contiguous operations, even if done with sector granularity.
I think you mean contiguously.
(I mean on ARX generally. Agree about Speck.)
1) Are built from constant time operations, which means they are naturally resistant to side channel attacks (timing, cache, power, etc).
2) Are far simpler in their construction. This makes them easier to reason about and analyze.
3) Related to #2, this also makes them really easy to implement, which means less likelihood of some coding mistake.
Beyond that, most recent ARX ciphers also have a few other advantages over AES. For example, Threefish has a built-in tweak field, which makes using it infinitely easier in practice.
EDIT: In case you're hungry for more detailed explanations, I highly recommend reading the papers for Salsa/Chacha and Threefish. They're very well written, easy to understand even if you don't have a lot of experience with cryptography, and they have sections that explain the design decisions in enlightening detail.
* The CPUs that have better-than-ARX (like, fast constant time multiplication) can do better than ARX
* Fast ARX ciphers are still slower than Intel AES hardware.
I like Salsa/ChaCha more than AES, but there's a reason AES is so popular, and it's not incompetence or a conspiracy.
Well, my points still remains. You need to store IVs/keys/etc. which makes it pretty unsuitable for a file system.
TFS uses Speck in XEX mode (edit: I just realized you're the author, so you already know that :). tptacek wrote a nice post about it: https://sockpuppet.org/blog/2014/04/30/you-dont-want-xts/
and I quote:
"If you’re encrypting a filesystem and not disk blocks, still don’t use XTS! Filesystems have format-awareness and flexibility. Filesystems can do a much better job of encrypting a disk than simulated hardware encryption can."
Edit: check out this presentation on how encryption was bolted on ZFS: https://www.youtube.com/watch?v=frnLiXclAMo (slides: https://drive.google.com/file/d/0B5hUzsxe4cdmU3ZTRXNxa2JIaDQ...) It's not perfect, but it provides data authentication by reusing checksum fields for storing MACs.
Edit 2: also check out bcachefs encryption design doc: http://bcachefs.org/Encryption/ (also not perfect, but uses proper AEAD — ChaCha20-Poly1305. I sent some questions and suggestions to the author, but received no reply :/)
That said, I'm very inclined to distrust a very young cipher... released by the NSA, no less[2]. Which is an add-rotate-xor cipher[2], for which we already have the more reviewed ChaCha20[3], suggested for TLS 1.3[4].
[1] https://github.com/redox-os/tfs/blob/master/README.rst#faq
[2] https://eprint.iacr.org/2013/404
Here is a reason why you can't trust US Gov for creating secure encryption: https://en.wikipedia.org/wiki/Dual_EC_DRBG
At the end of the day the odds of this filesystem gaining any traction are basically 0. I'm just pointing out that not supporting an encryption algorithm that can meet the FIPS requirement makes it basically a non-starter for any commercial application.
In the US. In many other countries, it might gain lots of traction.
- AES-256-CTR is about ~25 % faster with AES-NI than ChaCha20. The gap widens when we consider their respective AEAD constructions.
- AES construction has had simply different design goals than ChaCha20. Software implementations took the back seat, Hardware impls were important and AES is unproblematic for those.
- AES is arguably a rather conservative design. To some extent even twenty years later.
See that does sound like a good idea - I've always observed HDDs with an SSD cache to have a phenomanly useless caching system
Without that, I think one should focus on getting the on disk structures right and the code manipulating them solid. That is enough of a challenge.
A good design would look at the state of the art and use the best techniques available. If the aim was research, then try one new thing, not a thousand.
For actual promising new filesystem efforts I'd be looking at HAMMER2, Tux3 and F2FS.
That's what it does: It takes from many sources (although mainly ZFS).
What is the difference here between a hardware failure and a unexpected power failure?
TFS provides following guarantees:
- Unless data corruption happens, the disk should never be
an inconsistent state \footnote{TFS achieves this without using
journaling or a transactional model.}. Poweroff and the alike
should not affect the system such that it enters an invalid or
inconsistent state.
Provided that following premises hold:
- Any sector (assumed to be a power of two of at least
\minimumsectorsize bytes) can be read and written atomically,
i.e. it is never partially written, and interrupting the write
will never render the disk in a state in which it not already
is written or retaining the old data.
Data corruption can break these guarantees or premises, and TFS
encompasses certain measures against such corruption, but they are
strictly speaking heuristic, like any error detection and correction
method, as the damage could be across all the disks.Sadly, a lot of hardware has much more complicated behavior than that. A lot of RAID cards in particular will lie about what was actually written, so in a power-off scenario later writes might have made it while earlier ones didn't, writes can be incomplete, etc. I don't mean this as a knock against TFS. It's more of a suggestion that the fault model be expanded to include at least a few more possibilities.
- historic versioning (like in git)
- deduplication
- authenticity through cryptographic hashing like in block chain
- distributed delivery like bit torrent
I don't think, that we need another ZFS or btrfs competitor. Just my 2 cent.
[1] https://ipfs.io/
It is very similar to ZFS.
Really poor choice of name. If they wanted a modular replacement for ZFS then call it MFS or ModFS or something...
Everyone makes name conflicts in independent domains to be a much bigger problem than they are in reality.
There’s Amazon rainforest, Amazon the ecommerce website, and Amazon the cloud company. How often do you have any problem differentiating between them?
Though, they've started to refer to the source control protocol implementation as TFVC, since TFS supports git as well now. It does seem to have some conflicts, and even notes another file system called TFS themselves.
In this case, I'm pretty sure another name might be a better idea. Hell, TFS the version control system and the other file system are more well known than Firebird the database when Mozilla renamed their shiny new browser.
I hear C is very search engine friendly. Probably why it’s so popular.
That said, I wouldn't expect a "new" programming language called "Java-Script" (not the current ES/JavaScript) to gain traction. Or for that matter a programming language called "Coffee" to be very successful either.
If you want Microsoft’s TFS just search for “Microsoft TFS.” There is no confusion.
But yeah, if I wanted a set of three letters than suggested quality and reliability, TFS would be near the bottom of my list.
All the other ones I know require creating scripts that will manage branches and merges into the main branch.
On TFS I select a check box and go off doing something else.
I don't want to have a CI build that is green only a few days when all planets are aligned.
Surprise! Now you have 2 versions of an app to support, with twice the resource requirements.
I and everyone else on the team enjoys the fact that broken code produced by offshore developers that should have been fired in first place, never touches the official development repository.
Those developers will eventually produce something that passes the unit tests and gets merged, instead of borking the build for weeks.
Should the process done in another form with code reviews and such, yes but that isn't how many enterprise projects are managed.
Note that I am only referring to those that can't learn to code even if we ELI5 them.
There are others on the offshore teams that are highly skilled, but like us onsite, cannot do anything to change the rules of the game.