File systems unfit as distributed storage back ends: 10 years of Ceph
muratbuffalo.blogspot.com
muratbuffalo.blogspot.com
It seems like they could go one further and eliminate the ext4 underneath.
Talks about Facebook, Instagram, S3 and other Object Store services and how they deal with storage at scale.
C.f. What linus is saying in https://news.ycombinator.com/item?id=21673372 except turn it around. When an interface has devolved into two sides hating and Postel's-law-enabling each other ad infinitum, and a statement like his is actually justifiable, it's time to close up shop and move on. Nothing good will ever come from POSIX-like storage ever again, and any storage system built around it is doomed to be a mess of too many layers and also too many layer violations. Utter hopelessness.
Except that nobody will sign on.
Look at what happened to FreeBSD in the 5.0 timeframe when they reworked their storage layers into GEOM. It was a NIGHTMARE. Most people agreed it needed to be done, but there was an excruciatingly loud segment who complained incessantly. It took some gigantic brass balls and asbestos-lined flamesuits on the part of FreeBSD heavy hitters to drive it through.
If the system in Linux is to get fixed, Linus would probably have to step in and pronounce.
I've been thinking about transitioning entirely to sqlite for all my data.
You can use something like libsqlfs [1] for POSIX file heuristics with sqlite as the backing store.
One HA single primary/multi-master solution to use sqlite may be drbd.
[0] https://www.sqlite.org/fasterthanfs.html [1] https://github.com/guardianproject/libsqlfs
Bonus points for the `lsm1` extension of sqlite3 which allows you to use it as a key-value store, which I used with mixed success (if I could remember the key names that seemed the most logical thing in the world last week, lol).
There's nothing to it, really. sqlite3 is a very mature software and save for a mechanical failure of your storage drive, the odds of it losing your data are practically zero.
For even more bonus points, encrypt your sqlite3 storage. That way you can freely distribute it on Git hosting services.
- AFS as a federated posix file system for user home directories. My impression is that a distributed posix filesystem is... well... hard, for basically the reasons listed in the link. We're actually trying to phase it out, starting by reducing the size of the federation by cutting off access outside the CERN network.
- A few in-house developments like xrootd [1] (basically CERN's version of an object store) and EOS (a posix file system built on top), to store data. These projects have their roots in a time when CERN was at the forefront of "big data" and it made sense to develop an in-house project. These days there are a number of alternatives and my impression is that the reasons for continuing the projects are mostly historical.
- For read-only data we have cvmfs [2], a FUSE module which is synced to some other file system a few times a day. Making it read-only simplifies the metadata handling considerably: it's actually quite nice for a CERN project.
- Some people have started using Ceph for more experimental things, but in general these "industry" projects are only starting to replace the home-grown ones.