Lustre, BeeGFS are HPC filesystems, i believe. No data locality. Amazon EFS - feels like a MapR clone, but they expose some tunings - higher latency for higher throughput and larger clusters. We don't make that trade-off. I don't know enough about Qumulo and Objective FS to comment on them.
Then you haven't done your homework, since Gluster has had that architecture for a decade. And it supports real POSIX semantics, too - not just the HDFS subset.
Reference: http://www.admin-magazine.com/Articles/Comparing-Ceph-and-Gl...
It might also be useful to point out some semi-useful distinctions regarding levels of distribution. The PVFS2 lineage has had fully distributed metadata for even longer than Gluster. OTOH, both Lustre and Ceph have separate metadata servers, more than one but still fewer than there are data servers. That doesn't necessarily limit scalability, but it does increase operational complexity. I'm honestly not sure about BeeGFS or GPFS; I don't pay much attention to them since one is only fake open source and the other doesn't even try. If we want to include proprietary offerings we could also add EFS, ObjectiveFS, OneFS, PanFS, etc.
BTW, for those here who know me, at $dayjob I've switched from working on Gluster to working on a more HDFS-like data store. (They chose not to call it a filesystem because it isn't one, which I consider a refreshing bit of honesty in a field full of pretenders.) It scales way beyond any of the things mentioned here, but it's proprietary so I can't compare notes as much as I was able to with Gluster. :(
What Ceph and many others do is support partitioned metadata. One volume == one metadata partition. But cross partition metadata operations do not have the same semantics as intra-partition metadata operations. So, moving a file/dir between volumes (partitions) is not atomic - they do not solve the hard consensus problem. We do. Gluster is also not solving the consensus problem. The end result is leaky filesystem abstractions. If you move a file between two directories, it may go fast. But between two different directories (across volumes), it can go very slow. Renaming a file in Gluster can cause it to move between hosts.
No. It does not. Within a volume, files will always be renamed in place on the replicas where they currently reside. There are other problems with rename in Gluster, having to do with how the locations of files are tracked when they're not where they're "supposed" to be according to the hashing scheme, but movement of data as part of rename is not one of those problems.
"Across volumes" is not an interesting or relevant case, because the whole point of separate volumes in Gluster (and most other distributed filesystems) is to provide complete isolation. You seem to have this idea that users would want to have lots of volumes and still move data around between them. Is there something weird about your architecture that forces them to do that? Small volumes are operationally complex and stand resources. Generally, you'd have just one volume. Access control, quota, etc. could still be applied at the namespace or directory level, below volumes. It worked rather well at the world's largest Gluster installation, which I helped run for a year and a half, and I believe most other installations are similar.
Please try to learn about other systems before you make wild claims about them, lest your audience perceive them as deliberate lies to benefit your business. I don't think your claims about Ceph are anywhere near accurate either, but I'll leave those for one of my Ceph friends to address.
That's a very unique definition of "distributed". To many in the field, "spans multiple hosts" alone is sufficient. Some might refine it to preclude single nodes with special roles or by adding requirements such as a single view of the data etc., but it's a joke to add atomicity or implementation details such as how a consensus algorithm is used. There are plenty of distributed data stores that don't provide atomicity for all operations, and users are happy with that. There are plenty that use leader election instead of consensus to deal with consistency issues, and users are happy with that too. Inventing a novel and highly specific definition to exclude them seems more than a bit disingenuous. If we want to go that route, I'll point out that HopsFS is misnamed because it's not a filesystem according to commonly used definitions. There are many hard problems that it wimps out on, that a real filesystem must address.
"Operations that cross volumes do not solve the consensus problem, so they do not execute operations like move atomically."
Again, you seem to be using "volume" in a very unique way. Your architecture might misuse "volume" to mean some sort of internal convenience that's practically invisible to users, but others don't. To most, a volume is a self-contained universe, much like a virtual machine. Users are well aware that moving data between either will not have the same semantics as moving data within.
Such desperate attempts to disparage systems you obviously haven't even tried to understand are not helpful to users, and are insulting to people who have spent years solving hard problems to address those users' actual requirements. Believe it or not, systems exist which don't work like yours, and which prioritize different features or guarantees, and there's demonstrated demand for those differentiators. If you want to learn about those differences, to compare and contrast in ways that actually advance the state of the art, please begin. I say "begin" because I see no evidence that you've attempted such a thing. Does your academic employer know that you're misinforming students?
I didn't mean to imply otherwise, just to emphasize the hat in the way of the talking, assuming words mean what everyone knows. (I don't know if Gluster has moved too far from its roots to be an HPC filesystem too.) Good observations about distribution and misleading marketing.
For anyone still following and interested, OrangeFS uses Giga+ [1] to distribute metadata. I don't know where to read about the design for the other free systems mentioned available for study, and I've heard people being rude about Lustre's design anyway.
1. http://www.pdl.cmu.edu/PDL-FTP/PDSI/fast2011-gigaplus.pdf
Isn't that the case for S3 as well? I haven't seen AWS mention it explicitly, but the consistency guarantees and the "unlimited scalability" claims seem to point in that direction.
reference: http://docs.ceph.com/docs/master/cephfs/add-remove-mds/
The Ceph team considers configurations with multiple active metadata servers to be stable and supported; where is your evidence that this is false?