SeaweedFS – A simple and highly scalable distributed file system with S3 API
github.com
github.com
> Usually distributed file systems split each file into chunks, a central master keeps a mapping of filenames, chunk indices to chunk handles, and also which chunks each chunk server has.
> The main drawback is that the central master can't handle many small files efficiently, and since all read requests need to go through the chunk master, so it might not scale well for many concurrent users.
The chunk server architecture has been first put into production with the Google File System AFAIK. And it has been designed specifically for large files (what search needed at that time). So no surprise.
But that's only one architecture for a DFS. There are also block-based DFS (like GPFS), object-based DFS (Lustre.), cluster file systems (OCFS), and other architectures. They exhibit other characteristics.
Telling from the architecture and Wiki, it does not seem to be a file system at its core, but an object store with a file translation layer. One of the core problems of this approach is that in-place updates usually mean read-modify-write (if the object store has immutable objects, like most have, with Ceph being a notable exception).
From the replication page:
> If one replica is missing, there are no automatic repair right away. This is to prevent over replication due to transient volume sever failures or disconnections. In stead, the volume will just become readonly. For any new writes, just assign a different file id to a different volume.
This sounds like the architecture and implementation is still pretty basic. Distributed storage without redundancy (working redundancy!) is not that interesting.
Sorry to be that critical (great that someone writes a distributed file system!), but I think it is important to add some context. And the seaweed auther seems to have problems with bold statements either...
Disclaimer: I also work on a distributed file system (with unified access via S3 ;)
The blobs are read-modify-write, but a file can have many blobs. So the whole file does not need to be read-modify-write, unless it is a small file.
The redundancy is managed via scriptable admin commands.
I understand your unit of smallest write is a blob. How large are blobs? How does their identifier look like?
By default, the filer client uses 8MB.
Each identifier is as <volume_id, file_key, cookie>. With volume_id, you can locate the volume server. So the volumes are portable and can be moved around.
b) It does have redundancy through data replication unless I am missing something ?
Additionally, something more complex than N-replication (eg. M.N erasure coding) is also a strong 'redundancy working' requirement for me.
b) It is completely unclear how the replication works, and which properties it has (split-brain safe?). From what it states (see citation), the "repair" is not automatic.
On the other hand, SeaweedFS is has no API access fees and faster than S3 with your own hardware. Not sure what Amazon may do.
https://adoptopenjdk.net/sponsors.html
https://en.wikipedia.org/wiki/List_of_Java_virtual_machines#...
It is Google's own version of J++ that needs to be taken care of.
This is one of the biggest problems with Java, it’s free but not open source (Eric Raymond has been fighting this for a while now that free does not mean open source). Oracle can come after anyone using Java if there is substantial money to be made.
So Java is only free as long as money keeps coming in for Oracle, once it dries like for SCO, they will go after anyone using Java API to extract money.
Also this shows subtle difference between open source and free. In a legitimate open source licensed software copyright and right to change is granted and company cannot sue others just based on copyright like Oracle can sue anyone using Java API.
I was specifically implying block storage solutions that offer S3 API combability (i.e. use our AWS S3 competitor with matching S3 API).
[1] https://www.digitalocean.com/products/spaces/ [2] https://cloud.ibm.com/docs/cloud-object-storage?topic=cloud-...
They did it well:
> Most other distributed file systems seem more complicated than necessary.
> SeaweedFS is meant to be fast and simple, in both setup and operation. If you do not understand how it works when you reach here, we've failed! Please raise an issue with any questions or update this file with clarifications.
https://github.com/chrislusf/seaweedfs#compared-to-other-fil...
However, since I never had to touch hdfs after installing it in the first place, I wonder what the difficulties in operation are, that they tried to overcome here?
Also when you lose datanodes you will need to rebalance the data which SeaweedFS apparently does not need to do.
Agreed ! I'm just very glad, they didn't use the word "lightweight", it's a pet irritation of mine, when ppl "advertise" their software as "lightweight" !
The "first few" versions are always "lightweight" once they included all the bug fixes and corner cases of the other solutions over the next few releases, they seldom stay "lightweight" :/
Yet I do try very hard to keep each layer separate from each other. So each layer can be "lightweight". :)
https://github.com/mogilefs/mogilefs-docs/blob/master/HighLe...
It's an old system from the folks @ Danga, but the mailing list still sees random activity now and then...
There are many problems with large number of small files. It's more efficient to batch them together.
Only ambry(linkedin) AFAIK, but has no erasure coding.
> SeaweedFS Filer metadata store can be any well-known and proven data stores, e.g., Cassandra, Mongodb, Redis, Elastic Search, MySql, Postgres, MemSql, TiDB, CockroachDB, Etcd etc, and is easy to customized.
I'm not very familiar with other DFS's but at the very least glusterfs stores metadata as xattrs on an underlying filesystem and so has no need of an external data store.
Also, SeaweedFS has a "master" server (single centralized with failover to secondary) and "volume servers" (responsible for data).
You can move metadata to any faster scalable data store. And/or move the data to any cheaper/faster/larger storage, in cloud or in your garage.
Actually I do not understand why MinIO is slow here.
Is there anything like that out there (other than Unraid, which I kinda don't like)?
It is archiving and serving more than 40,000 images on a webapp I built for the small team I work with.
I run SeaweedFS on two machines and it serves all images I host.
I wanted to kick the tires because I was always fascinated by Facebook's Haystack.
It has been simple, reliable, and robust. I really like it and hope if one of my side projects ever take off at some point, I get to test it with a much bigger load.
But note that it doesn't work well on Kubernetes and POSIX FS interface isn't perfect but then I haven't found any that are.
But I don't know if there's a major shop that uses it. Anyone knows?
There are a few other companies, but not as famous as Apple. :)
And no I'm not particularly fond of the namr CockroachDB either.
Make it as silly and unique as possible, I'll be able to search for solutions later
Silly? Not my preference. Unique? Absolutely. Something made up or completely out of context is likely to be a whole lot easier to search for.
"Golang" works much better, but sometimes results in a pedantic correction "The language is Go not Golang", to which I can only reply that the project is hosted on golang.org. Weird a.k.a. unique names are useful.
CockroachDB: Cockroaches are hard to kill. The vendor wants to imply this is also the case for their platform.
It will be nice if you can integrate it with pyfilesystem [1]. It’s a small library providing unified file interface in python for local or cloud storage without fuse. It can already support seaweedfs using S3 API, it will be nicer to just directly have a driver for seaweedfs removing S3 translations.
I like dried crispy seaweed wafers, they do produce umami flavours, so it is a good name, reminding them. The quality of software isn’t based on name but by what it does, so don’t bother about people talking about name.
I am not good at Python at all and would not able to add Python specific support. However, SeaweedFS has gRPC APIs that you can use to do all file operations.
Take a look at Gasper (https://talhof8.github.com/gasper).