Updated MinIO NVMe Benchmarks: 2.6Tpbs on Get and 1.6 on Put
blog.min.io
blog.min.io
This benchmark uses large batch size, 64MB, to test. There is nothing new here. Most common file systems can easily do the same.
The difficult task is to read and write lots of small files. There is a term for it, LOSF. I work on SeaweedFS, https://github.com/chrislusf/seaweedfs , which is designed to handle LOSF. And of course, no problem with large files at all.
LeoFS last release is about 3 years ago (v1.4.3 February 20th, 2019), and seems more complicated. SeaweedFS is still growing and has being released on a weekly basis. Just ask if you need any new feature.
Also, think about how to manage. For example, grow capacity. For SeaweedFS, you just need to add one server and point it to the master. That is it!
SeaweedFS is portable, the data file and the metadata.
https://blog.min.io/scaling-minio-more-hardware-for-higher-s...
It reads like they roughly doubled the read (GET) throughput for 32 nodes. Though I don't know how much of that would be AWS improvements vs MinIO improvements.
It was only when GCS reached 1.0 Tpbs that we started shifting workloads over.
However, these benchmarks are only relevant within the same region. As you move further away, roundtrips and speed-of-light will become significant enough that I've seen some people do benchmarks in Tpbms.
I would imagine 1kb objects to be far more common...
64MB is far, far closer to our usecase than 1kb. I don’t even think our schema files/metadata which are arguably the smallest part fit into 1kb.
I would call an object store in which the objects are mostly < 1 kiB a key value store.
1kb object would be terribly inefficient if you have too many of them. It's slow even with regular file system, copy 1 million 1kb files would take forever. For small objects, it's best to pack them into larger block before storage, like how a database does with its data.
I think the document should have a clearer guideline about best practices when deploying a SeaweedFS cluster for different situations: small cluster, large cluster, H/A cluster etc..
For filer metadata, you should just pick the one you are most familiar with.
There is a wiki page for production setup. https://github.com/chrislusf/seaweedfs/wiki/Production-Setup
Small objects, 4K ~ 4MB, are very common in machine learning area(audio, video, text), or surveillance, game assets, etc. SeaweedFS saw a lot of usage in these areas.
I was also surprised by many of the comments. I had only played with minio a bit, and was considering using it. The durability comments were concerning, and there were other issues too.
One example is this bug: https://github.com/minio/minio/issues/8873
Basically, it was treating object names like foo//bar in them the same as foo/bar, except for sharding, which thought they were different.
Their fix was just to disallow '//' in an object name, even though other s3-like implementations allow it.
We have a lot of products that use S3 API and Azure was the only cloud that did not offer that out of the box - at first I could not believe this was the case, because even smaller providers like Linode or Digitalocean offer S3 API compatibility. But then I found the post linked below.
[1] https://cloudblogs.microsoft.com/opensource/2017/11/09/s3cmd...
Still, I wish Azure just added a built-in compatibility layer, like virtually every other cloud provider - I'm not a big fan of having to spin up a container just for this reason.
It doesn't have to use local disks. A big part of what it does is to provide a S3-compatible API, which you can use with any number of backends. Even other cloud providers. Say you want to be able to deploy on prem, on AWS, and GCP. You can add MinIO in front of all of these and you app won't care.
Of course, if you are writing to an actual disk you'll have to figure out the backup part.
The other people I tend to end up with always try to sell people on minio+rook-ceph and then offer support along the way. So basically you buy their kubernetes deployment and then you have to pay them for the rest of your life for troubleshooting. The TrueNAS seems cheaper to me.
I don't see a good backup story, but maybe I just don't know it.
Mainly use it for gitlab object storage current. Logs docker images etc
Also by the time most people account for 2/3 copies of the data + 1 completely offsite backup of data I think their eyes might start to water at the cost if you want reasonable performance as well. I experiment with this kind of stuff a lot of the time and always am surprised at how much more than you'd expect it costs to get drive, node, and region level redundancy with backups. You need essentially at a minimum 3TB for every 1TB of usable storage assuming regular RAID1.
Plenty of good use-cases.
I do agree with the Ceph being much more complex which is why I hope OpenEBS implements zfs-localpv send recv for a poor man's replication.
A pure kubernetes deployment is more complex(although it's all part of the same binary i think).
I could be completely wrong though.
My biggest wishes for Truenas Scale:
1. NVMe-oF support (over RDMA or TCP)
2. aarch64 uefi iso
I could see Truenas Scale overtaking Proxmox soon in the KVM space, their API and UI are already more enjoyable to use IMO.
Where do data-management SaaS companies like Rubrik.com and Druva.com fit in? Are they not popular enough solution to secure min.io deployments?
Now, my seaweed setup is a 3 vm cluster with 3 disks per vm (1TB) each. I configured a wireguard mesh (https://github.com/k4yt3x/wg-meshconf) between the VMs and configured master and volumes server to talk to each other via wireguard IPs securely. I also configured ufw to only allow communication between http/gRPC ports. I also configured a filer (using leveldb3) to use wireguard IPs (master and volumes) and let it communicate with some specific servers on the outside (ufw).
After that i mounted the filer via weed.mount on that specific server and tried to copy over the same files/folders. after 2 days i copied over about 1.5 TB of the data via rsync. There was also no problem with file listing and accessing the filer from different machines while uploading stuff. But there is a overhead when reading and creating lots of small files. File listing is even faster than local btrfs file listing.
chris is also very nice and fast fixing bugs.
btw: If you use "weed filer.copy", it should be much faster. Rsync needs to go through FUSE mount.
Not everyone can use s3 directly,for various reasons mix of legacy/ existing NAS systems or compliance/ regulatory requirements, not on AWS stack but most libraries/clients support S3 APIs even if you use GCP/Azure you could have S3 compatible API with minIO.
Even for something like just querying logs, you can churn through a few TB pretty quick by searching in parallel
AWS S3 recently launched strict consistency.