MinIO: A Bare Metal Drop-In for AWS S3
tech.marksblogg.com
tech.marksblogg.com
I've been running Ceph for about a year now, and the start up was a bit rough. We are actually on second hand hard drives, that had a lot of bad apples, and the failures weren't actually very transparent to deal with, which was a bit of a disappointment. Maybe my expectations were too high, but I was hoping it would just sort of fix itself (i.e. down the relevant drive, send me a notification, and ensure continuity). I feel I had to learn way too much about Ceph to be able to operate it properly. Besides that the performance is also not stellar, it apparently scales with CPU frequency, which is a bit astonishing to me, but I've never designed a distributed filesystem so who am I to judge.
I was looking for something that would scale with the company. Now we've got 70 drives, maybe next year 100 and the next year 200. Now all our drives are 4TB, but I'd like to switch them out for 14TB or 18TB drives as we go along. We're not in a position to just drop 100k on a batch of shiny state of the art machines at once. Many filesystems assume the number of drives in your cluster never changes, it's crazy.
What were some of the other hurdles you faced in your Ceph deployment?
So RBD is easy, S3 is somewhat more complicated as you need to run multiple gateways, but still very doable. The FS stuff also needs extra daemons, but I have not yet tested it.
I found Ceph's error messages very hard to debug. Just google around a bit for the manuals of how to deal with buggy or fully defective drives. There's a lot of SSH'ing in, running vague commands looking up id's of drives and matching them to linux device mount points and reading vague error logs.
To me as a high level operator it feels it should be simple. If a drive supplies a block of data, and that data fails its checksum, it's gone. The drive already does its very best internally to cope with physical issues, if the drive couldn't come up with valid data, it's toast or as close to toast as anyone should be comfortable with. So it's simple, fail a checksum, out of the cluster, send me an e-mail, I don't get why Ceph has to be so much more complicated than that.
If I recall correctly there isn't really a way to modify an existing file via HDFS, so you'd have to copy/edit/replace. Append used to be an issue, but that got sorted out a few years back.
Erasure coding is available in the latest versions. Which helps with replication costs.
I think, HDFS may just be a simpler setup than other solutions. (which is to say its not all that simple, but easier than some other choices). And I wouldn't use HDFS as a replacement for block storage, which is something I've seen done with Ceph.
Asking since we're in the same position as yourself w/ high double-digit disks trying to figure out our plan moving forward. Right now we're just using a very large beefy node w/ shelves. ZFS (via TrueNAS) does give us pretty good guarantees on failed disks + automated notifications when stuff goes wrong.
Obviously a single system won't scale past a few hundred disks so we are looking at alternatives including Ceph, GlusterFS, and BeeGFS. From the outside looking in, Ceph seems like it might be more complexity than it's worth until you hit the 10s of PB range with completely standardized hardware?
There are loads of issues like this on their github: https://github.com/minio/minio/issues/8873
Sounds like the shard hashing happens before or after object name normalization depending on the operation. Ouch.
I know this bug has hampered our use of ceph at singlestore. Note that this is not an eventual consistency issue. When it happens the list command will permanently miss files.
I'm pretty sure there was something weird going on with how minio was reading the config state, as I definitely was not the only one hitting it. Luckily I only had to use it for local testing in the project, but the whole thing didn't leave me feeling good.
I guess this makes minio "fast". But it might eat your data. Please use something like Ceph+RadosGW instead. It might be okay for running tests where durability isn't a requirement.
Their attitude about it isn't great: https://github.com/minio/minio/issues/3536
That's too bad, as it seems well thought out in other areas, like clustering.
You might want to look at other options as well like SeaweedFS [0] a POSIX compliant S3 compatible distributed file system.
Both should default to fsync on, with the option to turn it off. So not a great choice of defaults. Again, it probably looks good in benchmarks when people naively compare S3 stores. But it just shouldn't eat your data per default.
For anyone who wants HA and horizontal elastic scalability, checkout SeaweedFS instead, it is based on the Facebook "Haystack" paper: https://github.com/chrislusf/seaweedfs
This comes across as slightly condescending.
As I'm sure you'd agree, secure by default is very important, and it's what most responsible distributions aim for (i.e., Debian/Ubuntu). Starting up a daemon should not launch it in the most open way possible, but instead the most restricted way possible.
A reasonable expectation is that you should not have to pass the -ip flag; daemons should default to a secure configuration (which probably means defaulting to -ip 127.0.0.1, which you should be able to easily override, if that is your intention, and achieve the default behavior by simply passing -ip 0.0.0.0/0).
As does:
> As I'm sure you'd agree, secure by default is very important
I just meant it in a practical sense, you (as in people), need to read a guide in order to make it production ready instead of seafweed being production ready by default. I checked, there even is a guide in the repo, so I guess people need to read it.
I did not heard of this "kitchen" sink before. :)
But adding one more sink should be trivial.
It is also very easy to run. Just run this: "docker run -p 8333:8333 chrislusf/seaweedfs server -s3"
I know I'm asking quite a biased source but are there any shortcomings of SeaweedFS that are well known? Any hangups/weird corners that you can think of just off the top of your head?
For example, if you batch upload data, and the default 8 volumes happen to fill at the same time, you get transient errors until it has managed to create new volumes.
It'd be nice if there were an issue to explain this shortcoming, does the project know it happens and it's like a "hopefully we'll have auto scaling/adjusting volumes/online configuration update in the future" or something? How would one mitigate that?
The project is still growing and there are different edge cases for each features, especially new ones. However, in general, I feel the project is structured layer-by-layer, and should be easy to fix the problems.
Some parts are complicated, e.g., FUSE mount. It's hard or impossible to be fully POSIX compliant. SeaweedFS has come a long way and has improved quite a lot, but maybe do not run your database on it just yet, until SeaweedFS supports block storage later.
The only limitation is that you don't have all the IAM access rules that you get with AWS.
Oh wait, that's exactly why I love it.
And, the mc command line client is awesome.
And, it all runs inside dokku which is incredible.
It’s a complicated tool for sure, but it comes from the natural complication of dealing with auth in a very flexible way.
So you can say it's a "natural" complication, and you'd be right, but that says nothing about usability, which is where "issues" tends to come in.
It causes weird and overly broad privileges though usually, because you need to give permission to do any possible thing the job or user of the credentials COULD need to do, all the time.
This happens because any action to limit the scope usually causes more human friction than it is worth.
Ideally, when it is requested they do something, they get handed a token for the scope they are doing it in, which only gives them access to do the specific things they will need to do on the specific thing they need to do it, and only for the time they plausibly will need to do it for. This is a huge hassle for humans, and adds a lot of time and friction. For machines, it can be as simple as signing and passing along some simple data structures.
So for example, Alice would get a token allowing access to Q4 ‘20 only if that was plausibly correct and necessary, and then only for how long it took to do whatever she was doing. Bob would only get a token to access the specific EC2 instance that needs him to log into it because of a failure of the management tools that otherwise would fix things - and only after telling the token issuing tool/authority that, where it can be logged.
It makes a huge difference in limiting the scope compromises, catching security breeches and security bugs early on, identifying the true scope and accessibility of data, etc.
Also, since no commonly issued token should probably ever provide access to get everything - where the IAM model pretty much requires that a job that gets any random one thing, has to be able to get ALL things, then you also end up with the potential for runtime performance optimizations, since you can prune in advance the set of possible values to search/return.
Not really, unless you mean it needs permission to assume all the roles it could need in order to have the permissions it requires.
It’s the difference between ‘can query the database’ and ‘can retrieve user Clarice’s profile information because she just made a request To the profile edit page’
Does that make sense?
The 'because' isn't there, but I'm not really sure what that would look like, at least in a meaningful (not trivially spoofable) way.
Tokens can do that.
Don't get me wrong, I do see the hypothetical benefit, I'm just having trouble envisaging a practical solution. Is there something else not on AWS (or third-party for it) that works as you'd like IAM to?
Your token grantor is just taking in whatever request state you have (session, permissions granted, whatever), and stuffing them into the token that gets passed around. Then the various back ends and client calls also do that, and where there is a permissions check (or conversion) necessary, say on a backend API call to access something, it checks it the callers token has the right permission.
I don’t want IAM to work that way. I don’t want to use IAM for this? It’s the wrong tool.
There are a ton of various signed token frameworks, all with various trade offs. JWT (ugh), Gaia Mint (internal Google), etc.
It tends to work best where there are multiple layers of services, as you have an abstraction layer you can do checking/audits, etc. at.
Fine, but then that backend handler has permission itself to do whatever it is, regardless of whether or not the token does, since as you said, it COULD need it to service that request?
> I don’t want IAM to work that way. I don’t want to use IAM for this? It’s the wrong tool.
Yes.. I.. agree, I'm confused now why we're even talking about IAM, it's solving a slightly similar problem at a different level; isn't particularly useful here.
And I don't know why, you're the only who responded to my reply to 'What specifically do you have issues with when it comes to IAM?'
Haha
Essentially I suppose I disagree that it's only good for human users with long-lived roles, but I'm not saying it's the right tool for per-request granular authn, and I'd be surprised to learn that anyone is saying or (trying to be) using it like that. IAM's not even for end human users, (as in of your application) nevermind breaking further down into different types of request from them or on their behalf.
Using that as the sole way to Scope process/machine access though IS a weird fit in an automated environment for the reasons I laid out. You either come up with a broad scope that covers everything the job/process/machine could ever need to do or access (and then hope there is no exploit or bug that results in it accessing more), or build something like a token system that lets you get/scope access or permission in the context of the work it is doing on behalf of someone else. Which requires investment, but fits what should really be happening better. That is more of the ‘zero trust’ model, but certainly not all of it.
But overall I’m not sure constantly reaching out to IAM to retrieve scoped permissions for every single action makes much sense. Aside from the obvious latency issues the master set of credentials needs to have permissions to be able to request these scoped time-bound keys, and so them being leaked is just as bad as they can be used to just re-request access to “Q2 data”. Ok, so we need some logic to say “Alice should only be able to request these keys once a day” or some such, and these arbitrary requirements are much more complex to implement and a lot more fragile.
So it only makes sense if you’re expecting it to be materially more common for a service to somehow leak these time-bound single access keys but not leak any other credentials. Which isn’t an assumption that would hold up I think.
So what’s the point?
Minio's client side libraries appear to be packaged separately, and Apache licensed: https://github.com/minio/minio-go
https://github.com/minio/minio-js
(Etc)
That seems very confusing
https://www.dataengineeringpodcast.com/minio-object-storage-...
The notes don't mention this and the audio is over one hour, would you mind clarifying?
And, being S3-compatible at an API level would be a big bonus for a company the size of AWS, especially if it had nearly native compatibility with the aws-cli tool.
This first example from the article sounds very valid, but is still personally funny to me, because it's related to the first use I made of S3, but in the opposite direction (due to different technical needs than the article's):
> If an Airline has a fleet of 100 aircraft that produce 200 TB of telemetry each week and has poor network connectivity at its hub.
Years ago, I helped move a tricky Linux-filesystem-based storage scheme for flight data recorder captures to S3. I ended up making a bespoke layer for local-caching, encryption (integrating a proven method and implementation, not rolling own, of course), compression, and legacy backward-compatibility.
That was a pretty interesting challenge of architecture, operations, and systems-y software development. And the occasional non-mainstream technical requirements we encounter are why projects like this MinIO are interesting.
I've also used MinIO to mock an S3 service for integration tests, complete with auth and whatnot.
One of the advantages of MinIO would be the wide compatibility with other S3 storage services. If my NAS had downtime while on holiday I could spin up a new bucket on S3/Backblaze/Wasabi and backup everything in a few minutes.
One of the challenges we had was "pre-filling" the MinIO server with testdata. Some tests require reading a testfile from our mocked S3 API. We wanted to have those testfiles instantly available at MinIO startup, but couldn not get that to work with docker volume mounts, MinIO simply would not recognize the files and serve a 404.
Has anyone got that working or a proposal for an alternative solution? We resorted to uploading the files via the application on startup (if it is in "testing mode"), but that does feel like a dirty hack.
For example if you have a project which you store objects to S33, At CI pipeline you don’t want to store temp files into S3 for cosr purposes. So instead you store at minio. A company must be crazy to use minio as their real data storage.
One example: missing error handling for interrupted uploads leading to files that looked as if they had been uploaded, but had not.
Both ceph and MinIO's implementations differ from AWS original S3 server implementation, in subtle ways. ceph worked more reliably, but IIRC, both for MinIO and ceph, there is no guarantee that a file you upload is readable directly after upload. You have to poll if it is there, which might take a long time for bigger files (I guess because of the hash generation). AWS's original behavior is to keep the socket open until you can actually retrieve the file, which isn't necessary better, as it can lead to other errors like network timeouts.
I got it working halfway reliably by splitting uploads into multiple smaller files, and adding retry with exponential backoff. Then I figured out that using local node storage and handling distribution manually was much more efficient for my use case.
So for larger use cases, I'd take the 'drop in' claim with a grain of salt. YMMV :)
Can the author let us know if he is happy with the business results of these articles?
Some of my applications rely on S3 in production, but I don't want that dependency when running running the application on my machine - I use minio as a drop-in replacement for development.
Since I use docker compose to handle my app's services (postgres, rabbitmq, etc), adding minio into it is a perfect fix.
This is misleading. While there are no bare metal projects I'm aware of, there are 10+ S3-API compatible S3 alternatives, such as Wasabi, Digital Ocean Objects, etc., to name a few.
I believe that it addresses many of the OP's reasons for deciding against S3.
Can they scale down to "I need to spin up an S3 thing for local testing" for the cost of the storage and CPU?
Am I locked into a multi-year agreement, or can I just go and throw it away in a month and stop paying?