Amazon Elastic File System – Production-Ready in Three Regions
aws.amazon.com
aws.amazon.com
Finally we can start using standard posix semantics for things.
Yes people might say its primitive, but currently all I can see is people trying to re-create shared files systems over protocols not designed for it cough cough HTTP
Finally I can have a shared home, with a shared environment.
A single readonly binary source (great for docker by the way.) also great for ensuring one version of scripts, without having to ansible everything.
"HAHAHA. YES! So much this."
part. Though I can't down vote so I wouldn't know.
S3 is high latency, HTTP semantics, no locking mechanism.
EFS is low latency, high throughput shared file system with standard posix interface. everything can write to standard files. Not many things can write using S3 effectively.
I can see people using this to deploy code by deploying onto the shared volume so that all their instances get an "instant update". Which is great when it works, but lord help you when EFS goes down and every app server you have is hung on a broken NFS connection.
The rational argument here is that if uptime is important to you, the solution is to utilize multiple regions, of which EFS is now in 4. You are using AWS because of the SLA and the assumption that when something goes wrong there are legions of technical folks trying to fix it as soon as possible.
With EFS, it's exposed as NFS, which hooks in much deeper. If it goes down, there isn't anything you can do to work around the problem, unless you start hacking your own file system kernel modules.
With soft mounts (does EFS support this?) you can exchange SIGKILL for religiously checking return values for all I/O syscalls, retrying partial writes and whatnot.
In all of those cases, anything which tries to access something on the NFS mount will block in the kernel (i.e. “kill -9“ won't work) and the mount cannot be unmounted normally.
I wrote https://github.com/acdha/mountstatus awhile back – if memory serves, 2004 or so – because we found that on Linux a lazy unmount would still work in this case and so you could have a process monitor the mount status (fork() a child to check the mount, alert if it doesn't get a response within a set interval) and a watchdog could respond to an alert by issuing a “umount -l” and remounting, which doesn't fix the blocked process but is less disruptive than rebooting and that new processes won't block because they tried to access that mount.
> Down time is still on you
If you can say "It's not our fault, it's Amazon's", there are plenty of boards and customers who are fine with that, in my experience.1) Actual customers: They want you to fix it. They don't care who's fault it is, they are dependent on you.
2) Internal customers: They want someone to blame. They don't really care that it's down, they just want to point the finger at someone when their report goes out late.
You bring up a good point on the EFS being a point of failure though. What use case do you think the EFS is good for? It seems like even worse for data storage since IO can easily become a bottleneck.
Out of band big data/data pipeline processing. Something where you want an easy way to sync your data but can handle extended downtimes.
Keep in mind (and this comes from hard experience in 'traditional' NFS web server architecture) - if you mount everything on an NFS volume, you ensure that
1) If something goes wrong on that NFS mount, everything goes wrong. (bad code deploy? All nodes are down!)
2) If you rely on an NFS mount to store everything (e.g. trust keystores for JVMs,etc.) your entire infrastructure is dependent on the I/O capabilities of that NFS mount.
3) No matter how clever you are (or how much you trust NFS clients/versions) you will deal with file locking if you are doing a fair amount of read/write from multiple nodes to a single NFS mount.
Short story - EFS will make some of the 'hard' things with distributed nodes possible, but don't make the easy things impossible to troubleshoot.
NFS is actually multi homed, so there is no real reason why it can't be HA/clustered, apart from the block store and the underlying file system.
It's not really, and you should write your software to account for SNS outages, etc.
That being said, this is presented as NFS. NFS has a nasty habit of freezing your system if it breaks. If SNS breaks, you get errors and timeouts in your logs, but can keep going. If NFS breaks, you pretty much just sit around waiting for it to come back.
This is the underlying tech that powered their Docker volume plugin as well [2]
I'm using in production for serving up small images to some web servers, and I'm currently in playground with docker stuff.
So far, so good. I'm impressed.
[1] https://azure.microsoft.com/en-us/documentation/articles/sto...
[2] https://azure.microsoft.com/en-us/blog/persistent-docker-vol...
edit: footnotes
Disclosure, work for AWS.
[1] https://azure.microsoft.com/en-us/documentation/articles/sto...
Since EFS was not available, I was trying to develop a FUSE filesystem that writes data to both S3 and local filesystem (EBS) for durability and read only from local filesystem for performance. Even though it's cheaper than EFS, it's slow since each file needs to be sent to S3 individually. Our requirements are not that much because the files are immutable, there are not many files (I heard that EFS doesn't work well with many small files.), no need for concurrent access. I think I will give it a shot since it also fits our use-case.
I ended up going with a shared volume, which works fine and performance is great, but I can only attach the volume to a single instance at a time which prevents me from running multiple instances. That said, it looks like there has been some performance tweaks to EFS so maybe it might be better this time.
On I/O front also EBS (gp2) and (st1) are several times cheaper.
So most of EFS advantages are in flexibility and sharing same file system across many instances. This may be great e.g. for sharing readonly binaries, but large scale data pipeline likely can be done much cheaper using different technologies.
I'll have to double check my working out, but you're only paying for what you use, not provision, which is one big things.
Also, if you have > 3 machines working from the same dataset (ie docker image/video/large binaryblob) there is an instant saving, without factoring in provisioned vs useable space with EBS
but, your original point, as a 1-1 replacement of a properly sized EBS mount, is correct, its more expensive.
"large binaryblob" is literally "large binary binary large object"
Isn't I/O for EFS free?
It's really a top notch product; couldn't be happier.
But I also can't afford S3's bandwidth pricing.. I could, however afford to run an S3-API compatible service myself with bandwidth pricing I could afford.
But then you wouldn't let me use it because then I'd be "enterprise".
Depending on how they wrote their code, if the Lambda could seek to the right part of the file then this would work, but if it has to load the entire file into RAM anyway, then it wouldn't gain them much except perhaps a slightly easier way to get the file into RAM.
From a caching perspective, you wouldn't seem to gain much in lamda. You can cache in memory/tmpfs until your container eventually dies, and then you likely aren't on the same machine so there's no EFS level caching to take advantage of.
EFS/NFS makes a lot of sense I think when you have a lot of random/unpredictable access across large numbers of files. You could do the same thing with FUSE/S3 but across lots of smaller files/accesses I would think the overhead adds up faster.
We have N upload endpoints (as part of an autoscaling group), it's a bit of a pain to reaggregate the data since each server uploads its own .tgz to S3. I'm very happy to have EFS now! It makes scaling out our upload endpoints much simpler, as they can just dump their data into a common directory (which can still get archived and persisted to S3 later).
Immutable files => Object storage => S3
Mutable files => File storage with POSIX semantics => EFS/NFS
Thinking of small file size, EFS would probably work a bit faster (since you'd have some HTTP overhead) but I doubt that it would be significant, unless users upload hundreds of files at once.
EFS advantage in this case would be that you'd no longer need a S3 library to do it, since it's POSIX compliant. Given that most languages have simple, solid libraries for interacting with it, it's not a huge difference but still :)
All in all, I'd go with S3 unless there are some special requirements that are hard to satisfy using current S3 features.
OFS feels nearly like an attached disk, and uses S3. You only pay per mount. Their support has been awesome and very personal.
It provides a fuse-based full POSIX filesystem backed by S3 and can be mounted from multiple places in multiple regions. Some basic things like snapshotting they said are coming soon too.
I gave up on EFS and even though it's finally out of preview, I think I still am going to prefer OFS based on what I've seen so far...
That would be an interesting feature.
I would be very wary of running a database from that.
There are workarounds for the stateless mess of NFS in modern linuxes, but if you accidentally access a NFS-mounted database from multiple AWS instances, you might get into big trouble, especially when the machines in question run different OSes (e.g. Windows and Linux) or OS versions.
It'll never be as fast as local disk but the flexibility it provides is very nice.
Our NetApp can saturate 10GBe without even trying.
Would love to see this in comparison to a 10Gbps (or 1Gpbs) attached spinning rust or SSD NSF drive. That would really help me understand what tradeoffs are happening here.
Even though we don't use AWS much due to lack of IPv6 support, we've wanted to have a giant bottomless FS for various purposes. It'd be easy to set that up in here and export it over a virtual network.
It's an interesting start to a project, but the public features feel very much lacking.
The neat trick would be if NFS 4.X grew support for communicating interesting operations like snapshotting.