Reading NFS at 25GB/s using FIO and libnfs
taras.glek.net
taras.glek.net
I learned nothing about FIO, libnfs, NFS, or your patch.
No feature comparison with other efforts. No benchmark comparison with other efforts.
I wasted my time here.
But it’s good to know there are ways to get much more perf out of it.
Using vectored IO and spreading across multiple connections greatly improved throughout. However, metadata operations cannot be parallelized easily without application side changes.
In more modern kernels, NFS supports ‘nconnect’ mount option to open multiple network connections for a single mount. I wonder if the approach of using libnfs for multiple connections is even required.
I've tried this with a read-only NFS server, across an AWS multi-gigabit connection, and I still found that I couldn't get anywhere near this level of performance for the workload of "make -j$(nproc)" in a Linux kernel tree. As a quick baseline, some numbers from the last time I tested this, with a c5.12xlarge (48 CPU) client and server: a defconfig local build was 40s, and a defconfig build with a Linux kernel tree in read-only NFS (with a tmpfs overlay on top for writability) was 6m55s. That's a 10x slowdown. System stats during the build showed 4-5MBps net recv and net send, and 1-4MBps disk write.
Is there some well-known method to getting reasonable performance out of off-the-shelf NFS servers and clients?
Performance turning NFS is difficult mostly because the information on how to do it isn't readily available. The two things that most people run into:
- 'Close to open' cache consistency. In practical terms this means that open() is a round trip to the server to validate any data that might be cached already (unless you use delegations), write() goes into the page cache (as a writeback cache), and close() flushes all dirty data. Building a kernel tree reading and creating tons of small files, each of which requires two serial round trips over the network. Compare that to a local fs where neither open(O_CREAT), write() or close() actually go to disk and therefore run at memory speed (unless you use things like O_DIRECT or fsync/fdatasync()).
- Per-TCP flow throughput limitations. On the AWS network the per-flow limit is 5 Gbit in general and 10 Gbit within a placement groups. To work around this, people use the 'nconnect' mount option. (which does not work currently with EFS). Local networks might have different limitations, but single TCP streams will typically always have some bw limit lower than the physical network bandwidth. I believe that this (very cool!) fio plugin works around this by using multiple connections.
The actual data write latency of NFS servers isn't terribly different from local file systems.
Today, the best way to get the most performance out of NFS is to either use large files and/or keep files open, or use high concurrency. By default, the 4.1 client will issue up to 64 concurrent requests, which can be increased by increasing the 'max slots' NFS kernel module parameter. In your example of a kernel build, you could -j much higher than the number of CPUs because the compile jobs will be IO bound on reading input and writing output. This will amortize the round trips over more threads, and in theory (barring any other bottlenecks) reduce your build times.
> - 'Close to open' cache consistency. In practical terms this means that open() is a round trip to the server to validate any data that might be cached already (unless you use delegations), write() goes into the page cache (as a writeback cache), and close() flushes all dirty data. Building a kernel tree reading and creating tons of small files, each of which requires two serial round trips over the network. Compare that to a local fs where neither open(O_CREAT), write() or close() actually go to disk and therefore run at memory speed (unless you use things like O_DIRECT or fsync/fdatasync()).
That definitely sounds like a concern for writable NFS filesystems, but I was benchmarking reads to a read-only NFS mount.
Related: Is there some option I can pass to make it clear that the data on the server will never change and thus no possible write-to-read or close-to-open consistency issues can arise?
> - Per-TCP flow throughput limitations. On the AWS network the per-flow limit is 5 Gbit in general and 10 Gbit within a placement groups. To work around this, people use the 'nconnect' mount option. (which does not work currently with EFS). Local networks might have different limitations, but single TCP streams will typically always have some bw limit lower than the physical network bandwidth. I believe that this (very cool!) fio plugin works around this by using multiple connections.
Interesting! I've never seen the per-flow limit mentioned before. Is that documented somewhere?
I'd be concerned about that if I were getting anywhere close to that limit, but I was experiencing 4-5MBps network throughput. It seemed like individual file operations (like stat) were taking an excessive amount of time.
> In your example of a kernel build, you could -j much higher than the number of CPUs because the compile jobs will be IO bound on reading input and writing output.
I'm writing output to a local tmpfs (via overlayfs), not to NFS. And I'd love to tune the NFS setup to the point that reads (and stats) from NFS aren't causing a 10x slowdown.
As far as I know the NFS client does not support such a mount option today. I should have mentioned this, but there /is/ a way to eliminate the 'close to open' cache check for repeated open() operations, which is to use NFS delegations. NFS read delegations are supported by both nfsd and the NFS client. They are not perfect, as they are best effort, but can typically keep the core data set of your workload fully local. This would not work for your first build but would work for the second.
> Interesting! I've never seen the per-flow limit mentioned before. Is that documented somewhere?
https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ec2-inst...
I really only know the name, nothing on how to configure it.
https://developer.nvidia.com/blog/doubling-network-file-syst...
Otherwise, yeah, you incur network latency on every file, plus, as you say, symlink "pings" if those are present. So its "dir walk"+"number of files"*"ping"+"symlinks"*"ping", which adds up.
Batching high-latency ops is one of the only cases where I like async.
I wrote a tool that issues raw READDIRPLUS requests to list a directory:
I haven't yet instrumented memory bandwidth on my amd machines, but it feels like I'm at the limit.
Metadata is more complicated to model, need realistic directory structures etc.
https://www.spec.org/sfs2014/ (2020 version) is a good test of a metadata-heavy workloads, but isn't open source like fio :(.
I wrote a suite of tools that does all this and dumps the NFS transactions as JSON, maybe it can be useful:
Oh the horror.
What? I presume "multiple NFS connections per host." was meant. Not sure what those NFS connections are supposed to be though. NFSv3 is (stateless) request/response protocol on top of TCP/IP connections. NFSv4 introduced sessions, but here NFSv3 was used, wasn't it?