That's exactly what happens. The server ACKs data until it fills its write buffer, and then stalls unresponsive until the entire buffer is flushed to disk. If it takes longer to flush the buffer to disk than the client's timeout, it gives up.
I have personally watched this happen via wireshark where the server doesn't ACK for more than 10 minutes.
You can confirm by watching /proc/meminfo and watching the Dirty and Writeback numbers.
Changing up the vm.dirty* settings can help as described here:
https://lonesysadmin.net/2013/12/22/better-linux-disk-cachin...
About 4-5 years ago, i was working on a project, and part of that was copying big amounts of data to a system via nfs. At 30 minutes exactly, nfs would croak, transfer fails.
I think this buffer fill and empty flow was fucking killing it. Its a shame i dont work there anymore, id definitely wanna try tweaking these settings and see if i could solve it
And it's the vm.dirty* settings to change to fix it as described here: https://lonesysadmin.net/2013/12/22/better-linux-disk-cachin...