Presumably you mean "Transparent remoting isn't a thing"?
Paper authors were at Sun, so had experience with NFS (the design of which assumed that transparent remoting was a thing).
Presumably you mean "Transparent remoting isn't a thing"?
Paper authors were at Sun, so had experience with NFS (the design of which assumed that transparent remoting was a thing).
NFS locking up is the most common answer to the typical interview question “what could cause a load average in the hundreds on a machine with zero CPU usage?”: because there’s a stuck NFS server and the pending IO operations put the processes in the run queue.
I’m glad there’s a paper that articulates my thoughts about this subject better than I can.
These things have nothing to do with distributed systems, they can happen on any machine. That's what the "Not ready reading drive A. Abort, Retry, Fail?_" prompt is for.
And write is write-once, from the start of the file.
Blocking incurs a lot of problems, because the kernel ends up waiting a ridiculously long default timeout for the server to respond, and the calling thread is blocked the whole time. Making a remote-access API asynchronous (callback-based), or at least having O_NONBLOCK being the only allowed option, would help a lot here.
But really, the other problem is that filesystems have too much baggage around assumptions of speed/locality/reliability that just blindly treating a remote server as if it’s a local filesystem is just a bad idea in the first place. Too much client software is written with the assumption that filesystems are fast (ie. you don’t need to cache it, stat()’s are cheap, etc) that it’s just better to not conflate the two.
What's interesting about GFS is that it's a purely RPC protocol with no kernel and the access library is linked into the application. That allowed a lot of freedom for both the server developers and the clients. It also forced the client developers to explicitly acknowledge they were working with a remote filesystem.
(BTW, I work for an enterprise that uses NFS; we have 20PB in there and lots of HPC jobs and interactive users pounding on it. But we spent $$$ for a very good NFS server)
A typical issue was jobs which ran a bash script which calls perl a bunch. Perl is on NFS. Oh and perl isn’t just on NFS, but the path to it is a symlink, which resolved to another symlink, which resolved to another symlink, which finally resolved to the perl binary. So that’s a half-dozen stat() calls per each call to `perl`. (None of these are cached, because metadata wasn’t cached, only the file contents were…) Said bash script would use ` | perl -pe` instead of grep, etc…
Oh and typical perl programs called by these scripts had a bunch of modules they needed. Guess where the modules are? NFS! Not just NFS, but behind their own sets of symlinks (gotta be able to version your perl code! Why not use symlinks!)
Some of these symlinks would trigger an automount. So now your bash script’s call to `perl -pe ‘s/foo/bar’` had to go mount something from another tools server. Yay.
Needless to say we paid the storage vendor a lot for support.
We put tons of artifacts and modules on NFS but our filers are designed to scale. Once, I even had to evaluate a product (Gear 6) which was a read-only cache of another NFS server, just to offload the heavy reads caused by cluster scripts.
With all that said, the best cluster I ever ran was diskless- all the worker nodes mounted / and /home from a master node, booting straight into a root NFS filesystem.
IBM's automountd was single threaded.
Someone's code (specifically, a systems programming class assignment) would lock the FS with their home directory. After a period of inactivity the automounter would try to unmount the directory, blocking. The machine is now nearly useless.
Eventually we got a patch (custom for our enterprise) that limited the outage to anybody who had a current working directory that was the parent of the rm -rf'd dir. And we dropped NetApp shortly after.
But the real point is that the filesystem, as an abstraction, has too much baggage from clients that make local-first assumptions, like that stat()’s are cheap (let’s use lots of symlinks! Hey, why is the NFS server CPU usage so high from stat() calls?), or that the filesystem can be used for easy cache (which on NFS is kind of the opposite of what you want), etc… silently and transparently making your filesystem an NFS mount is a bad idea unless the software using that mount is aware that the other end is a (potentially slow) NFS server and not local disk. The two really shouldn’t be conflated.
Under Linux, a high load roughly tells you that the system is bottlenecked on something, where that something isn't necessarily CPU cycles. If your system is bottlenecked on pending NFS requests, the load will be high, and that's not the result of any kernel shortcomings (unless one blames blocking file I/O on the difficulty of correct use of exposed async/non-blocking file I/O syscalls).
The high load average is a matter of accounting[0]. Under Linux, tasks (threads) in I/O wait are considered runnable, although they're not scheduled for CPU time. Other OSes (e.g. most Unix descendants, probably WinNT) make other accounting decisions, but either way, tasks blocked on I/O don't affect CPU-bound tasks. The Linux manner of load accounting isn't what I would have implemented at first, but if you're going to collapse system load into a single metric, then it seems best to include blocked tasks. If you've got high load, then it's time to dig deeper into more nuanced metrics anyway, so in retrospect, I agree with Linux's more holistic load accounting.
[0] https://en.wikipedia.org/wiki/Load_(computing)#Unix-style_lo...
https://news.ycombinator.com/item?id=20007875
>There used to be a bug in the GatorBox Mac Localtalk-to-Ethernet NFS bridge that could somehow trick Unix into putting slashes into file names via NFS, which appeared to work fine, but then down the line Unix "restore" would totally shit itself.
And here's another cool party trick you can baffle people with (not related to NFS, just an obscure MacOS quirk):
>I just tried to create a file name on the Mac in Finder with a slash in it, and it actually let me! But Emacs dired says it actually ended up with a ":" in it. So then I tried to create a file name with a colon in it, and Finder said: "Try using a name with fewer characters or with no punctuation marks." Must be backwards compatibility for all those old Mac files with slashes in their name. Go figure!
https://news.ycombinator.com/item?id=31820504
>NFS originally stood for "No File Security".
>The NFS protocol wasn't just stateless, but also securityless!
>Stewart, remember the open secret that almost everybody at Sun knew about, in which you could tftp a host's /etc/exports (because tftp was set up by default in a way that left it wide open to anyone from anywhere reading files in /etc) to learn the name of all the servers a host allowed to mount its file system, and then in a root shell simply go "hostname foo ; mount remote:/dir /mnt ; hostname `hostname`" to temporarily change the CLIENT's hostname to the name of a host that the SERVER allowed to mount the directory, then mount it (claiming to be an allowed client), then switch it back?
>That's right, the server didn't bother checking the client's IP address against the host name it claimed to be in the NFS mountd request. That's right: the protocol itself let the client tell the server what its host name was, and the server implementation didn't check that against the client's ip address. Nice professional protocol design and implementation, huh?
>Yes, that actually worked, because the NFS protocol laughably trusted the CLIENT to identify its host name for security purposes. That level of "trust" was built into the original NFS protocol and implementation from day one, by the geniuses at Sun who originally designed it. The network is the computer is insecure, indeed.
>And most engineers at Sun knew that (and many often took advantage of it). NFS security was a running joke, thus the moniker "No File Security". But Sun proudly shipped it to customers anyway, configured with terribly insecure defaults that let anybody on the internet mount your file system. (That "feature" was undocumented, of course.)
https://news.ycombinator.com/item?id=31820891
>Stewart, I think "sucks" is a pretty fair description of a protocol that actually trusted the client to tell the server what its host name is, before the server checked that the host name appears in /etc/exports, without verifying the client's ip address. On a system that makes /etc/exports easily publicly readable via tftp by default.
https://news.ycombinator.com/item?id=31822138
>Speaking of YP (which I always thought sounded like a brand of moist baby poop towelettes), BSD, wildcard groups, SunRPC, and Sun's ingenuous networking and security and remote procedure call infrastructure, who remembers Jordan Hubbard's infamous rwall incident on March 31, 1987?
https://news.ycombinator.com/item?id=31821646
>Another reason that NFS sucks: Anyone remember the Gator Box? It enabled you to trick NFS into putting slashes into the names of files and directories, which seemed to work at the time, but came back to totally fuck you later when you tried to restore a dump of your file system.
When I read this paper (it's been a long time) it seemed to overlook this angle on the problem.
> Rather than using those resources in attempts to paper over the differences between the two kinds of computing, resources can be directed at improving the performance and reliability of each.
> ... it is a mistake to attempt to construct a system that is “objects all the way down” if one understands the goal as a distributed system constructed of the same kind of objects all the way down.
(E.g. in E the objects are the same kind; it's the references to objects that have two kinds.)