Unix shells and the current directory
utcc.utoronto.ca
utcc.utoronto.ca
Related reading: Lexical File Names in Plan 9 or Getting Dot-Dot Right (https://9p.io/sys/doc/lexnames.html)
The current directory is managed with SetCurrentDirectory/GetCurrentDirectory, however the cmd.exe command-line shell also stores the current directory for each drive in an environment variable like "=C:", and the CRT and shell hides all environment variables that start with a "=".
It gets mightily confused if these two concepts of current directory ever diverge.
I don’t mind cmd.exe and it launches instantly (same reason I frequently use notepad.exe for quick edits). That latter quality is very hard to find :)
Edit: but if you meant for scripting, yeah, batch files are terrible.
Which has the odd result that '..' behaves differently between shell builtins and normal commands. `cd ..; ls` uses the text version, but `ls ..` uses the filesystem version. `cat < ../x` uses the text version, but `cat ../x` uses the filesystem version.
I like the text behavior in theory, but this inconsistency is weird enough that I question the benefit of having the text behavior at all.
The shell already seems to track it, so presumably, that logic could have been part of the standard library, and get tracked from user-mode.
If the kernel has to track the current directory (e.g. for performance reasons, to make accessing files relative to a particular directory more efficient), wouldn’t just remembering the device ID and inode be easier for the kernel?
Alternatively, there could be kernel calls taking (device, inode) pairs, and the kernel could be completely ignorant of the ‘current directory’ concept.
That can work; except for naming them ‘directory ID’ instead of ‘inode’, that’s what the first Mac OS hierarchical file system did; paths were second-class citizens here.
But you can't punt this entirely to the shell - the shell has to look at the kernel's idea of the current directory name at startup; all the in-shell tracking can only be done for subsequent changes.
One major caveat is that the kernel's API stupidly relies on a single-step global record and is limited to one page (usually 4096 bytes), rather than reconstructing it component-by-component. So if you change into a deeply nested folder, `getcwd` falls back to the `open(".."); readdir` loops. Of course if `PWD` is set correctly it can be used, but if it's not canonical you might have to do the nasty version later.
A more subtle caveat is all the possible end cases:
* you reach the current mount namespace's sense of `/`
* you reach the current mount namespace's sense of `//`, if your environment supports such a thing (note that `readdir` likely fails at the last level though!).
* you reach some other sense of `/` (e.g. from an FD kept open across `chdir`, or an FD passed across a Unix socket from a different mount namespace)
* the directory was not found in the readdir loop (a directory moved due to a race condition, or special filesystems that aren't fully enumerable - this includes /proc/ if you use a thread ID directly - this has a different inode than the main PID which it mostly acts like!).
The kernel knowing the name for the current directory is not specific to current directories; it is part of a general system of caching the name mappings for directory entries ('dnodes' in Linux, a 'name cache' in FreeBSD). Unix kernels added these caches because Unix programs spend a lot of time looking up names, making the operation worth optimizing in general. Once you have a general name cache, you might as well pin the entries for actively used entities like current directories and open files so that they don't get expired out of the cache and you always know (some) name for them.
(One useful complexity of name caches is that you can cache negative entries, ie that a given name is not present in a directory. In the modern Unix shared library environment where shared libraries may be probed for in a whole collection of directories every time a program starts up, I suspect this saves a nice chunk of kernel CPU time.)
* That $PATH is just an ordinary env var, that many programs use by convention
* That CWD isn't, and is in fact a first-class kernel concept. I had assumed that it was just a conventional envvar that stdlibs prepended before passing absolute paths to syscalls
I'm sure there's good reasons why the other way wouldn't work, it just amused me that I'd got it wrong in both ways
However, if you want things like DTrace, eBPF, or even just reading the /proc/PID/cwd symlink to be useful, it helps to cache the actual path in the kernel. A DTrace/eBPF script will not be able to loop to chase ..s, much less will it be able to do the I/O needed to work out the cwd.
The same applies to the names of the files that each FD refer to.
Caching these things is just for observability.
This must be tracked by kernel, because not all syscalls go through libc, you can issue the open syscall directly from a process.
There might be other reasons, but I'd bet it's the main one.
ls -l /proc/`pidof foo`/fd
to see what files a taciturn process is working onI mean, the article talks about the "what", but it makes me wonder about the "why".
Unless one only wants the current directory name at the start of a script why not just use the builtin pwd command, $(pwd). Or getcwd() if it's "other code".
`$PWD` is always set accurately at startup (for both interactive and non-interactive shells) in bash, dash, zsh, ksh93, mksh, and busybox ash.
So I'm really not sure where this assumption can be violated.