Not knowing the /proc file system
admccartney.mur.at
admccartney.mur.at
The sample code won't work reliably, because it assumes that the "files" won't change while being read. If you read /proc, you must "read" each file with one unbuffered kernel read to be free of race conditions. See [1].
The Linux kernel has no standard mechanism for delivering a variable-sized result from a system call. I/O comes closest to that, so this was hammered into the file system API. The dents show. /proc does not have standard file semantics.
[1] https://stackoverflow.com/questions/5713451/is-it-safe-to-pa...
I think the biggest issue with /proc at the moment is that it leaks in core Linux components, and getting rid of it is becoming impossible. Some years ago I worked on adding `$ORIGIN` rpath support in OpenBSD and looked at how Linux did it. Mind you, the dynamic linker (ldlinux.so) grabs the running executable path from /proc...
One thing though, you focus a lot on syscalls, but I think the challenge in recent years has been to find new ways to have kernel <-> userspace interactions _outside_ of syscall, which are cumbersome to use, rigid in structure, and practically speaking, there are only so many entries you can store in an IDT.
The work on netlink is going in the right direction IMHO. It allows userland to communicate bidirectionally with the kernel, register on specific events, etc; while using a familiar tooling with bsd sockets that is convenient from both shell and programming languages.
The number of syscalls isn't limited by this, they all go through the same interrupt and are dispatched to a specific handler function based on the value of a register.
For example this is how it works on arm64: https://elixir.bootlin.com/linux/v6.5/source/arch/arm64/kern...
And on x86: https://elixir.bootlin.com/linux/v6.5/source/arch/x86/entry/...
There's always ioctls. You could have something a lot like /proc and /sys, with a directory hierarchy modelling various interesting things, but rather than the files having textual contents, you interact with them via ioctls. ioctls are identified by a device-specific integer, so you have masses of namespace (about 2*32 operations per device type, and you can have a lot of device types). They take or receive a block of memory in one go, so there is no parsing or formatting or worrying about inconsistent reads. I think this is even simpler than protocol-esque interfaces like netlink.
It is still not perfect for all use cases though, especially because it forces you to operate in a _polling_ fashion.
Programs like top/iotop/htop are a good example (also, who is not guilty of abusing one liners a-la `watch -n1 cat /proc/meminfo`?). Instead of polling /proc or ioctls, you would really like to register on events, and get notifications pushed.
Netlink is well suited for these kinds of scenarios, though mostly to get notified of network interface changes last I checked.
> Netlink is often described as an ioctl() replacement. It aims to replace fixed-format C structures as supplied to ioctl() with a format which allows an easy way to add or extended the arguments.
Not good. Not so long ago I ran into a problem with mdadm which instead of reading sysfs used ioctls which simply didn't carry enough information for it to function properly. So, it still has bugs because of that, but it being written in C, reading from a file and parsing stuff from a string seems to cause a lot of melancholy in its maintainers, so the bug has been on slow burner for years now.
Today there is another option for doing the kernel-userspace communication: using an io_uring like interface. Yes, it's a syscall, but it's one with enough extensibility and performance characteristics. It's extensible because it's close to an IPC with the kernel (and it could be changed to be a new IPC mechanism for process-to-process communication if it isn't already). And it's performant because it's based in shared memory and is assyncronous by default.
But, more generally, if we want general-purpose communication, why not press further and have relational database-like interface, with transactional semantics, triggers, fine-grained ownership?..
`getauxval(AT_EXECFN)` is "better" but there are weird edge cases with `fexecve` where you only get `AT_EXECFD`, even ignoring the inevitable TOCTOU.
I ended up pushing the executable filename through the aux vector to the dynamic linker, which is simar to what you mention AFAIU.
Never managed to get that merged upstream though. Glad to know there's now a _standard_ way.
>If you read /proc, you must "read" each file with one unbuffered kernel read to be free of race conditions.
you make this sound like it's some arcane and tricky incantation; it's totally normal, like saying "to walk, put one foot in front of the other". It's unix file I/O 101, bufferless file operations are always atomic, it's a feature. There is absolutely no problem making atomic reads from files, it's a massive fucking feature.
The point is not that this is the best conceivable way to get this information, but that by making this information available this way by default, you can write and leverage all sorts of shell scripts while you are noodling at your keyboard. You don't very often need to turn your keyboard noodling into a 6-sigma 5-nines server, but if you know how to rip through this type of keyboard noodling, I'd trust you a lot more when it comes to uptime than all the people who show up at meetings ranting about 6-sigma and 5-nines while they cite Dave Cutler.
btw linux the /proc text file tree has already been supplanted by a binary "restatement" of the same and more data. I think this is more in keeping with the standard unix/solaris /proc approach, but it's not something I've had a need to look into much.
I got an SDF account recently, and was surprised to find it in NetBSD. OpenBSD has great resistance to it.
...looking at the wiki, many more kernels implement /proc:
"Many Unix-like operating systems support the proc filesystem, including Solaris, IRIX, Tru64 UNIX, BSD, Linux, IBM AIX, QNX, and Plan 9 from Bell Labs."
A case in point: FreeBSD's /proc is very different to Linux's /proc, and most of what one would go to /proc on Linux for is obtained via sysctl() on FreeBSD, with a lot less in the way of machine readable → human readable → machine readable busywork formatting and re-parsing involved.
https://www.usenix.org/sites/default/files/usenix_winter91_f...
Also there's no guarantee whatsoever that /proc/123 is going to refer to a specific process, or to remain existing while you get all the data you need. The whole API is full of race conditions.
does using openat solve that or is there some kind of directory inode reuse occurring?
In NetBSD you can create XML serialized system calls for this. :-)
FreeBSD is the most general purpose of the 3 and the most popular. It is also the only BSD out of the big 3 that doesn't utilize a global kernel lock, allowing for modern symmetric multiprocessing similar to Linux.
NetBSD is aimed at being extremely portable. It's sort of like the "Can it run doom" of the OS world. Just take a look at their list of ports: https://wiki.netbsd.org/ports/
OpenBSD is aimed at being secure. Exactly how realized this goal is is somewhat controversial. But regardless of that, security is the stated highest priority of the development team.
There's also DragonflyBSD, which was forked by Matt Dillon from FreeBSD following some personal and technical disagreements. It's since diverged pretty heavily from the rest of the BSD family. Given its very low market share in this category of already niche operating systems, it seems more like a pet project of Matt Dillon's, though I'm sure it has serious users.
Not sure if it’s related to this case or not but it was once much more popular even for sometimes human written files.
Maybe we can expand the userspace networking interfaces, and make a fully featured messaging system. Some day Linux may even reach feature parity with L4!
AFAIK, as that StackOverflow page also says, that’s not guaranteed to be possible. https://man7.org/linux/man-pages/man2/read.2.html:
“RETURN VALUE
On success, the number of bytes read is returned (zero indicates end of file), and the file position is advanced by this number.
It is not an error if this number is smaller than the number of bytes requested; this may happen for example because fewer bytes are actually available right now (maybe because we were close to end-of-file, or because we are reading from a pipe, or from a terminal), or because read() was interrupted by a signal”
I think you can avoid the “because read() was interrupted by a signal” part by not installing signal handlers, but even if that’s an option for you, that list isn’t exhaustive.
Your best bet is to pass the maximum supported number for count to the call (SSIZE_MAX in POSIX, 0x7ffff000 on Linux, and hope that the call returns all bytes.
Luckily, I think that will be fine most of the time, but that doesn’t mean /proc is a good idea.
And then just hope your read buffer doesn’t overrun? Because 2GB is a lot of buffer to allocate for a read…
For example, `timerfd_create`'s returned FD guarantees that `read` will always return exactly 8.
Naturally, only source reading will answer whether the proc file you're interested in is vulnerable to this or not.
At the moment. To be sure, you also have to keep track of changes to that source code to look for regressions, and have a mechanism to update all your deployed code if there is a regression there.
It could require multiple attempts at reading, and still doesn’t ensure progress, but I think /proc would be a lot more reliable if file contents included a good checksum.
I'm not disputing there are issues but that's a very misleading description of the implementation.
It's not an error, per se. (The ioctl is literally erroring, but that's an expected possibility for the calling code, and it handles that.)
The reason there's an ioctl is documented in the docs for `open`:
> buffering is an optional integer used to set the buffering policy. Pass 0 to switch buffering off (only allowed in binary mode), 1 to select line buffering (only usable in text mode), and an integer > 1 to indicate the size in bytes of a fixed-size chunk buffer.
You're not passing the `buffering` arg, so the subsequent text applies:
> When no buffering argument is given, the default buffering policy works as follows:
> * Binary files are buffered in fixed-size chunks; […]
(That doesn't apply, as you're not opening the file in binary mode, so it's the next bullet that applies)
> * Interactive” text files (files for which isatty() returns True) use line buffering. Other text files use the policy described above for binary files.
That ioctl is the underlying syscall that isatty() is calling. It's determining if the opened file is a TTY, or not. The file isn't a TTY, so the ioctl returns an error, but to our code that just means that "no, that isn't a TTY". (And thus, your opened file will automatically end up buffered. The flow here is a good default, for each of the cases it is sussing out.)
root@74c03a282fbe:/# ed
a
ls /proc/[0-9]/status | xargs -n 1 cat | awk '/^Name:/ { name = $2 } /^Pid:/ { pid = $2 } END { print "cmd: " name ", pid: " pid }'
.
w prc.sh
132
q
root@74c03a282fbe:/# chmod +x ./prc.sh
root@74c03a282fbe:/# hyperfine --warmup=100 "./prc.sh"
Benchmark 1: ./prc.sh
Time (mean ± σ): 2.2 ms ± 0.3 ms [User: 1.3 ms, System: 2.7 ms]
Range (min … max): 1.8 ms … 5.0 ms 880 runs
Warning: Command took less than 5 ms to complete. Results might be inaccurate.
Warning: Statistical outliers were detected. Consider re-running this benchmark on a quiet PC without any interferences from other programs. It might help to use the '--warmup' or '--prepare' options. $ echo /proc/[0-9]/status | wc
1 7 105
$ echo /proc/[1-9]*/status | wc
1 483 8778
It also will sporadically print error messages due to all the race conditions. Here's my stab at it: $ cat dumbps
#!/bin/sh
case $# in
1)
;;
*)
echo "Usage: dumbps user" >&2
exit 2
;;
esac
2>&- find /proc -mindepth 2 -maxdepth 2 -type f -name status -user "$1" | awk '{
while ((getline li < $0) > 0) {
if (li ~ /^Name:/) {
split($0, fn, "/")
print fn[3], substr(li, 6)
break
}
}
close($0)
}'
$ ./dumbps "$USER" | grep -w 'bash$'
1616678 bash
$ ./dumbps 0 | grep -w 'systemd$'
1 systemd
$
It's pretty fast too: $ ls -ld /proc/[1-9]*/status | awk '$3 == "root"' | wc
445 4005 29367
$ time ./dumbps >/dev/null
Usage: dumdps user
real 0m0.001s
user 0m0.000s
sys 0m0.001s
$Perfectly possible, even likely, for a process to end between "find" finding it and awk opening it... especially given the pipe buffer between the two.
Personly I'd rather an "expect errors" approach.. break it out so that you can easily treat a failed open as a "continue next" scenario and hide just those errors
while ((getline li < $0) > 0) hides races in awk
in practice Name: is the first line so cmd is not truncated
Of course I got my test wrong though :D
$ time ./dumbps 0 >/dev/null
real 0m0.030s
user 0m0.012s
sys 0m0.022sIt's particularly helpful in larger infrastructures where tool variability means differences in available commands, their output, and cli options. I'm sure /proc iteration has its own issues of variability across large infrastructres, but I haven't seen it. It's a fairly consistent API. Or at least it was, since I haven't touched a large infrastructure in some time.
When I got tired of `lsof` not being installed on hosts (or when its `-i` param isn't available) I ended up writing a script [1] that just iterates through /proc over ssh and grabs all inet sockets, environment variables, command line, etc from a set of hosts. Results in a null-delimited output that can then be fed into something like grafana to create network maps. Biggest problem with it is the use of pipes means all cores go to 100% for the few seconds it takes to run.
Do you mean like using `fp` for a file pointer or `fd` for a file descriptor? That is idiomatic, and I would consider calling them `filePointer` or `fileDescriptor` to be an obvious smell that the developer doesn't know what they're doing.
I'll do something like:
int epollFd = epoll_create1(0);
So yes, I don't think I have ever typed out "fileDescriptor" but I do label these things to make it more legible!OP has a point. We could do with more discipline when it comes to naming conventions in C. We're not using punch cards any longer.
Somewhat related, there's another comment about the 80 char max... clang-format keeps to this. I'm kind of OK with this particular remnant of punch cards because I can get four editor windows open side by side on my ultra wide monitor.
The nginx codebase is pretty freaking fantastic and I've learned a lot from just randomly browsing through the source, but I mean, come on:
https://github.com/nginx/nginx/blob/master/src/core/ngx_arra...
if ((u_char *) a->elts + a->size * a->nalloc == p->d.last) {
What's p? Oh, ngx_pool_t. Object pool? Memory pool? What's a again? elts? Is that elements?If that read:
if ((u_char *) array->elements + array->size * array->nalloc == memoryPool->d.last) {
It's time for the old school conventions to change a little bit. C ain't going anywhere. Let's embrace a world where descriptive variable names are not subject to cost/benefit analysis like they were in 1977.Even after that you had compilers with limits like 6 character names for a symbol.
Those constraints went away, but names made to be comfortable for the users of teletypes and ancient compilers stuck around, and people made more of them because it fit the preexisting theme.
And now it's just plain inertia, where it keeps on going because that's what it looks like in the books, so students imitate it.
That's more than inertia, that seems to cross the threshold for ceremony. When was the last time that actual printers/teletypes were used, the 60s?
The inwrtia part of it is people imitating what they already see on a project (cause why would you disrupt something as a newcomer on a project) and line length rules which I've seen come up from time to time along the years (not sure if a 120 character line length is enforced nowadays within the kernel)
After you've grokked it the first time, e.g. "dgemm" is much more convenient than "double precision general matrix multiplication". This is akin to speaking aloud "DC" versus "The District of Columbia". Uses far outweigh learning.
Consider even Python calls it a "dict" not a "dictionary" because the latter is a mouthful. Though what was wrong with "map" I often wonder.
A conflict with the map() function maybe?
No. I like these names as it's just easier to read. fp or fd is short and to the point, and "filePointer" or variants thereof add noting except more characters to read, more "wall of text"-y code that's harder to scan, etc.
And I don't really want to have a discussion about it as such; whatever your preference is that fine. I'll be happy to adjust to whatever works well within a team. But I wish people would stop spreading nonsense like this to invalidate other people's preferences.
I definitely make a point of using only single-letter variables in most of my C and python programs.
It is a very common usage in scientific computing. In math, all variables are single letters. Always. If a variable has more than one letter, you read it as the product of several variables, one for each constituent letter. When you are translating a formula that reads "y=Ax", you want to write something like
y = A * x
Writing this formula as output = operator * input
or, god forbid, as something like output = linalg.dot(operator, input)
is completely ridiculous to any mathematician.Mathematics itself used to be like that in ancient times. But after centuries of distillation, we arrived to the modern efficient notation. Some "programmers" want us to go back to the ancient ways, writing simple formulas as full-sized English sentences. But they will take single-letter variable names from our cold, dead hands!
Of course, the first appearance of each single-letter variable must be accompanied by a comment describing what it is. But after this comment, you can use that letter as many times as you want. Encoding that information in the variable name itself would be disturbingly redundant if you use the variable more than once (which will be always the case).
However that particular use-case has become dramatically less significant.
When I code, I prefer comments explaining the general idea of what’s happening and why in a given folder/file/section (not explaining each line) combined with terse code.
Terse notation has nothing to do with manual handwriting. Modern math books and articles are still written by computer using a very terse symbolic notation, which has been developed during the last six centuries. Originally, the symbols +, = were shorthand abbreviations of Latin words.
I guess computer scientists want to re-invent everything from scratch. How long will they need to evolve from "LinearAlgebra.matrixVectorProduct(,)" to the empty string? I hope it's less than six centuries!
Mathematicians seem to miss the difference between math and code quite often. The former provides the solution to a well-understood problem in a straightforward manner. The latter transports an abstract concept, a plan, a state of thought to the reader. A neat side effect is making computers go beep. In that context, being as clear as possible is really important.
Example:
https://learn.microsoft.com/en-us/windows-hardware/drivers/d...
func copy(f *os.File) { … }
I think f is more than clear enough as a parameter name. The same can be said of variables where the type is easily inferred from the declaration or initialization.
Apple has one 82 characters.
https://developer.apple.com/documentation/contacts/cnlabelco...
On the other hand when you are dealing with a file pointer in a language that deals with file pointers constantly, “fp” is meaningful enough.
CN_LCR_BiaoMei = 77,
// either mother's sibling's daughter or father's sister's daughter
// aka female cousin involving at least one female parent
), while the other adds a constant tax on all uses of the variable (well, constant in this case) in perpetuity.A lot of lisp material features single letter vars, for example. (Actually, var length in prod code bases seems related to scope. So toy examples with obvious context see single letters strewn about. But digging through the classic books will rub off on the budding programmer...
Here's the very opinionated documentation: https://www.kernel.org/doc/html/v6.5/process/coding-style.ht...
I also want to note that the 80 columns limit was bumped to 100, and is no longer strictly enforced: https://www.phoronix.com/news/Linux-Kernel-Deprecates-80-Col
C - not so much. In C you are more likely to see acronyms. It's also interesting that for some reason, macro names are spelled out in full, but function names are abbreviated. So, typical C looks like:
LONG_SCREAMING_MACRO_NAME_(uh, oh);Initially I wrote a python program for flexible querying & summarizing of what the threads of interest are doing (psn) and then wrote a C version to capture & save a sampled history of thread activity (xcapture) [1]. I ended up spending too much time optimizing the C code - as just formatting strings taken from /proc pseudofiles and printing them out took very little time compared to the kernel-dives when extracting things like WCHAN and kernel stack via the proc intereface.
That's why I've since built an eBPF prototype for sampling the OS thread activity. The old approach still works even on RHEL5 machines with 2.6.x kernels without root access too :-)
It would be cool if they just said "Everything here will be TOML" or something.
But I like to stick with higher level tools anyway and avoid touching the low level stuff on Linux, so it's fine in practice.
char fname[BUFSIZE];
Also better to pass that in as a variable I don't think you can return a char array allocated on the stack like thatThis requires a GNU xargs that supports NULL termination (or something compatible); as I understand it, this cannot be done with POSIX tools.
$ cat shps
#!/bin/dash
for path in /proc/*/cmdline
do p=${path#*/} p=${p#*/} p=${p%/*}
case "$p" in *[!0-9]*) continue;; esac
c="$(xargs -0 echo < "$path")"
[ -n "$c" ] && printf %6d\ %s\\n "$p" "$c"
done*You can do this with `tr`, right?
#!/bin/dash
for path in /proc/*/cmdline
do
p=${path#*/} p=${p#*/} p=${p%/*}
case "$p" in *[!0-9]*) continue;; esac
c="$(tr '\0' ' ' < "$path" | sed '$s/ $//')"
[ -n "$c" ] && printf %6d\ %s\\n "$p" "$c"
done
Those escape sequences are covered in the standard: https://pubs.opengroup.org/onlinepubs/9699919799/utilities/t...(But please correct me if I'm missing something!)
This also avoids (POSIX-)undefined behavior with `echo` in case one of the processes' `argv[0]` happens to begin with `-n` or any argument contains a backslash.
This part can determine if the descriptor is seekable or not. It's a noop if it is and returns an error if it isn't.
> It was quite clear from the strace output that the overhead of the interpreter costs practically the same as running the program itself.
This kind of idea keeps repeating and... it's misplaced. You can't use a high level API like glob and expect it will do the same minimum of work as your trivial implementation. This has nothing to do with the interpreter itself.
It’s probably easier to just use getline(), it does basically the same thing and is in POSIX.1-2008[1].
[1] https://pubs.opengroup.org/onlinepubs/9699919799.2018edition...
A general review of the stdlib https://nullprogram.com/blog/2023/02/11/
A review of scanf https://sekrit.de/webdocs/c/beginners-guide-away-from-scanf....
Originally I implemented `getLine` (terrible name!) as a way to get multiple lines of input from stdin in a relatively safe way. The implementation borrows almost all of its ideas from the `sekret.de` post. Because I knew the implementation and it was close to hand on my machine, I used it for this version and just swapped out `stdin` for just any old file.
Edit: here's the implementation for reference, https://github.com/adammccartney/algorithms/blob/master/libs...
The actual solution is to use C++ instead of C :)
int s_isdigit(const char* s) {
Name does not fit implementation. Should be "contains_digit".Also I would say almost any "issomething" method involving a loop has an opportunity for an early return.
int s_isnum(const char* s) {
int result = (*s != '\0');
while (*s != '\0') {
if ((*s < '0') || ('9' < *s)) {
result = 0;
}
s++;
}
return result;
}
otherwise, it will choke on directories like `/proc/etc64/` or `/proc/net6/` if any such directory is added. (Plus other bits like early return, but that's not a correctness bug.)