Decoded: GNU coreutils (2019)
maizure.org
maizure.org
Check man7.org for good, though brief info on many of them.
I had explored many of them a while ago.
Maintained by Michael Kerrisk, author of The Linux Programming Interface, a kind of reference bible for Linux APIs and system calls.
Edit: many of which are used in making such utilities.
https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...
perl -e '$! = shift and print "$!\n"' ERRNO
works to decode (nonzero) ERRNO if Perl is installed (per strerror(3) in the C runtime library linked to Perl, so YMMV on non-POSIXish platforms).Speaking of portability, Microsoft's err.exe[1] is a conceptually similar tool that references a considerably larger collection of Windows error messages (user and kernel mode Win32 error codes, COM HRESULTs, etc) and is therefore far more useful on Windows platforms than anything naïvely implemented in terms of C runtime errno (e.g., my stupid Perl one-liner).
In GNU coreutils, it's not really "another" ls; it's the same source code with a preprocessor macro set differently (ls, dir, and vdir all share the same source).
Edit: Also, some easter egg looking thing at the bottom right of the page:
<div class="copyright col-md-6">
##*#**##****#*#**/\##*###*****#**#*#*#**#******#**#*#*####*#*##*
</div>Edit: Fixed asterisks, I think.
##*#**##****#*#**/\##*###*****#**#*#*#**#******#**#*#*####*#*##*
HN's formatter is having a time with that many asterisksI don't recognize the format, can someone help me out?
First part is maizure
Ps: hn messed up with stars :)
## *# ** ##** **# *#* * == maizure
I can find some long words with a dictionary approach in there, like: ARRIVE -> .-.-..-......-.
CLEVER -> -.-..-......-..-.
DESTINY -> -......-..-.-.--
FENCED -> ..-..-.-.-..-..
MEMBER -> --.---.....-.
MISTER -> --.....-..-.
(etc)
But, too many variations that direction too.for the curious: https://www.jbowman.com/remorse/
Maybe an email address?
Some ignorant and probably cliched musing: when I look at small utilities like these I am always struck by a seeming distinction between best practices for little C programs versus best practices for large C applications (the author of the post touches on this ad well).
In particular, the explicit flow (including goto) and “pedantic” style is actually quite appropriate for something < 1000 lines and where the expected behavior is extremely well understood. In cases like pwd, mkdir, etc, trying to abstract too much is arguably a mistake for maintainability and understanding.
I say all this as an immutable functional-first dev who hasn’t done much native code :) And I think the various type-safe / memory-safe / etc versions of these tools are worth developing. But there’s something to be said about well-optimized native code that clearly “does what it says on the box” in a way that’s accessible to anyone who understands basic Linux programming - even if they can only contextually read C code.
(My only real gripe is typographic / linting related, mostly due to being a whippersnapper).
I'm not criticizing it in context -- a lot of this code dates back to the mid 80's if I'm not mistaken. But always write new code using scalable idioms.
There's two 'allowed' uses in C that are common and represent good code even today. goto error cleanup stubs, and goto in virtual machine dispatch loops.
The size of the codebase doesn't really matter for those cases; they're largely considered the idiomatic way to go about the problems they're trying to solve.
Even in C, if you're writing Microsoft only code, seh is probably a better mechanism than goto error.
I'd argue that the defer statement in go (and the surprising side effects of it, like that it's function instead of block scope like you might otherwise expect) ultimately come from trying to wrap this idiom in a construct that's better supported by the language.
My point though is that in relatively standard, portable C, there are valid, idiomatic use cases of goto, and it's not quite so easy to say 'eww goto' in those very specific circumstances.
Search for retry: or handle_itb: in https://github.com/torvalds/linux/blob/master/fs/ext4/resize...
Or fixleft: or copy: in https://github.com/torvalds/linux/blob/d158fc7f36a25e19791d2...
And frankly, the fix_left style code you see just isn't modern idiomatic C, IMO. In a code review I'd have them either combo of write a block comment explaining why it's necessary to be weird and a lot of test cases for when someone inevitably tries to rewrite it, or just rewrite it in the first place.
Some of the areas of the Linux kernel aren't exactly known for being the best written C (unfortunate as that is) and you're seeing some of that.
POSIX and similarly stuffy requirements (even if “soft”) means that this code is fairly static. While there is some bloat in the pragmas, etc., these applications are necessarily slow to change and I think it’s reasonable to say that they won’t suffer from feature bloat anytime soon. So the normal software risk considerations are a bit different here. Further, any changes to the code will be fiercely reviewed, and the individual programs are small enough that increases in complexity will be quickly spotted. Relatedly, these programs are small enough that, if a refactor to more structured code were necessary, the work would be quite feasible. So while the risks of goto are real in any C program, in practice I think they’re quite minimal here.
And I do think you’re missing an advantage. These are core userspace functions that perform safety- and security-critical kernel interactions. So I definitively agree there is a strong argument to use safe code, modern abstractions, and so on. This is especially true for modern PCs that really can afford to spend a few extra cycles creating a folder.
But a modern code construct, correctly applied, is only as safe as the compiler. This is not guaranteed! A common “gotcha” with buggy C compilers is inappropriately pruning instructions because the compiler optimizes away a loop or else statement. It is hardly a frequent issue but similar bugs have shown up in recent gcc/clang releases. And in particular core developers who are working on operating systems are more likely to be using shaky C compilers.
Using gotos and ugly global state has the distinct advantage that generated assembly tends to have less “surprises.” If there is a bug in the compiler it will be less well-hidden; if there is a bug in the program then there is less mental work between analyzing the C and analyzing the disassembly for debugging.
Again, in general I think you’re correct and that my argument is ultimately more of a judgment call.
EDIT: I didn’t really want to address any structural advantages of goto for, e.g. exception handling via breaking loops earlier, etc. I am not a domain expert enough to comment appropriately but it does seem there are cases where properly abstracted cleanup code in C is more spaghettified than a goto: https://lkml.org/lkml/2003/1/12/203
I do agree that the specific use of goto to jump cleanly out of several loops is appropriate: the problem is that C lacks clean constructs for exiting named blocks. That would be preferable to general goto and doesn't harm optimization, the flow graph is still easy to analyze, convert to SSA form and the like.
Having a handful of global variables reduces the amount of stuff being passed around from function to function; utilities don't need to worry too much about free()ing dynamic allocations, since that gets cleaned up on exit anyways; none of the code has to be re-entrant, because each invocation of the utility is running in its own process.
Decoded: GNU Coreutils - https://news.ycombinator.com/item?id=20328650 - July 2019 (55 comments)
What a coincidence!!!
Truly an amazing resource on GNU coreutils