Following the Unix philosophy without getting left-pad (2021)
raku-advent.blog
raku-advent.blog
The other point is that GNU utils are self-contained binaries and don't usually include transitive dependencies 20 levels deep.
Even when the problems these packages solve are not 100% trivial, unless you trust that library to be continuously maintained, maybe consider not adding the dependency, instead taking inspiration from the code and writing your own, maybe even simplified version of it (careful with that, you don't want to write your own crypto, for example - but when we're talking about a three-file "compare XML files" test util that hasn't been updated since 2014, just copying that code into your project might be the better approach).
[0]: And it really is JavaScript - I don't know of any other language with the same kind of explosion of micro-packages. I'm not saying that other languages never experience dependency hell - of course they do - only that it's usually not as horrifying as it is with JS.
The biggest difference here is that distros vet their maintainers, often with a lengthy process similar to hiring someone at a company, while NPM and similar repos are entirely self-serve, where anyone can register an account and just publish whatever packages.
This is very surprising to me. During the debacle over Rust in the Python cryptography package maintainers were saying it's unreasonable to expect them to read dependency changelogs.
They do not provide precompiled files for everyone and suddenly gcc isn't enough to build.
Debian (and I guess ubuntu) are stuck on a year old version because of using rust before packaging rust stuff has been solved.
They aren't the most careful of maintainers, it seems.
It was end users of the package that didn't know, had a dependency like "<=X" and with the big change in a minor version got the rust version they could not compile (and thus run).
On the plus side it gave the impetus for Python to define wheels for musl and other arch’s which is a nice outcome.
The AUR wants a word :)
If I wanted to do my own work I'd go back to Gentoo. ;)
Mostly because UNIX programs don't actually call other programs all that much. Composability happens on the shell level, but not so much in the program code level.
e.g. The implementation of 'cat -n' isn't 'cat | pr -n' or the equivalent C, and pr usually isn't in a separate package from cat.
It is about “coding the permitter, not the area”, so if you have NxM features to build, don’t write N*M individual programs/functions, instead write N+M programs/functions that can be composed together to get all NxM features. This video explains it well: https://youtu.be/3Ea3pkTCYx4
If anything, the UNIX philosophy is more like "do one thing, and do it to a text file".
You can do a surprising amount with flat files. For decades HP / HPE has been using flat files for Serviceguard / SGLX which is a high-availability cluster software that gets used in some very large Enterprise environments.
For example, tables in relational databases like sqlite are representationally isomorphic to tabular files like the output from `ls -al`. Sqlite has a lot of useful performance optimizations (including eg, not having to encode strings as "\x20" or the like if you want arbitrary bytes), but those come at the cost of a data format that you can't easily (not even notice that you needed to and did) reverse engineer, when either writing new software to consume or emit it, or even just visually inspecting it in a editor that doesn't already speak that format.
The main thing you can do with them though, is tell yourself that you don't really need a spec or documentation because you can just eyeball the output, and not having to follow a spec makes things feel simple.
Relatedly: Brian Kernighan's Unix: A History and a Memoir was an enjoyable read.
https://www.amazon.com/UNIX-History-Memoir-Brian-Kernighan/d... (not an affiliate link).
what you get if you follow this pattern is git, not left-pad. git high level commands were derivative higher order combinations of underlying programs that are quite conformant to this aspect of the unix philosophy. their newer replacements in many cases no longer compose from single purpose programs, instead importing aspects of their function as code, and specializing new behavior in order to better meet user interface demands. This is a direct repetition of the spell checker UX story.
At the language level rather than system utilities, think of something like boost. You install one dependency, and you get hundreds of individual libraries, each of which does only one thing and one thing well, and you can link only what you need into your own program. UNIX philosophy followed, without turning your dependency tree into combinatorial explosion.
The micro-packages idea would require us to take glibc and break it out into thousands of different packages. Perhaps one packaged library for each individual system call? That's about the level that the "left pad as a package" idea operates at.
If you want a larger javascript library that provides fills for many common string functions that are missing in the broader language and keep those separate from a group of array functions, that would be a reasonable middle ground and would approach more closely the actual reality of a glibc based distribution of utilities.
It's either that, or you have to accept that these small cases really do need to be folded back into a "standard library" that gets distributed along side whichever JavaScript environment you're using. This is something that languages like ruby, python and go were much better positioned to solve while still having a large repository of third party dependencies you can easily incorporate into your project.
An expressive language almost inevitably leads to more than one way of doing things and those different ways of doing things are often fundamentally mutually incompatible (i.e. a better way of writing the code will not save you). That essentially fragments the standard library and any single utility packages.
See e.g. the absolute profusion of alternative "standard libraries" for Common Lisp, Haskell, etc. Even ones that rein things in a little bit like Clojure still fragment when it comes to things like concurrency, with a lot of 3rd-party libraries purporting to replace core libraries (e.g. core.async vs manifold).
1. parse command line arguments
2. parse an input format
3. do its one job
4. format output
Two or three out of these are superfluous and could be shared with a larger program if it were just a function in it.
For instance, regarding (1) if the program just has a "dry run -n" and "verbose debugging -v" flag, and it becomes a function in the larger program which already has those flags, all the function has to do is refer to the configuration.
Regarding (2), the function probably has access to some data structure constructed by the larger program rather than raw text.
Regarding (4), the function may be able to put out a data structure for whose type the larger program already has an output format.
Unix itself showed that the one-program-for-one-task is a weak idea when Awk was developed: a programmable utility that allows a single program running in one process to avoid using numerous Unix tools.
They don't "parse" command line. They use getopt/getopt_long, which does the parsing. Besides using library functions, GNU coreutils commands do share code.
Wouldn't be nice to let the compiler package only the functions really used in the code and discard the rest? Do any language provide this? (C/C++?)
In Javascript, webpack (as well as other module bundlers) supports it, though they call their dead code elimination mechanism "tree shaking".
Yeah. No.
Dependencies come with risk and cognitive load: deleted like left-pad or loaded with malware like ua-parser-js. Sometimes that is worth it. But if you need 10 lines to pad a string or validate a number, copy some code off of stack overflow like a real developer. Nobody has to know.