Moreutils – Unix tools that nobody thought to write (2012)
joeyh.name
joeyh.name
But what immediately stood out to me is `vidir`. I really like the idea of editing file names with an editor. Using loops and regex in a shell for mass renaming can be a mess. It should be way easier with `vim`. This tool made me install moreutils.
Also, this can be done in emacs using wdired.
Edit: Or use a more portable terminal FM, nnn pops to my mind
https://pubs.opengroup.org/onlinepubs/9699919799/utilities/f...
https://pubs.opengroup.org/onlinepubs/9699919799/utilities/v...
I often used :r !ls for that, thanks for the tip + it's shorter
Eshell is interesting, but I can't use the bashisms I'm used to, and I can't copy commands back-and-forth between the prompt and a standalone script.
shell-mode lives in-between these two extremes: it runs a normal shell, but Emacs manages the buffer. I don't know about ansi-term or eshell (or alternatives mentioned by others), but shell-mode handles emoji fine, as well as progress bars and colour codes; I think it defaults to TERM=dumb, so many programs won't output colour, etc. unless you override it to something like TERM=xterm-256.
Regarding curses, I found that there was two solutions. The first is to automatically spawn curses apps in a proper terminal emulator, you just have to setup the `visual-commands` variable properly. The second alternative is to replace curses apps with Emacs apps, e.g. htop to helm-top. Personally, I ended up going the second route after a while, as I realized that there are actually very few curses apps that are important to me, and that Emacs apps are better integrated if you use Emacs for everything else.
If you rely on a lot of curses apps, a “real” terminal like emacs-libvterm may however suit you better. Renders curses apps and emojis as well as any other terminal I’ve used. It’s also much faster than ansi-term and friends at rendering.
Also if you happen to use fern[0] you can mark off multiple files and directories, hit a hotkey and now you can edit their paths in a Vim buffer.
I just use ranger for that.
Outside of `moreutils`, one of the utilities I always install on all my machines is `atool`, which is just so much nicer and more intuitive than trying to remember the command-line options needed to handle the various tar.*, rar, zip, 7zip, lzip, etc. archive formats from the command line.
https://github.com/skx/sysadmin-util
Later they were replaced by a busy-box style collection of utilities written in golang (mostly to ease installation):
* tree - print directory structure
* bat - cat, but actually designed for reading files. Syntax highlighting, line numbers, automatic paging.
* rg - better grep
* direnv - local environment variables
Never going to happen because it's written in rust. ag (https://github.com/ggreer/the_silver_searcher) on the other hand is written in C.
(I have some patches to improve things but I have not been able to submit them to LLVM yet, and with those patches I did manage to get a working x32 build of rg on my system. I hope to be able to do so in the future.)
https://doc.rust-lang.org/nightly/rustc/platform-support.htm...
AFAIK, rust is still marked as "guaranteed to build" on these platforms, but assume only Linux. BSDs, not so much.
If Rust being added to the Linux kernel[1] isn't far-fetched, I don't think adding a utility in Rust is crazy either.
[1] - https://lore.kernel.org/lkml/CAKwvOdmuYc8rW_H4aQG4DsJzho=F+d...
Isn't it the case that LLVM doesn't support all architectures that Linux does?
https://github.com/fishinabarrel/linux-kernel-module-rust/is... is a chart of Linux architectures and whether Rust and LLVM support them.
Why is bat better than less?
I admit to catting files for reading as much as the next guy, but I almost always have the tiniest regret that I didn’t feed it into less and kept my terminal cleaner.
As for RG and AG. other than being faster, how are they better? I thought they were api complainant drop-in replacements so it’s not like they fill in a missing role.
To grep? Not that I know, I don't think either respects POSIX grep to say nothing of GNU grep. Though the trivial usage is identical, and many basic options carry over (e.g. -C and friends).
> other than being faster, how are they better?
Way better defaults for one. `rg` defaults to being recursive, ignoring binary files, respecting various types of ignore files, not being limited to BREs, coloring, …
They also have useful convenience features like "file categories" (e.g. `-t` will expand to a bunch of predefined include glob patterns so you don't have to input them by hand, which can be tedious), or parallelism (recursive ripgrep will parallelise searches across files, with grep you have to remember to combine `find`, `xargs -P` and `grep` for that to happen).
Could you please let me know how you came to that understanding? Because I'd really like to fix it. ripgrep was never intended to be POSIX compatible. It would have been a straight-jacket over its functionality. See: https://github.com/BurntSushi/ripgrep/blob/master/FAQ.md#pos...
> As for RG and AG. other than being faster, how are they better?
For ripgrep at least, a lot of it comes down to speed, better output formatting by default and suppressing results you probably don't want by default (but that can be turned off quite easily).
But there are other features, like better Unicode support, built-in support for searching compressed files, support for file types and a few other things. The README includes a bit more. ;-)
I have actually read a few of your (great) articles about ripgrep. I hope I didn’t come off as dismissive of your impressive work - though rereading my comment I should perhaps have chosen my wordS a bit more thoughtfully.
I did know that RG would skip more files than normal grep, so I guess I knew it’s not a safe replacement for all shell scripts.
But other than that, I’ve never thought of AG or RG as tools that brings something new to the table. I’ve always thought of them as a drop ins, with slightly new defaults and much faster speed. And the discussion related to TFA is tools that are missing, not tools that are better.
Maybe it’s a mix of the name and countess blog posts telling readers to use RG instead of GREP, that lead me to consider RG as a better grep and not a different tool.
`less` not good enough?
(Also, yeah, let's ignore that Rust in the standard libraries is still a problem.)
I agree, tree is great. It can be somehow replicated by using "find .", if you are in a hurry.
> bat
If you need that functionality, why don't you open the file with vim, for example?
> rg - better grep
Grep is pretty nifty, I don't see how could it be improved. What is the main advantage of of rg over grep?
> direnv - local environment variables
I read the manpage for direnv and I was really scared. What is it for? I never used an .envr. If I need to see my environment I can simply use "env". What is it missing?
The main advantage of rg is how much faster it is than grep. It’s WAY faster.
I don’t use direnv but I have colleagues who do. I think the idea is that it can set environment variables/run arbitrary commands when you cd into a certain directory (such as activating a Python virtual env, for instance).
I figure you can also use it in conjunction with `module load` to automatically load the right module when you enter a project's directory. Modules are used in many computing clusters to manage environments.
Also direnv will never execute an .envrc without permission, you need to run `direnv allow` whenever there is a new or newly modified .envrc.
function lenv {
path=$PWD
while [[ ! -f "$path/.env" && "$path" != '/' ]]; do
path=$(dirname "$path")
done
if [ -f "$path/.env" ]; then
source "$path/.env"
fi
}
function cd { builtin cd "${@}"; lenv; }It lacks the security but it works recursively and allows more than just environment variables. A common script I have is to automatically build for instance:
here=$(dirname ${BASH_SOURCE[0]})
function autobuild {
(
mkdir -p build
cd $here/build
../configure
while true; do
make && make check
inotifywait -qr -e close_write,delete $here/src $here/test $here/Makefile.am
done
)
}
I've tried generalizing the latter but between different targets, different build systems, different projects structures etc, copy and pasting something like this per project is the least worst option.If there is no .envrc in the directory it will look for one in the parent directory, but it won't activate more than one.
Also one thing it does that your hand-rolled version does not is that it remembers what the environment variables were before and restores them to that state when you exit the directory.
My above example wasn't running the command, just defining it to be run manually. The .env becomes a dumping ground for all sorts of project specific stuff. The docs say direnv doesn't support this.
> Also one thing it does that your hand-rolled version does not is that it remembers what the environment variables were before and restores them to that state when you exit the directory.
I've thought of adding this, but apart from a couple of things I reset manually I haven't really found the need.
You're right, you can't define aliases/functions using direnv. That's a shame, that does seem useful.
jot -- print sequential or random data
rs -- reshape a data array
vis -- display non-printable characters in a visual format
and from Unix >V7: apply -- apply a command to a set of arguments
mc -- multicolumn print
and from AT&T: tw -- file tree walk
And then the obligatory personal tools that have stood the test of time: align -- align text columns
crop -- crop lines to a width
dabl -- delete adjacent blank lines
dtb -- delete trailing blanks
emboxxen / deboxxen -- convert to/from box drawing characters
eol -- convert line endings
field -- simple line field extraction
freeze / thaw -- cross-shell synchronization
mergl -- merge lines into previous lines' blanks
pad -- pad lines to a width
put / take -- cross-shell pipe
uni -- unicode character properties and searchAnd seq in Linux is like jot for sequential data, at least.
And from my blog: An Unix seq-like utility in Python: https://jugad2.blogspot.com/2017/01/an-unix-seq-like-utility...
I've found I'm using the `put`/`take` pair, for splitting a pipeline across shells, more again in the WFH era where I often have multiple ssh sessions into a machine, when I would probably have used the clipboard in a GUI session. Also useful for things like seeing stderr from different parts of a pipeline in different windows. `put` is just `cat >>$(mkpipe "$1")` and `take` is `cat $(mkpipe "$1")` where `mkpipe` is
fifo="${TMPDIR:-/tmp}/fifo-$(id -u)-$1}"
test -p "$fifo" || mkfifo "$fifo"
echo "$fifo"Not all Unix tools are simple. E.g. make. But I know what you mean - the Unix philosophy.
https://en.m.wikipedia.org/wiki/Unix_philosophy
The TAOUP book by Eric Raymond (The Art Of Unix Programming) has a lot about that.
https://en.m.wikipedia.org/wiki/The_Art_of_Unix_Programming
And my IBM developerWorks tutorial / case study on Developing a Linux command-line utility (in C) may be of interest to people who want to write their own Unix tools that play well with others.
https://jugad2.blogspot.com/2014/09/my-ibm-developerworks-ar...
In fact the docs even say:
> make sponge buffer to stdout if no file is given, and use it to buffer the data from pee.
Translation, pee into a sponge!
This is by the way the reason why GraphicsMagick is better than ImageMagick (I still use the latter, though, because it doesn't cause the same problems for me as moreutils does, and it's just more popular than GM).
would it solve the problem if the author namespaced the commands with a hyphenated prefix? curious about these considerations, which it seems like you've spent time thinking about. what do you see as the "best practices" regarding a set of utilities that are maintained and published together?
I think if tools are absolutely unrelated, they just should be distributed as a separate packages. GNU coreutils is tolerated mostly because it's so ubiquitous (so much, that it causes Stallman to grumble about "you mean GNU/Linux, not Linux"). moreutils is late to the party, it isn't ubiquitous, and the usefulness of any tool in the package is questionable, so the author really shouldn't be so brave to assume that if he thought he needed all of them, everyone will.
If there is a reason enough to distribute a package with several separate callable binaries (as with ImageMagick), I think git or GraphicsMagick are perfect examples of how it should be done. Hypenated prefix is also ok. Even if your tool is supposed to be used by somebody 50 times a day, too long of a name isn't really a problem, since user can always just make an alias (as I do with most of the tools I use frequently).
Sure you can always combine binaries from 2 packages manually, but as I said, it just requires some tinkering, so I cannot simply have in some textfile a list of utils to install on a new PC in a matter of minutes.
The GNU build/install process let you set a prefix, so it was customary to specify `g` get prefixed, so `gmore` and `gcd` etc. if you were afraid of breaking old user scripts that used the sun versions.
The build for this looks pretty simple. I cloned the source, ran "make". I got some XML errors about docbook, but it did create a `sponge`
otool -L sponge
sponge:
/usr/lib/libSystem.B.dylib (compatibility version 1.0.0, current version 1281.100.1)
looks like one could copy it to `~/bin/`But in general it is bad practice to package a bunch of unrelated tools together. You should be able to install them separately to only have what you need. Youd want different packages for each command or sets of related commands with common dependencies. And then have a meta package that includes all of your packages for people who want to install it all in one go.
% sed "s/root/toor/" /etc/passwd | grep -v joey | sponge /etc/passwd
If you tried this as: % sed "s/root/toor/" /etc/passwd | grep -v joey > /etc/passwd
The shell would overwrite /etc/passwd with an empty file before sed has a chance to read from it.https://imagemagick.org/script/command-line-tools.php
Edit: I wasn't quite right, the commands still exist, but they are symlinks to "magick" and can be called using the subcommand mentioned above; you still have to delete the symlinks to clean up the namespace.
% sed "s/root/toor/" /etc/passwd | grep -v joey | sponge /etc/passwd
It looks like this allows an in place modification to the same doc. If you used % sed "s/root/toor/" /etc/passwd | grep -v joey > /etc/passwd
Bash will process the redirection first, then execute the commands in the pipe. So, it will have cleared out /etc/passwd with the non-appending redirection ">" before sed operates on it. You'd end up with a blank /etc/passwd file. (ymmv, I don't know how other shells would handle this)Sponge, it seems, would cause the pipe to allow the preceding commands to complete before it outputs.
(i'd welcome a method of doing that in zsh.)
If you install parallel & moreutils you get the GNU parallel as /usr/bin/parallel and moreutil's parallel as /usr/bin/parallel.moreutils. If you only install moreutils, it provides /usr/bin/parallel.
You can use the debian alternatives system to flip the order.
Debian also has namespace-collision policies and an 'alternatives' facility for deciding which of multiple implementations of a tool (gawk, nawk, mawk; vim, vim-tiny, nvi; python3, python2; etc.) is primary on a given system.
https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=597050
https://www.debian.org/doc/debian-policy/ch-binary.html#virt...
https://www.debian.org/doc/debian-policy/ap-pkg-alternatives...
It seems like it shouldn't be too much work to fork the repository, and tweak the packaging logic to only include this one binary. You'd still have to compile the whole thing every time, but rebasing would be straightforward.
The right way for distros to package this is to split the moreutils source distribution into moreutils-* packages, plus a meta package that pulls in all of the commands. So your complaint actually belongs in your distro's bug tracker -- it has nothing to do with the moreutils project.
It’s discussed on stack-overflow. [0]
Someone even wrote a nodejs tool for this functionality [1], but I would rather have something written in a compiled language.
[0]: https://stackoverflow.com/questions/10574794/how-to-list-onl...
> find . -maxdepth 1 -type f
A little clunky, but there's certainly no need to mess about in JavaScript land. If you use it often, create a shell macro.
$ mkdir -p tmp && touch tmp/a tmp/b tmp/c
$ find ./tmp -maxdepth 1 -type f
./tmp/a
./tmp/c
./tmp/b
$ ls -1 ./tmp
a
b
cfrom the man page:
> The format is identical to that produced by ``ls -dgils''.
$ find ./tmp -maxdepth 1 -type f -printf '%f\n' $ find ./tmp -maxdepth 1 -type f -printf '%f\n'
find: -printf: unknown primary or operatorI also use MacOS and I have setup the gnu userland on my local so it matches the linux environment on our servers and containers.
It's not about convenience. It is about having a development environment that matches the target runtime environment.
You can run into nasty surprises when the default behavior of a tool in your dev environment is different than in your production environment.
I've been writing shell scripts professionally for over 20 years and I have always taken this approach and it has served me well.
Works for me on OSX.
+ $ find /tmp/x ! -path /tmp/x -prune -type f -exec sh -c 'for p; do printf '\''%s\n'\'' "${p##*/}"; done' - {} +
c
b
a
Seems to do the trick and is POSIX only (so works on busybox and mac). You can probably just tack `| sort` at the end if you want it sorted.People value software that's easy to use. That command you typed is a Frankenstein's monster.
fls () {
find . -maxdepth 1 -type f
}
vs fls () {
find . ! -path . -prune \
-type f \
-exec sh -c 'for p; do printf '\''%s\n'\'' "${p##*/}"; done' - {} +
}
The user would just type `fls` in the directory they were exploring. $ (cd tmp && find . -maxdepth 1 -type f) $ find ./tmp -maxdepth 1 -type f -exec basename {} \;
b
a
c
I don't know the downsides of using basename though.Thanks to all the helpers, who took their time to propose solutions.
But what I was trying to say is: I would be ok with adding a new command to the tools I use on my machine, VMs and containers. fzf, fd, rg are tools which make my life easier. I would prefer to have the features of fls in ls itself or a tool similar to fls in coreutils, moreutils or another package .
ls -1p | grep -v '/$'
You may wrap it in an alias or .bashrc function. I use the reverse (list only directories): function lsd {
ls -1p $* | grep '/$'
}But I don't think it's too hard to solve this problem. A command which filters should also be able to invert the filter. One example for this is grep -v.
Coming from a more low level background but doing front end development from time to time, I am always amazed by the sheer amount of seemingly useless javascript reimplementations of what would be bash one liners.
For your question, how about "find <dir> -type f - maxdepth 1" ?
You can put that in an alias or script in your path
xclip -o | vipe | xclip
This one in particular, let's you edit your clipboard with $EDITOR.2016: https://news.ycombinator.com/item?id=12023277
2015: https://news.ycombinator.com/item?id=9013570
2015 (1 comment): https://news.ycombinator.com/item?id=9004302
sponge can be replaced with dd
Other than that, yeah nice ideas. And very much in the unix spirit of one tool to do one thing.
By the way, I think you’d be in even more trouble if you wrote:
cat /dev/zero > ased -i 's/root/toor/;/joey/d' /etc/passwd
But I get your point.
* https://cr.yp.to/daemontools/setlock.html
* http://cr.yp.to/daemontools/multilog.html
* http://cr.yp.to/daemontools/upgrade.html
* http://jdebp.uk./Softwares/nosh/guide/commands/cyclog.xml#CO...
sed "s/root/toor/" /etc/passwd | grep -v joey | sponge /etc/passwd
I think it can be rewritten as: sed -n '/joey/! s/root/toor/p' -i /etc/passwd(Would love to be proved wrong on this)
sed -i.bak 's/foo/bar/' filename
[1]: https://stackoverflow.com/a/22084103/3266847 sed -i ‘’ ...
works on BSD sed but not GNU. Meanwhile: sed -i’’ ...
sed -i ...
both work on GNU but not BSD.My guess is sponge buffers all the input and then sends it to output once stdin is closed.
> Unlike a shell redirect, sponge soaks up all its input before opening the output file. This allows constricting pipelines that read from and write to the same file.
Some commands (eg. sed) have an in-place option, but many don’t.
The defaults are a lot saner, and it's easier to pass arguments how you want.
Like I have subdir "/media/tnoko/sdc1/capture/Roinaa/".
"ncd Roi" should be totally sufficient and unique command to go there.
Yes, the find -type f focuses on files; nonetheless it's easy enough to right-click-select the relevant directory to open it rather than the file.
#!/usr/bin/env rc
if (~ `{find -type f | 9 grep $1 | wc -l} 1)
plumb `{find -type f | 9 grep $1}
if not
find -type f | 9 grep $1"rc - implementation of the AT&T Plan 9 shell".
Do I have to install that?
#! /bin/bash
d=$(find / -name "$1" 2>/dev/null | head -1)
echo Jumping to $d, Control-D to return
cd $d
bash
Actually it is now perfect. If the guess is wrong, you can return and add letters to you search string.But I already made my own. It is quite perfect:
#! /bin/bash
d=$(find / -name "$1" 2>/dev/null | head -1)
echo Jumping to $d, Control-D to return
sleep 1
cd $d
bashI have it bound to ctrl-e for edit: https://gist.github.com/FeepingCreature/5f575b6fcda041a48f58...
It's like adding the vscode file finder to bash :)
The additional code is to make bash "act as if you typed in the command manually". So you can arrow-up to find it in the history buffer with the filename expanded.
> sponge: soak up standard input and write to a file
I do that all the time...
cat > file.txt
Update: reading the comments on here, apparently sponge sucks up all content before opening the destination file which allows editing an input file in place. Minor advantage there.
grep foo file.txt | sponge file.txt
If you do this with redirections then file.txt will be truncated before it's been processed, leaving you with an empty file instead of what you wanted. Sponge collects its input first and then writes everything out at the end, so you can output to a file that was used as an input.
(Parent updated while I was writing. Oh well)
Commands that process file in place (sed -i) write to a temporary file in the same filesystem and then rename to the target file, which works if you want to process files that don't fit into memory.
If you do stick with bash, I would recommend giving shellcheck a try. It won't catch everything, but I've learned a lot from running it on my code.
On another note
> pee: tee standard input to pipes
Goddammit guys!
$ update-alternatives --list editor
Why not use > file?
> mispipe
In bash there's PIPESTATUS for that.
I'm assuming sponge buffers everything in memory so you don't get concurrent modification bugs.
$ cat test
moo
$ cat < test > test
$ cat test
$ $ cat test
moo
$ cat < test | tee test > /dev/null
$ cat test
moo
...seems to work. Am I just getting lucky with a race condition? $ seq 1 99999 > f.txt
$ cat < f.txt | tee f.txt > /dev/null
$ wc -l f.txt
23696 f.txProbably this is buffering, either from dd or from the kernel.
$ dd if=/dev/urandom bs=1M count=1 of=test.bin 2>/dev/null
$ cp test.bin test2.bin
$ cat < test2.bin | tee test2.bin >/dev/null
$ diff test.bin test2.bin
Binary files test.bin and test2.bin differ
$ wc -c test{2,}.bin
131072 test2.bin
1048576 test.bin
Notice the file got truncated at exactly 128k; a nice round size for a write cache. tac | tac
is a safe bet since must be reading all file before by definition.Sponge lets you do a series of piped transformations on 1 or more files and overwrite the result back to the original files.
Using > things will just break.
I’m guessing this also means you can’t pipe something bigger than what your memory can hold.
There's a good Stack Overflow answer here: https://unix.stackexchange.com/questions/207919/sponge-from-...
This, I learned the hard way, and I am not the only one. So in order for you not to make that mistake: the first thing that happens is that "file" is truncated, that happens before any command is run, so
sed s/foo/bar/ file > file
will always result in an empty "file", because it will be truncated before "sed" is run. I lost a couple of hours of work like that, once, and then I learned.From the manpage of sponge
sponge reads standard input and writes it out to the specified file. Unlike a shell redirect, sponge soaks up all its input before opening the output file. This allows constricting pipelines that read from and write to the same file.
So, the command sed s/foo/bar/ file | sponge file
will do what you expect sed s/foo/bar/ -i file
to edit file.This, btw, is why I hate shell scripts. There are so many variants of bourne shells and UNIX tools that writing a portable script is a minefield, as if properly dealing with spaces wasn't tricky enough...
I really like many of the moreutils tools, but dealing with constant parallel breakage (kind of a Hadoop for dummies) is annoying.