Linux utils that you might not know
shiroyasha.io
shiroyasha.io
Especially looking down the list of all time greats: http://www.commandlinefu.com/commands/browse/sort-by-votes
...In hindsight, of course...
The gem I just plucked out is something I've been curious about for a while but never looked up:
CTRL-X e
The shell will take what you've written on the command
line thus far and paste it into the editor specified by
$EDITOR [then run it when saved]
Similar to `fc` except you don't need to run the command before invoking the editorRun:
set -o vi
once after you log in (for ksh / bash and compatible shells only, maybe, not sure about csh), or (better) put that line in your .bashrc or similar startup file, so it runs each time you log in. (I used to use "ksh -o vi" earlier, before I knew about "set -o vi" or before it existed, but in that case, it has to be the last line in your startup file, otherwise the other lines below will not run until you exit that (sub)shell.)
Then, when typing a command at the command line, just press ESC then v ; it does the same as what you said.
You can also do ESC :q! (in the editor, if it is vi) to quit without running the command you just edited, or save the command to another file for editing later at leisure, then quit without running it right now.
In fact, "set -o vi" also enables limited editing in vi mode right on the command-line, after you press ESC - you can use the command-mode commands of vi (h, l, b, w, f, F, and more) to move around, change characters or words, can also overwrite or append or insert text, etc.
You can even use / and ? and n and N to search backward and forwards in the commmand history to find (by substring) a previous command, to edit it. Once you find the right commmand, just press v.
Great for productivity.
open http://example.com/<up>, <ctrl-a>, type sudo and enter is nearly as quick and much more explicit for me.
> echo "Test"
> sudo !!<enter>
> sudo echo "Test"
You can then modify the command if you want to.
- ifdata: do not parse the output of ifconfig/ip anymore. Just use this tool.
- sponge: when you need to overwrite an input file at the end of pipe. sponge will wait for the pipe to end before overwriting, preventing any data loss
- vidir: edit a given dir with your EDITOR. Awesome for mass renames/deletes.
- ts: add timestamps to a command
- parallel: C implementation of GNU parallel (in perl). Very small, very fast, just does the core (running commands in parallel), and does not try to take over xargs.
There are others of course.
For personal use these are great. But if you write a shell script to be executed on different machines where you can't install anything, these tools won't be available to you.
And if you don't have that requirement, i.e. you can install everything, just install appropriate Perl/Python/Ruby libraries and use a proper scripting language.
(BTW, I tend to assume that Python is installed on almost all modern Unix systems by default, hence I prefer writing portable tools in Python instead of portable Shell code.)
No?
It is often possible to install things locally, i.e. not using apt. The idea that everyone has sudo privs is unrealistic. A version of apt that added a per user subset would be very useful.
I regularly run "one-off" stuff in a `guix environment` that includes all the required packages for that command, without polluting my user (or system) profile.
For example, I recently wanted to transfer a file over HTTP from one machine to another, and did not have the "python" executable in PATH. So I started it in a `guix environment` and "containerized" it (using user namespaces) just for show:
guix environment --container --network --ad-hoc python -- python3 -m http.server
That starts the HTTP server from the current directory, but the process can not see anything else from the "real" system.And having a system using Python on startup is a pretty good indicator that Python will be preinstalled.
However, I see that the situation may be different in specialized distros.
This bit us in our ass with centos 6.8 , when another package required python 2.7... If you install 2.7 directly, your system fails with a thousand cuts. A chroot is needed for that, unfortunately.
As a general rule, never try to replace any program or library supplied by your distribution.
This installs it to /opt/rh, and you can run commands that need them with "scl enable python27 mything" or by sourcing /opt/rh/python27/enable in a script to set up PATH, etc.
Nevertheless, statically linked binaries are indeed a good alternative, and provide good-enough portability in a wide range of cases.
And from your list, Go doesn't really belong there, it outputs statically linked binaries. You can just plonk them anywhere.
When you're using Python/Perl/Ruby you'll often have to use pip/cpan/rvm for dependencies, and you're back to square one.
combine file1 not file2
1. http://manpages.ubuntu.com/manpages/zesty/en/man1/combine.1....Ah, HN downvoting neutral comments again and I don't like it, but I have a policy of not upvoting anything out of pity.
PS: do not worry about karma ^^
1. http://manpages.ubuntu.com/manpages/zesty/en/man1/comm.1.htm...
* Newer versions of sort have an option for running sorts in parallel. If you're using an older sort, you can split the files, sort them individually with GNU parallel, and use --merge to combine them.
* If you have scripts that read data files, process them, and output more files, consider using a Makefile.
* tmux is a good way to leave a development session on a server that you come back to at the start of the day, or that you run long-running processes in.
* Setting a soft ulimit system-wide is a good way to avoid accidentally running the machine out of memory, causing your C libraries to be paged out to disk, and usually requiring a reboot. Since it's soft, you can override it at any time.
* strace and gdb can be run on just about anything to see what it's doing. If it's an interpreted language, you can often attach to a live process, call a function to invoke some interpreted code, and not have to kill & restart the process to fix a bug or oversight. Or you can attach to a Postgres worker process to get a stack trace, which often tells you why your query is slow if it runs so long you can't get an explain analyze.
* Look in /proc to find tons of useful information about running processes and settings. For example, you can tell how many filehandles a process has open, and sometimes what file it corresponds to. Just yesterday, a grub update caused a server to boot the kernel with the command line replaced with 2 random characters and a newline, which would be harder to debug without /proc/cmdline.
* pushd and popd are useful in shell scripts to temporarily change directories without forgetting where you were.
Dear sweet hypnotoad, don't use Makefiles for this. Scripts in a pipeline are perfectly well suited for ETL. They have the advantage of using the same language as the command language (Makefile is not shell, and when it differs it's a significant surprise). Plus, you can drop a script into a dir in your PATH and it will work in any project.
Use make if you have a project which expensively computes assets and reusing partial outputs is possible.
> pushd and popd are useful in shell scripts to temporarily change directories without forgetting where you were.
I would start with
cd - # go back to where I just came from
Which handles most of the pushd/popd use cases (outside of writing a script).Really, there's a whole separate set of skills for effective interactive shell use & writing great bash scripts (e.g. parameter expansion and history manipulation).
is that an advantage? Do you have the time to explain this a bit more?
I feel that new users need less expressiveness, to avoid decision overload, and keeping one automatic directory save point is easier to mentally manage than a stack of them.
I do recommend using pushd/popd in shell scripts (always) and interactively (if you must), but I think 'cd -' should be the first thing you introduce to newcomers w.r.t tracking working directory changes.
That said, I don't disagree that cd - is a good introduction to the idea that there's more to the cd command than meets the eye. (I'm also a big fan of the CDPATH variable, despite its issues.)
$ units '4123412312312 bytes' 'tebibytes'
4123412312312 bytes = 3.7502217 tebibytes
4123412312312 bytes = (1 / 0.26665091) tebibytes
$ units "50 miles per gallon" "liters per 100 kilometers"
reciprocal conversion
1 / 50 miles per gallon = 4.7042917 liters per 100 kilometers
1 / 50 miles per gallon = (1 / 0.21257185) liters per 100 kilometersExample: Created two files (1.txt and 2.txt)
[efg@fedori ~]$ cat 1.txt
file1: line1
file1: line2
[efg@fedori ~]$ cat 2.txt
file2: line1
file2: line2
In order to merge the lines of these two files we can use paste:
[efg@fedori ~]$ paste 1.txt 2.txt
file1: line1 file2: line1
file1: line2 file2: line2
Also check out 'man fold'
https://github.com/sahib/rmlint
There are great GUI visualizers for disk space, like WizTree/Baobab; for the command line I've had success with ncdu.
join -t$'\t' <(cat c <(comm -13 <(cut -f1 c) <(cut -f1 d) | sed -e 's/$/\t/') | sort -k1,1) <(cat d <(comm -23 <(cut -f1 c) <(cut -f1 d) | sed -e 's/$/\t/') | sort -k1,1)
Then I stumbled on an easier version with using just join join -a1 -a2 -o auto f1 f2
There are really too many unexplored things on the linux command line for a typical dev.https://www.nrl.navy.mil/itd/ncs/products/mgen
ps. ...I know, I know, but sadly no way to make them change the protocol used.
All-in-all it's arbitrary anyway. :P https://xkcd.com/1073/
bratch@serenity ~ $ LC_TIME=en_GB cal | head -n2
May 2017
Mo Tu We Th Fr Sa Su
bratch@serenity ~ $ LC_TIME=en_US cal | head -n2
May 2017
Su Mo Tu We Th Fr Sa for extra in fun
do
watch -t -n 1 "factor \$(date +%s)"
done alias primetime watch -t -n 1 "factor \$(date +%s)" 1) write to journal "I'm going to overwrite this file with this data (zeros)"
2) commit journal
3) write data to file
This is typical for journaling filesystems-- step 3 can be interrupted by a crash and replayed later (by re-reading the journal).For filesystems with CoW data (ZFS, btrfs), the in-place data will probably not be overwritten.
I would assume that the concern with ext3/4 in data=journal mode is that shred does not guarantee that the records of previous writes are evicted from the journal.
Note that the ext3/4 journal is a redo log, not an undo log. Old file contents are not copied into the journal on a write.
Thus, I don't see why shred should be less effective in data=journal mode compared to the other journaling modes.
CoW file systems are a different story. They don't allow you to overwrite physical file contents. You have to set the +C (FL_NOCOW) flag, which is, by principle, only effective for a file that does not have any contents yet. Thus, you can't set +C on an existing file and overwrite it's contents.
For one thing, if you have a contiguous file and you update some (but not all) bytes, putting them back in the original location allows the file to stay contiguous.
Also, if you write the data back into the original location, you don't have to update metadata such as inodes. Now, you may say that's less data, but on a spinning disc, there is some threshold below which the amount of data written doesn't matter much at all and it's the number of seeks that matters more. That is, if it's a choice between a single 50k continuous write or two separate 1k writes in different locations, the single write is probably quicker. (But this falls apart eventually of course.)
Whether these reasons are enough to prefer updating in place is another question, of course. But it's not like there isn't any benefit at all.
Because with data=journal, content that was previously written to a file (and made its way through the journal) might still be in there if the journal has not been replayed or garbage-collected in a while.
Many disks have a full erase command, which asks the internal controller to do the erase.
In general though, I use a hammer.
.30-06. >:)
Source: https://www.gnu.org/software/coreutils/manual/html_node/shre...
$ curl -s http://datascienceatthecommandline.com | pup '.sect3 > h3 text{}' | head
alias
awk
aws
bash
bc
bigmler
body
cat
cd
chmodReturns just the filename plus extensions e.g.
basename a/b/c/d/foo.php returns foo.php
Very useful for shell scripting.
$ basename a/b/c/d/foo.php .php
foo
The complement to basename is dirname: $ dirname a/b/c/d/foo.php
a/b/c/dEDIT: fixed the substitution.
The reason I use python whenever I need something "shell" like. Cryptic symbol salad that makes code golf fanatics green with envy.
Gotten a lot of mileage out of the various seldom used options of uniq, diff and cut too.
https://joeyh.name/code/moreutils/ https://linux.die.net/man/1/sponge
Now I found that full text search is man -K, but it searches through sources so it can also be quite useless. Is there a desktop full text search for man pages? It should use rendered man pages. It could be a fun project. Google doesn't count and I use it already.
Going through the core utils manual [1], I found a couple other useful looking commands I didn't know about. Previously I used awk in bash scripts primarily for printf, but apparently it's part of coreutils too. The timeout command also looks useful.
1. https://www.gnu.org/software/coreutils/manual/coreutils.html
https://marc.info/?l=openbsd-cvs&m=147272695616766&w=2
https://marc.info/?l=openbsd-cvs&m=146826185625800&w=2
Hey its part of bsdGAMES not bsdGETSHITDONE :)
locate /\*xyzGlobPattern | xargs grep -E '\<abc[(]'
The problem with sector access still stands, so you're not sure it's actually overwritten. You should use sfill(1) in addition to shred. sdmem(1) also, if you don't want to physically power down your machine.
People thought that you needed multiple overwrites for traditional drives but in reality that data was gone after a single overwrite of 0. There were no "advanced techniques" that could recover the data. Some specs suggested multiple passes, but this was because they were using precautionary principle.
> This is not the case for SSDs, so a single pass of shred (either random or zero) is enough here. Multiple passes will just kill your SSD faster.
SSDs do weird things with the data, so if your data is important enough you should destroy the drives. There's some data tucked away in odd places.
It's not about people thinking unrealistic things. The command was created for defending against a demonstrated attack.
It's not possible, and it's never been possible unless we're talking about 1970s 24" platters, and we're not talking about those.
You really didn't. It's a really persistent myth that there's some secret technique to recover overwritten data. For PC hard drives "blobby bits" has never been a thing. It might have been a thing on 1970s style 24" disc platters, but we're not talking about those.
No one has ever recovered data from a drive that's had a single overwrite of zeros - no software claims to be able to do this, no recovery service claims to be able to do this, there are no published papers that claim to be able to do this (there's one that gets an accuracy of 50%-55% per bit, ie useless).
Since the early 1980s a single overwrite of 0 has been sufficient to destroy data.
And if you're worried about a well funded government agency that has the money to spend on exotic techniques you can do a single pass with random data, or NIST 7 passes, or you destroy the drive.
It does effectively shred your SSD. In that it will fail sooner and sooner the more you run that command.
SSDs have a sector erase command. I don't know how it's exposed on the CLI, but it exists and will effectively erase anything that was removed. Most of them will run it automatically on the background after you delete stuff, so it's normally just a matter of keeping it powered for a while.
This makes for good reading: http://porkmail.org/era/unix/award.html
At least with "in appropriate use of cat" (as some call it) you're literally just swapping the stdin file stream with a disk io file stream so there's no functional difference what-so-ever.
I'm not saying I agree with the GP either though as most of the time complaints about "in appropriate use of cat" are just showboating. Using `cat` "inappropriately" is arguably more readable for less seasoned shell script developer and it's certainly a more logical program flow for a human to parse. ie "open file, grep for contents, do something else, etc". But it's still sometimes worth a reminder that many string processing tools can accept file input directly without the need for piping it via stdin (or the files can be redirected directly from the shell via the less than, `<`, token).
amirite?
pgrep often fails me I think mostly for apps the edit their cmdline and it needs to whole word match I think?
ps | grep [g]rep
Though it is still a bit terrible.If you mean adding a 'cat' at the start to provide a clear entry-point for the data, you can actually put the redirect at the beginning:
< example.txt grep cheese | tac > cheesy_lines.txt
<input.txt tr -d \\n >output.txtYou get a blank screen, but if you do:
cat file-that-doesnt-exist.txt | less
You get a nice error message:
cat: 'file-that-doesnt-exist.txt': No such file or directory
$ < not_existant_file less bash: not_existant_file: No such file or directory
I just tried this on bash 4.3. Cat remains superfluous in this case.
du -hs * | sort -h
du -hs * .* | sort -h
$ find . -mindepth 1 -maxdepth 2 -print0 | xargs -0 -- du -hs
Not that I would write it in a shell. ;)
Also paste
factor 100000000000000000000000000000000000000
100000000000000000000000000000000000000: 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 5 5 5 5 5 5 5 5 5 5 5 5 5 5 5 5 5 5 5 5 5 5 5 5 5 5 5 5 5 5 5 5 5 5 5 5 5 5
http://www.wolframalpha.com/input/?i=factor+1000000000000000...
also numfmt bug:
$ numfmt --to=iec 999999999999999999939999999 828Y
$ numfmt --to=iec 999999999999999999940000000 numfmt: value too large to be printed: '1e+27' (cannot handle values > 999Y)
but 999999999999999999939999999 is ~1000YB : http://www.wolframalpha.com/input/?dataset=&i=99999999999999...
== 10 ^ 38
== (2 * 5) ^ 38
== 2 ^ 38 * 5 ^ 38
== 2 * 2 * 2 * 2 * 2 * 2 * 2 * 2 * 2 * 2 * 2 * 2 * 2 * 2 * 2 * 2 * 2 * 2 * 2 * 2 * 2 * 2 * 2 * 2 * 2 * 2 * 2 * 2 * 2 * 2 * 2 * 2 * 2 * 2 * 2 * 2 * 2 * 2 * 5 * 5 * 5 * 5 * 5 * 5 * 5 * 5 * 5 * 5 * 5 * 5 * 5 * 5 * 5 * 5 * 5 * 5 * 5 * 5 * 5 * 5 * 5 * 5 * 5 * 5 * 5 * 5 * 5 * 5 * 5 * 5 * 5 * 5 * 5 * 5 * 5 * 5
Are 2 and 5 not prime?
999999999999999999939999999 / 1024^8 827.18061255302767482177