Top Unix Command Line Utilities
blog.coldflake.com
blog.coldflake.com
Some issues:
- don't forget that /dev/random blocks
- It's easier to use dd_rescue to track progress than to signal dd
- Using dd to zero out a hard drive repeatedly doesn't increase security[1]. Using ATA secure erase does[2]
- an alternative for summing file sizes is
du -ch **/*.png
[1] http://en.wikipedia.org/wiki/Data_erasure#Number_of_overwrit...Also, for those wondering about the blocking of /dev/random, it will restrict the number of bits you can copy using dd, but this won't be apparent unless you attempt to copy more bits from the entropy pools than there are available for random number generation. For more information, see this question on Super User: http://superuser.com/questions/520601/why-does-dd-only-copy-...
If you fear the NSA seizing your disks, consider the tradeoffs of explosive disposal.
If you fear a technically-savvy reporter going through your trash bin, overwriting your disk three or four times with patterns will be fine. But it's probably faster to take a drill and make a couple of holes. Make sure you hit the platters.
If you are selling your old hardware and just don't want your unencrypted stuff to be recovered by a sixteen year old with no budget but lots of time, overwrite the disk once.
If you're trashing an SSD, make sure any patterns you use for overwriting are not compressed out of existence by the controller. Or pull off the controller and crunch it.
Zeroize the keys. Then you don't need to worry about the disk.
If your data is something that would be a problem if the physical disks were stolen, don't store it on-disk unencrypted.
http://computer-forensics.sans.org/blog/2009/01/15/overwriti...
http://www.howtogeek.com/115573/htg-explains-why-you-only-ha...
Multiple over writes is pointless. There's the Gutmann stuff, but that's ancient and the 35 passes was for multiple drive controllers, if you didn't know what drive controller was being used.
But then sometimes you don't have to do what works, but what other people tell you. Thus, if you're working to a standard it doesn't matter if DOD specifications are actually more secure than a single secure erase, you do what the spec calls for. And if you have to persuade other people that the data is provably gone it's easiest to just grind the drives.
Why? I just use: watch pkill -USR1 dd
ls, cd, git, ssh, make, e, cat, veille, rm,
wpa_supplicant, grep, evince, mv, x, dhclient,
cp, echo, todo, mplayer, scp, man, mkdir, ack,
pdflatex, apt-get, apt-cache, sed, less, feh,
racket, gcc, wget, xrandr, bg, svn, pmount,
for, gpg, halt, ping, tail, top.
"e" is an alias for emacsclient ; "veille" is a script which toggles between "xset s 5" and "xset s default" ; "x" is an alias for "xinit" ; "todo" is a script which manage a text file which I use as a todo list."ls", "cd", and "git" are far more used than any other commands : 14808, 13256 and 10078 times respectively, against 3919 times for "ssh" which is just behind.
I obtained these data from my .bash_history. Here are the place of the commands that are listed in the article :
"tr" is 76th
"sort" is 66th
"uniq" is 120th
"split" is there only one time like many other command so its rank is not relevant
Substitutions operations are what most "for" do so it is in my top 42. However see (1) below about the article.
Files size are a mix of "ls" for individual file and "du" for multiple files, "du" is 65th
"df" is 63rd
"dd" is 473rd
"zip" is 123rd (and funnily "gzip" is 122nd)
I didn't use "hexdump"
(1) About the following line for i in *.mp4; do ffmpeg -i "$i" "${i/.mp4}.mp3"; done
I have two remarks. First, using the "-vn" of ffmpeg would accelerate the conversion by making ffmpeg ignore the video entirely. Second, substituting with '/' in the Bash expansion is not the right way to do that. "${i/.mp4}" would be "$i" without the first occurence of ".mp4". It is '%' that you want to use here. cat ~/.bash_history | cut -f1 -d' ' | sort | uniq -c | sort -n -r
turns out my .bash_history is clipped at it's default size-limit (500 lines). So I'll change that to gather more data for next year. My results started with:
147 clang++
77 ls
54 cd
15 gs
14 rake
13 vim
...where gs is short for "git status" ...and thanks for the hint about the substitution! I updated it on my page.
$ hash | sort -nr | less
shows usage counts and full paths of executed commands.joe, ls, cd, time, cdbdump, tail, more, cat, rm, grep, wc, apachectl restart, find, curl, chmod, history, mv, locate, cpan, apt-get, pwd
But the most useful one is a command line Perl utility I called "flt" that executes a block of perl code for each in the stdin.
cat file.txt | flt ' $line=~ s|\s+| |gsi; print $line."\n"; '
That would compact free spaces
find . | flt ' if (-f $line) { print (-s $line)."\n"; } '
This would print the size of all files in the current folder and subfolders.
So it works like awk, but with full Perl, no need to learn awk syntax. You can do conditionals, loops and whatnot. I write 30% of my one time throw away scripts directly in the command line.
http://perldoc.perl.org/perlrun.html#*-n*
Those "flt" lines could be written
perl -lpe 's|\s+| |gsi' file.txt
find . | perl -ne 'if (-f) { print -s }'
(-l chomps the incoming newlines, and puts them back on the output)Of course, "perl -ne" is longer than "flt", and I appreciate all this implicit use of $_ is not to everyone's tastes.
Another VERY useful tool I didn't see on this list is iperf. From the Debian package description:
Iperf is a modern alternative for measuring TCP and UDP bandwidth performance, allowing the tuning of various parameters and characteristics.
Features:
* Measure bandwidth, packet loss, delay jitter
* Report MSS/MTU size and observed read sizes.
* Support for TCP window size via socket buffers.
* Multi-threaded. Client and server can have multiple simultaneous connections.
* Client can create UDP streams of specified bandwidth.
* Multicast and IPv6 capable.
* Options can be specified with K (kilo-) and M (mega-) suffices.
* Can run for specified time, rather than a set amount of data to transfer.
* Picks the best units for the size of data being reported.
* Server handles multiple connections.
* Print periodic, intermediate bandwidth, jitter, and loss reports at specified intervals.
* Server can be run as a daemon.
* Use representative streams to test out how link layer compression affects your achievable bandwidth.
I use iperf initially when I'm troubleshooting poor file server transfer speeds, for example. There's a pretty Java GUI too if you want that.
#!/bin/bash
#~/.scripts/runinbg
#Author: Khaja Minhajuddin
#Script to run a command in background redirecting the
#STDERR and STDOUT to /tmp/runinbg.log in a background task
echo "$(date +%Y-%m-%d:%H:%M:%S): started running $@" >> /tmp/runinbg.log
cmd="$1"
shift
$cmd "$@" 1>> /tmp/runinbg.log 2>&1 &
#comment out the above line and use the line below to get get a notification
#when the test is complete
#($cmd "$@" 1>> /tmp/runinbg.log 2>&1; notify-send --urgency=low -i "$([ $? = 0 ] && echo terminal || echo error)" "$rawcmd")&>/dev/null &http://unix.stackexchange.com/questions/9496/looping-through...
for i in *.mp4; do ffmpeg -i "$i" "${i%.mp4}.mp3"; done
will result in executing: ffmpeg -i foo foo.mp3
ffmpeg -i bar.mp4 bar.mp3
Which is obviously not what was meant. So it's a good habit to learn to loop over files in a directory in a different way. $ function echon { echo $# $*; }
$ >'foo bar.mp4'
$ for i in *.mp4; do echon "$i" "${i%mp4}mp3"; done
2 foo bar.mp4 foo bar.mp3 for i in `find -name '*.mp4'`; do # ...
or similar. In that case, the output of `find` is indeed split first, and `for` sees "foo", "bar.mp4", and so on.is something I use all the time to get sorted frequency tables.
Learn Linux the Hard Way talks about them quite a bit, is that the go to guide in 2012 or can I do better?
I also learned a lot from "The Linux Cookbook" (Second Edition) by Michael Stutz (this might be the first edition online: http://dsl.org/cookbook/cookbook_toc.html).
$ time find -ls | awk '{s += $7} END {print s}'
15970582120
real 0m27.721s
user 0m1.256s
sys 0m1.780s
$ time find | xargs wc -c 2> /dev/null | tail -1
604260969 total
real 0m0.332s
user 0m0.068s
sys 0m0.204s $ find -print0 | xargs -0 wc -c 2> /dev/null | tail -1
(Also, many uses of find | xargs can be replaced with -exec cmd {} \; or -exec cmd {} +, e.g. $ find -exec wc -c {} + 2> /dev/null | tail -1
although this isn't much faster in this case.)Also, the results are different - though I'm too lazy to figure out why right now :)
find -type f -ls|awk '{s += $7} END {print s}'
find -type f -print0 | xargs -0 wc -c | tail -1
find -type f -exec wc -c {} + | tail -1Any arguments specified on the command line are given to utility upon each invocation, followed by some number of the arguments read from the standard input of xargs. The utility is repeatedly executed until standard input is exhausted.
and
-P maxprocs Parallel mode: run at most maxprocs invocations of utility at once.
The way I interpret that is that you could run xargs in parallel mode, but by default the "utility is repeatedly executed" in the same process.
give me a break.
You can count scripts as commands (I did it in my other comment elsewhere on this page) but the way you do it you will miss a lot. For instance you won't count any "uniq", "sort", … that are almost exclusively used as filter and not in first command, you will also miss a lot of "less" and "grep" for instance.
sed 's/ *| */\n/g' ~/.bash_history | cut -f1 -d' ' | ...
(This fails on something like "echo 'a | b'" and doesn't split correctly on |& and || and doesn't split at all on &&.)