The Art of Command Line
github.com
github.com
Do not. Having a non-utf8 locale means you won't be able to handle utf-8 sanely ("that's why it's faster") and it will break at the most inexplicable times. Any non-latin1 character appearing in your prompt or command line with this will mess its spacing up for example. Do not do not do not.
Hell I even check for it in my .zshrc: https://github.com/jleclanche/dotfiles/blob/master/.zshrc#L3...
Good post otherwise.
You absolutely need to be aware of and manage language and locale settings diligently. This is a big topic, but generally, whenever possible, use UTF8. And you generally don't want to change your desktop OS or applications' locale or language settings other than that.
However, if you are processing data files and are aware of the implications, and can get away with treating data as binary instead of character streams -- for example, any of the command sequences that use sort/uniq for uniqueness or set union/intersection/difference -- then using C locale is way faster. As in, the difference between doing something in an hour on a big machine and rewriting a whole pipeline in Hadoop. A good fact to know.
It's obvious I (the author) didn't make the simultaneous truth of both these points clear. I'll rewrite.
It's not worth the "performance improvement".
$ echo $'a\naa\nb\nå'| LC_ALL=C sort
a
aa
b
å
$ echo $'a\naa\nb\nå'| LC_ALL=nn_NO.UTF-8 sort
a
b
å
aa
Even worse, some characters that are byte-different collate the same: $ echo $'∨\n∧'|LC_ALL=C sort -u
∧
∨
$ echo $'∨\n∧'|sort -u
∨
Handling Unicode sanely requires understanding some Unicode.In general, I'd never do it for performance only, it would have to involve one of the issues you've pointed out (there are some components I build that require it, erroneously IMO).
Knowing that "LC_ALL=C can improve performance but cause UTF problems" is awesome information to base a design on. A blanket statement like "never use LC_ALL=C" is not.
Again, not worth it -- if only for performance sake, generally speaking.
You can apply the same argument to any other obscure encoding scheme, not just UTF-8. The answer is really to pick the encoding which will suite the most usecases and won't be painful to use.
Fact of the matter is people have to get out of the habit of going "oh, that will never be relevant for me", because a) that's unlikely to be true for everybody (for example in a server environment you can't even write perfectly good North American names like Étienne let alone perfectly good older-than-the-country names like Þórr) b) it's unlikely to be true even for you as you will discover the first time you try to read a file written by somebody else c) it promotes intellectual laziness that makes later internationalization efforts difficult to impossible, because they're trying to convince people like you to care about stuff they find inconvenient but that is of vital importance to your customers.
I know of a man whose first name is Þórr who outright refuses to deal with any organization that will not let him write his name, and I have a hard time disagreeing with him.
I wish I had that luxury sometimes. Here's a pet peeve. I live in France, where people expect everyone to have a one-word family name (it seems). It's really common for people and especially computer systems to "correct" my name (which is of the form Me van der Somebody) to something like Me VANDERSOMEBODY or Me Van Der Somebody. A particular web page (my employer's, actually -- where I have to fill in my name for my "profile page") does the latter, and refuses to allow me to fix it. I have submitted bug reports, and just get WONTFIX back. :(
It's really rude, so I can imagine that for a person called þórr it's a nightmare everywhere except Iceland.
This is a horrible assumption. UTF8 can show up just about anywhere. 99% of the utilities mentioned in the article support UTF8, which means they will output UTF8 if they think that's a good idea. And that can show up at any point. Exotic filenames, error messages, fancy boxing or indicative characters in the output...
Don't expect Latin-1, ever. Respect specified encodings, and default to UTF8. Never, ever think you'll be fine with Latin-1 when dealing with an UTF-8 output just because you speak english.
The trick the original article specified is a hack to get better performance in very controlled situations, it should never be put in a rc file like it recommends because it will break and you won't know that is what's causing it.
Edit: Here's a few examples of very real potential breakage, to show you how bad an idea this is: You're debugging a python script. You open a python shell and start pasting some of its lines to figure out what's happening. OOPS, there's UTF-8 in there.
[4:08:48] adys@azura ~ % python -c 'print("I like the letter Š")'
I like the letter Š
[4:08:54] adys@azura ~ % LANG=C python -c 'print("I like the letter Š")'
Unable to decode the command from the command line:
UnicodeEncodeError: 'utf-8' codec can't encode character '\udcc5' in position 25: surrogates not allowed
(The input doesn't even have to be explicit, by the way, it could even be in your filenames).Should I keep going? See how "awk '{print tolower($0)}'" behaves when it's dealing with utf-8 input. Maybe that awk is deeply hidden in a 5000 line script. Maybe it deals with web data, filenames, anything. And suddenly you have unicode escapes in data that doesn't expect any.
Now how's that one line in your rc file, which you wrote maybe months or years ago, working out for you? Are you going to know the issue is your LC_ALL variable?
Second, you keep treating UTF-8 as some special encoding (ein OS, ein encoding?), but it's certainly not. Even if you are using UTF-8 as a system encoding, you are going to have problems dealing with non-UTF8 data, and it's actually going to be worse than dealing with UTF-8 when using LATIN1. Mind you, most systems don't use UTF-8 as default encoding: most Linux distributions do, but Windows uses UTF-16, Mac OS X has it's own somewhat incompatible version of UTF-8, Java is UTF-16, and so on. One will have problems dealing with all these systems if his encoding is set to UTF-8.
Lastly, the Python example is not really representative. Some language implementations have better support for multi-byte encodings, and some worse. Ruby, for example, has a very good support for dealing with UTF-8 data even when not using it a system encoding.
How does he fly? All flights I've ever been on have required me to ASCII-fy my name.
setenv LC_ALL en_US.UTF-8
setenv LC_COLLATE C # use the ASCII sort order
What have I broken, and how badly?And welcome to MITM haven. This is awful advice if you care about the first S in SSH.
On your home box, sure. But let's not pretend there's not a reason for the option to exist.
alias ssht='ssh -o UserKnownHostsFile=/dev/null -o StrictHostKeyChecking=no'
ssht floaty.vm # does not use host key checking
ssh my.bastion.host # validates host key
Really I ought to get SSHFP records populated when my vm's are created...
https://www.digitalocean.com/community/tutorials/how-to-crea...
There are two other modes to use. Either DNSSEC with SSHFP fingerprints, or sign your host keys with a CA key that you install on all your machines.
A decent configuration management tool could also keep known_hosts up to date, but it tends not to be worth the trouble. Do the above first.
set -o vi
Then you have vi keys in your shell. And it is marvelous.
Instead of that, create the following file:
~/.inputrc (tilde slash dot-inputrc)
And in that file put at least this:
set editing-mode vi
Now every program that you run that a) has its own command line, and b) uses the readline library (there are a lot) has vi commandline editing. psql, mysql, telnet, lftp, and many more.
man bash has a section on READLINE. For example, in man bash search for
editing-mode
which goes against his childish comment on emacs:
> Learn Vim (vi). There's really no competition for random Linux editing (even if you use Emacs, a big IDE, or a modern hipster editor most of the time).
Also, for more readline maneuvers: http://unix.stackexchange.com/questions/21788/how-to-delete-...
The way to think about it is that the prompt starts out in `vi insert mode` so whatever we type in will just be inserted at the prompt. However, we can hit the`Escape` key and move into `vi normal mode` where the keys are interpreted as normal `vi` commands. So, for example, once we hit `esc` we can navigate around the line using `h` and `l` and navigate history (previous lines) using `k` and `j`.
Commands I find very useful are:
`jk` : for history navigation
`bw`: for moving in the current line. These are word movement commands that are faster to navigate with.
`cw` : similarly, changing a particular word in the current line is also great. This works particularly well with history - retrieve a previous command and change a word or option
`dw` : for deleting the current word
`/` : is a fantastic command for searching history. `/cd` will bring up our last change directory command and just hitting `n` will cycle through all our change directory commands that are in history.
`AI` : to insert at the beginning/end of line
...and so on. If you have `vi` muscle memory it's a fantastic option for working on the shell.
(edit: Formatting)
v : to open command in a vi window. This is helpful for editing multi-line scripts/iterative composition.
I don't see this to be the case at all. Plus, having this be the opening line may cause people to equate archaic=I don't need this.
I'd suggest (would pull request but afk) that you remove "is a skill that is in some ways archaic".
Sure, GUI tools have crept in here and there. And I understand that some toolchains mandate all-singing, all-dancing IDEs. I don't work in those areas (the closest I've come was doing Java work for about five years and I still lived in vi, because Eclipse makes me homicidal). And I still see our iOS developers using `find' and `sed'.
Which isn't to say GUIs can't improve on the command line for some things. For example, the MacOS tool "A Better Finder Rename" (despite its inability to rename itself something less clunky) is a great tool for mass-renaming, e.g., photos. I have written such critters before as shell/perl/python scripts, but pulling EXIF data and renaming files programmatically is ticky and potentially destructive enough that I tend to write it once and then not modify it unless it breaks. The GUI preview and regex capture display makes this much less annoying.
"Command Line?! HAHAHAH"
Every time you have an issue with the shell, go there.
https://github.com/jeroenjanssens/data-science-at-the-comman...
The reason: http://arstechnica.com/information-technology/2015/05/source...
[1] http://www.amazon.com/Unix-Power-Tools-Third-Edition/dp/0596...
Also really handy is Ctrl-Y, which pastes the last cut characters by Ctrl-W, Ctrl-U or Ctrl-K.
I've never really had to use a console editor for more than 10s of lines.
It's not as hard as it first seems to learn, either. Well worth the time spent, vi is worlds more powerful than nano.
So you're saying "why would you use X, when you could avoid that by using Y"?
Well, because people prefer X, for a variety of reasons. Not to mention vi is installed everywhere, while sublime and atom are not.
There are other CLI editors. I'm quite partial to ne (the Nice Editor), which has more manageable shortcuts than emacs, a command line (it's still a non-modal editor, to be clear), macros and syntax highlighting. Also jed is not bad. While vim (and nvim) have advantages, and certainly benefit from being ubiquitous and having ther shortcuts replicated in other programs, users should shop around, even on the command line.
set -o emacs
Will give command line editing bindings compatible with Emacs (which is also the default editing keys in MS-Windows and many IDE's BTW).
HTH
http://ss64.com/nt/syntax-keyboard.html
Also, the original comment was on "console editors" and not on "command line editing", which is what I thought the subject was about. Looks like I struck out on the whole comment.
"Use vi, nothing beats it." == "I use x and I like it just fine."
If that was your intention, my mistake. It's just not apparent.
By whom? (honest question: even the most GUI oriented people I know reckon the power of command line skills)
[1] https://github.com/draios/sysdig/wiki/Csysdig%20Overview
I normally end up opening another terminal and using man there.
I think this will glob unless you escape the asterisks. Also on my system (debian 8) I need to put the directory to search first, or not at all:
find . -iname \something\
find -iname \something\
edit: hackernews ate all my asterisks
The majority of times, this is the case. At least one exception is when the shell pattern does not match one or more files in the directory which find(1) is executed. However, this can be "surprising" so it's best practice to either use escapes or ensure interpolation is disabled via single quotes (in Bourne shell syntax).
> Also on my system (debian 8) I need to put the directory to search first...
FWIW, you can specify more than one directory if desired. The primaries will be applied to the results of each.
Thankfully zsh has the NULLGLOB option which will do the right thing and expand to an empty list, of the zero matching files.
find . -iname "*file*"
zsh will cry that nothing matches the glob if you don't. bash will expand it if something matches, or pass it as-is to the client if something does - which will give you unexpected results. edit: hackernews ate all my asterisks
Well that solved the mystery. cat hosts | xargs -I{} ssh root@{} hostname
glad to know this one :-) echo y | xargs -Ix echo x
how tricky!As the article mentions, especially useful when combined with the -P flag for paralellism.
cat urlpaths.txt | xargs -I{} -P wget domain.com/{}
Another handy trick, is to use `sh -c` or `bash -c` to include multiple commands with xargs. cat dirs.txt | xargs -I{} sh -c 'cd {}; touch file;'# print 10 values from 1 to 10, with leading zeros:
for n in `jot -w "%04d" 10 1 10`; do echo $n; done
for n in {0001..0010}; do echo ${n}; done
# generate 10 random numbers between 1 and 1 million
jot -r 10 1 1000000
For us Linuxers: shuf -i 1-1000000 -n 10
touchegg & disown; exit function launch {
type $1 >/dev/null || { print "$1 not found" && return 1 }
$@ &>/dev/null &|
}
alias launch="launch "Thanks for this!
If you want an "installed everywhere" editor, learn ed. If you're willing to take an editor with you (or otherwise make sure a specific editor is everywhere you are), there's no reason it has to be vi.
Plenty of those didn't have emacs, especially on first boot, when you're most likely to need to modify network configuration files in order to get online to install new packages in the first place...
Edit: By "vi" I mean that somewhere in the path there is a program called vi. Of course, usually this is nvi, vim, evil, etc; I don't mean the original UNIX vi binary.
It's often/always part of a full install, but it isn't always there when you end up in rescue mode and/or situations where /usr won't mount.
However, ed is available, and vi is closer to ed than emacs is. And emacs won't be available either.
Default Ubuntu install.
obCompulsoryEdJoke: https://www.gnu.org/fun/jokes/ed-msg.html
I learned ed for a joke and found myself actually liking some features so much that I started porting them to emacs.
.
w
"Of course, we are not absolutely compelled to use ed, since V7/x86 comes with the full-screen editor vi (as well as the related line editor ex), but getting access to vi in this context would involve mounting the /usr filesystem and changing various environment variables and settings to allow editing in full-screen mode. Some basic ed skills are particularly useful in the V7 world, and most of all when one is involved on system administration tasks." [2]