Devops/Sysadmin Cheatsheet
rubytune.com
rubytune.com
- Instead of 'while true;' you can use the shorter 'while :;'. ':' is a null command.
- On OS X, dtruss is kinda, sorta like strace (it's a wrapper around dtrace, and the dtruss name comes from Solaris).
- For basic host to IP resolution, I prefer ping as it calls gethostbyname(), like most other programs will do. host/dig are suitable for querying/testing DNS independent of how the local box is configured as they call the resolver library directly. For example, host bypasses nsswitch.conf.
(Separately, OS X's stub resolver has some neat tricks. e.g., using /etc/resolver/<domain_name> to route DNS requests for a particular domain to a specific name server.)
- I had never heard of http://michael.toren.net/code/tcptraceroute/ before. I don't think I've ever encountered a situation where the issue was outbound ICMP/UDP packets being blocked, but rather it's the return ICMP Time Exceeded packets.
- 'find . -size +100M' is shorter than 'find ./ -size +100000000c -print'.
- ctime is NOT when a file was created, though this is a common misconception. Unix does not record when a file is created. Rather, ctime is the last time the file's status (i.e. inode information aka metadata) was changed. This differs from mtime which records the last time the file's data was changed. ctime is a superset of mtime. atime is the last time a file was accessed, unless the filesystem is mounted with noatime (common for NFS file ssystems). atime/mtime can be set to arbitrary values (permissions allowing) via the utime() system call.
- Modern find supports -delete instead of '-exec rm {} \;'. In any case, my muscle memory still defaults to '-print0 | xargs -0 rm'.
- 'dd if=/dev/zero of=file.txt count=1024 bs=102' seems like an odd way to do that. I think that 'bs=1024 count=102' might be more efficient depending upon buffering?
traceroute -P udp <hostname>
This is super useful if you're probing a network that's hostile to ICMP packets - it's a lot more reliable than ICMP in many scenarios, from my experience.
I was about to give up and install dnsmasq locally to get around a VPN-related issue. Thanks for the tip!
I just wanted to note that some Unix systems really store the creation time. See http://www.daemon-systems.org/man/fstat.2.html for an example.
Quick mysql "why is my database under so much load suddenly" that doesn't rely on the slow query log:
pt-query-digest --processlist h=<host> --interval=0.01 --print > /tmp/queries.out
pt-query-digest /tmp/queries.out
Tracking down what files exactly are filling up the disk. Go to /, and follow the big numbers: du -smc *
Combine strace (for the file handle no), lsof (for the file handle to ip/port) and netstat (to find the process on the remote host) to track RPC calls that are hung.For 'ps auxww', throw on an f to get a visual representation of child processes.
And finally, ssh tunnels (which can be chained together) to get past bastion servers easily:
Host <alias>
HostName <ip>
ProxyCommand ssh -q <bastion host, can be another alias in .ssh/config> "nc -w 3600 %h %p"The pt-query-digest function will cause additional load on the DB, but it won't interrupt communications.
If tcpdump can't keep up with the incoming traffic, the kernel will drop the packets in its buffer (or rather overwrite them with new packets).
Throw in TCP's flow and congestion control protocols, and dropped packets can have disastrous effects on your database.
Google has many references you might find useful on this subject.
Host <alias>
HostName <ip>
ProxyCommand ssh -W %h:%p <bastion host>
Which doesn't require nc (or anything else) installed on the bastion host.Unlike fundamentals (think books like SICP, Introduction to Algorithms, K&R, the Dragon book, et al), this is just a collection of useful commands that do not bring barely any collateral learning aside from learning the 'UNIX' way.
If you really want to put up the effort to learn this, and you don't have any projects or anything that requires this knowledge, I've noticed lots of tech offices have a copy of UNIX in a nutshell always have a copy around. I've checked it out myself and it's pretty useful (I already 'know'), not sure how good is it for learning this stuff from the ground up.
Also, talking with/demanding explanations/debating with geekier than thou friends has been really helpful with the bigger concepts.
I had the privilege of working with a badass russian unix hacker for some time, he taught me how to do black magic with find, grep, sort and uniq.
I have the feeling ops people can be somehow.. "measured".. by how many things they had exploding in their face, and then learned how to fix them.
For example, looking at TFA, the first thing I thought is "oh, `df -i` should be there too".
Because I ran out of inodes enough times that I learned to check that (though I am not a sysadmin).
> ps aux | head -1 && ps aux | sort
Is a wasteful construct and doesn't even do what is claimed "List the top 10 memory hogs". Depending on what you consider "memory" something like this is shorter and correct.
ps aux --sort=-resident|head -11
Other errors: They list same command "du -hs" for twice. I believe they meant "df -h" for "overview of all disks". Although, it's correctly overview of mounted file systems.
Deployment automation, monitors, backups, cleanup...
Wikipedia: http://en.wikipedia.org/wiki/Devops
In small companies it usually always exists by accident. I think the "hype" is more around large companies that have huge barriers and sometimes friction between development and IT or sustaining Operations.
Ben Rockwood from Joyent has a good conceptual talk at http://www.youtube.com/watch?v=h5E--QSBVBY , "DevOps Demysitified".
But yes, my impression is that it's normally a developer who "trends towards doing systems administration" – so someone who knows the codebase, but can provision and wrangle servers. So, someone like me (I write full stack apps, but my "speciality" is the server end of things).
Because, atleast in my understanding, it's totally in the working area of sysadmins to set up a server, install and configure software, configure monitoring, etc. The typical sysadmin will also write scripts and stuff to do that.
My bet is that it just came up by some angry sysadmins who felt like they need to differentiate from the dumber kind of admin who can barely touch a shell.
Of course a software developer may also write deployment scripts and similar stuff, so that's where it becomes fuzzy.
I never really thought about that term, but thanks to yyour question i just figured that i will drop the term from my vocabulary. It's too fuzzy, it's too much of a buzzword.
Things like "sudo !!" are INCREDIBLY dangerous, and I would never put that on a cheat sheet.
Obviously, this doesn't help when using Fabric or scripting, but you shouldn't need "sudo !!" in those scenarios.
Or put in a slightly nicer way - Blindly running commands is going to turn into a learning experience, don't do it on a production machine.
If you have lots of jitter and lag, i suggest using Mosh which I've found to be amazingly awesome when using it on Amtrak's spotty free wifi (3G uplink).
$ rm -rf /root
rm: cannot remove `root': Permission denied
$ sudo !! # Does not run this, but returns the line below.
$ sudo rm -rf /root # Then you say, "Wait I dont want to do this." $ rm -rf /root
rm: cannot remove `root': Permission denied
$ sudo !!
sudo rm -rf /root
# /root is now kaputIf and only if you meant to do that, you're fine.
Hopefully if you're in a NOPASSWD user/group, you know what root can do if you're careless; if you're not, there's yet another point where you'd have go out of your way to do damage with this shorthand.
user@host$ less /var/log/syslog
Permission Denied
crap... $ sudo !!
sudo less /var/log/syslog
Woo!Though...the only way it's really "rails focused" is that the commands are collected over time and sourced from my fellow rails developer/sysadmins. A lot of these are basic/general linux commands.
Note: not affialated with them; just one of the top Google results. Also, I'd take a screenshot (and then export to pdf if an image file wasn't acceptable) since the pdf conversion via pdfcrowd splits it into multiple pages.
host1$ while : ; do nc -l 6666 > /dev/null; done
host2$ pv /dev/zero | nc host1 6666
156MiB 0:00:17 [9.46MiB/s] [ <=> ]Will also correctly do a PTR query for the reverse entry for that IP. Also functions correctly with IPv6, so you don't need to remember to split the IPv6 address up into a lot of dots :P
And I concur with the comments about "sudo !!" - !! in general I've totally removed from my UNIX vocabulary. At least do something like !?string[?]
wget cachefly.cachefly.net/100mb.test -O /dev/null
check the write speed of the disk. Mostly used to check what kind of a write speed you get on the machine. dd if=/dev/zero of=iotest bs=64k count=16k conv=fdatasync && rm -rf iotestMost of the commands listed rely on various GNU extensions to the various utilities or are only applicable to Linux.
FreeBSD:
strace -> dtrace/ktrace/ lsof -> sockstat/fstat watch -> no idea
sudo -> Almost never installed by default on FreeBSD ps aux --sort=-resident|head -11 -> --sort is not valid ...
And the list goes on ...
But, yeah, noone’s stopping you to make your FreeBSD version of the cheatsheet, what with the open source spirit and all… :P
ps auxww -> ps auxww -H (H is hierarchy)
Faster than lsof, and only displays files - although most things in unix are files!
lsof -p -> ls -l /proc/$PID/fdRun something forever watch command
Overview of all disks du should be df.
Find files over 100mb find . -size +100M
Low hanging fruit for size ls -al | sort -nk5
Files created (modified) within the past 7 days: find . -mtime -7
Find files older than 14 days: find .gz -mtime +14 -type f This will break when you have more archived files in the directory than the shell's glob char can support. Use: find . -mtime +14 -type f -name '*.gz' and it will run quicker too.
TCP Sockets in use, "netstat -antp" will be faster and also lists the process id.
EDIT: formatting
ls -al | sort -nk5
I suggest: ls -larSEither way, have an upvote!
df -h
not du -shE.g. man find
-atime n
File was last accessed n*24 hours ago. When find figures out how many
24-hour periods ago the file was last accessed, any fractional part is
ignored, so to match -atime +1, a file has to have been accessed at least
two days ago.Small nitpick/question: why put the commands in inputs? Do you edit them on that page? Or is it just to format them?
I tried to do a copy/paste on the page contents to my personal offline notebook but the most important bits, the commands, didn't paste.
The plan is to have a pdf shortly; that will most definitely be copy and pasteable.
du -h --max-depth=1 -x
I use this frequently to keep an eye on disk used by user directories. It returns the disk space used by each directory in the current directory. du -sh *Problem solved.
Honestly, when I have to get fine grained enough to dig into disk usage using 'du', I'm rarely concerned with hidden directories (and I typically append a 'c' to the command as well, so if I do have to pay attention to . directories, I will notice the discrepancy between reported sizes).
What I find really useful is you can immediately type some characters to filter, and then tab to the command you want. It's automatically selected in that case.
I'll likely stick a "copy" button on hover shortly — had some initial problems with z-indexes (it uses flash) but this solves the "one click to copy" issue without affecting the default and expected interaction.