First 5 Minutes Troubleshooting A Server
devo.ps
devo.ps
https://github.com/BashtonLtd/whatswrong
The idea isn't too tell you the problem exactly, but more to stop you missing things that are obviously wrong.
I have the recording capability now (and, more importantly, playback too), but the constant influx of broken boxes is gone. Funny how that works.
Playback: a fork of term.js derived from the jslinux terminal with some minor adjustments, plus a wrapper of my own which plays back the data with timing intact. http://rachelbythebay.com/w/2013/03/04/jvt/
I've been using them to demonstrate how I go about writing certain bits of code, with the goal of eventually showing the creation of something larger.
Particularly useful when you're accessing via a serial console, as you'll get full boot from BIOS. You can both log and troubleshoot based on output.
Playback utilities can take a multiplier to speed (or slow) playback from realtime.
http://www.fduran.com/blog/quick-linux-server-review-for-mor...
User experience reports are nice, but rarely indicate something other than "load is high", or "server is unresponsive". vmstat and top not only can tell you that, they can start telling you why and where to look for your problem.
I'm pretty sure there is also a whole bunch of other (less famous) commands that would and should be added.
In this sort of environment you want to look at your pam logs (/var/log/auth.log on debian-like distros, and iirc, /var/log/secure on RH-like). This gives you not only a list of what was run, but who ran it and when.
I've seen inexperienced sysadmins remotely rebooting a server only to discover that it won't boot up anymore because the original problem was that the server had run out of disk space.
Same idea for services restart; don't do it unless absolutely needed. While it may be doing the trick in some cases it can also generate its own set of new false symptoms.
Take the example of :
- mysql is "slow"
- let's restart ! ...
- mysql init script has been puking dots (.) for the last 30 minutes on shutdown
- let's kill -9 it ! ...
- mysql db is corrupted ! Hurray!
Shame for me ...
http://imgur.com/CBTrdJw [server linux box] http://imgur.com/tEjBtNY [development os x box]
In medicine, it's commonly known that the interview with the patient (the 'history') is the first thing a doctor should be doing. Not just because it establishes a relationship with the patient, but because the diagnosis of most illnesses is guided primarily by the history [1] - even with modern MRI machines and DNA amplification techniques! At the very least the chat with the client provides context for the problem that you are investigating - you are now putting flesh on a skeleton of meaning rather than trying to create it on your own.
This article stresses the importance of first getting a verbal 'history' from the client - what the problem is, characteristics of the problem, time-course of the problem and co-incidence with other events (like software upgrades). There is also a parallel to medicine in that in this field a skilled practitioner may be able to diagnose the problem based solely on the history alone [2].
The second thing I noticed was the fault-finding mindset. As a medical student halfway through his second year of hospital placements this is something I took some time to learn. The initial approach to finding the reason for a problem is usually to (1)think of a possible reason for the problem, (2)try to fix that reason, and (3)if that doesn't work, goto 1. While this is a good because it shows you are actually thinking about the cause of the problem rather than its effects, it's not the most efficient way of going about things. One way doctors can narrow down problems is by restricting them to systems such as the cardiovascular system or the neurological system. A searing pain in your chest is more likely to be due to a problem with your heart or lungs than due to a problem with your kidneys or gonads.
This article takes exactly the same view of servers, classifying the individual hardware and software components that make up the vast majority of (linux) servers in the wild.
I don't fiddle around with servers much any more, but I'm bookmarking this page because it is such a useful illustration of a fault-finding mentality.
[1] http://archinte.jamanetwork.com/article.aspx?articleid=11058... [2] http://blogs.msdn.com/b/oldnewthing/archive/2012/08/29/10344...
Usually this comes as > ps -ef | grep something | grep -v grep | grep -v $myownpid
why not use one simple and concise awk statement which does it all in one go?
awk -F: '{print$1}' /etc/passwd ps -ef | awk '/[h]ttpd/{print$2}'
But apart from that: very nice summary of things to consider and the sequence for analysis.
IMO is easier to type the "grep/grep -v" statements than the awk ones. Usually we are more used to use grep and we add pipes and filters until we get the expected result.
I'm using zsh and thanks to the Global Aliases I can do, though: $ ps -ef G something GV grep GV ownpid
sysstat ('sar') reporting can also provide some of that much-needed history. Sar output is pretty readily visualized with utilities such as gnuplot.
<meta name='viewport' content='width=device-width, initial-scale=1, maximum-scale=1'/>
other than that, I learned some new tools - ss is pretty awesome!
Typically they're better when you can put them in context: you should run all these regularly on fully working servers which you know are operating normally so you have "something" you can compare your results to when the shit hits the fan.