1. df -h; notice that it hangs.
2. Log in to a new shell, 'strace' the old process or a new one doing the same thing, see what path it was choking on.
3. If the breakage is on an external/network filesystem, reboot the host in almost every case. Unless it was happily completing day 364/365 of some incredibly important task elsewhere, it's just not worth my time to remount a dead share and go clean up everything that was broken trying to talk to the old one. I've had database servers lose some random NFS share that the DB process wasn't using, then crash months later due to PID exhaustion because some monitoring script in cron that kept trying to talk to a somehow-corrupted mountpoint and hanging forever. Yes, timeouts and client programs should be able to handle these failures perfectly in theory. Given my experience, I have very little faith in theory matching up with reality.
4. If it's on an internal drive, check dmesg/syslog (if I can) for any smoking guns. Reboot and see if the problem goes away. If it does, unless I can find something blindingly obvious indicating that the issue was transient and unlikely to reoccur, I'm probably reprovisioning the system after a hardware diagnostic. Even if the server isn't critical and just serves a cat blog or whatever, it's not worth my time and repeated head-scratching to deal with issues like this more than once per host.
5. If I need data off of the questionable filesystem, I'll get it exclusively via a recovery environment; not worth the risk in the case of a flaky/failing drive otherwise (this applies even if the server itself was virtualized). I hope the server had some sort of LOM console set up so I can do that, otherwise someone's getting travel expenses for my trip onsite.
Edits: grammar.
TBH I get rather claustrophobic when I can't `w` with aplomb.
df -h &Downside is you need to log in with a new shell and tread lightly because it's easy for anything to get stuck in that state. Checking the syslog for NFS errors is a good place to start, or inspecting the fstab to see what is supposed to be mounted.
thereyago.
Their open source NCPA (Nagios Cross Platform agent) plugin for Windows and Linux hosts is pretty great, though we were stung when we found out it didn't natively support SPARC as of yet unless you compiled your own client from the source (I did). [2]
Plenty of other monitoring services equivalent to Nagios offer equivalent services. Nagios was just the flavour I suggested to my current employer because of my familiarity with the product over my last few employments.
If you don't want to fork out the licensing for the supported version (XI), their open source version is free and relatively easy to deploy using Ansible if you don't mind writing in PHP. [3]
Edit: words, grammar.
[1] https://www.nagios.com/products/nagios-xi/
[2] https://github.com/NagiosEnterprises/ncpa
[3] https://hobo.house/2016/06/24/automate-nagios-deployment-wit...
Deploying Zabbix roughly consists of setting up one (or more) Zabbix Master servers, and installing a 'zabbix-agent' process on each device you want to monitor. The agent process extracts a variety of statistics about the system and makes them available to the Monitoring server, either in push ("active") or pull ("passive") fashion.
The Master logs all of these statistics over time. You can then define 'Triggers' that apply logical tests to statistics, such as "in the last 5 minutes, was the free diskspace < 10GB". When this happens, it triggers an event.
Then elsewhere you have defined your Notification rules that act upon generated events, to send out an email or text message. https://www.zabbix.com/features#notification
This was obviously really simplified, and there is so much more you can do, but I hope this at least gave you a basic picture.
# node_exporter on the client to get metrics # prometheus on server to pull the metrics from node_exporter # alertmanager for raising alerts # grafana on top of it for pretty graphs(optional)
Still dont know what lesson I should learn from this story.
-m reserved-blocks-percentage Set the percentage of the filesystem which may only be allocated by privileged processes. Reserving some number of filesystem blocks for use by privi‐ leged processes is done to avoid filesystem fragmentation, and to allow system daemons, such as syslogd(8), to continue to function correctly after non- privileged processes are prevented from writing to the filesystem. Normally, the default percentage of reserved blocks is 5%.
-g group Set the group which can use the reserved filesystem blocks. The group parameter can be a numerical gid or a group name. If a group name is given, it is converted to a numerical gid before it is stored in the superblock.
> when a process running as root runs df (or any similar command or syscall), then it sees that capacity as available.
I have never seen that. I have only ever seen the “real” available space being shown by tune2fs, not df or anything similar, which have always shown the space available after subtracting the reserved space.
Most of the time some database connection is laggy
Writes to replayable binary logfile with 10 minute system-state snapshots and uses process accounting to ensure it see's every process during that time.
Gives more metrics than any other top; including network, disk and all the counters that you have to check the man page to know what they refer to.
Its counterpart atopsar lets you replay the data for specific stats in an easily viewable format; i.e; - atopsar -m - this shows the memory stats for todays logfile in 10 minute increments.
It goes on every single server I manage without exception. With atop you can actually see why it died instead of guessing from old log entries.
Screenshot of atop : https://www.atoptool.nl/images/screenshots/genericw.png
Example output of atopsar:
# atopsar -m
*snipped* 2.6.32-896.16.1.lve1.4.51.el6.x86_64 #1 SMP Wed Jan 17 13:19:23 EST 2018 x86_64 2018/03/27
-------------------------- analysis date: 2018/03/27 --------------------------
00:00:02 memtotal memfree buffers cached dirty slabmem swptotal swpfree _mem_
00:10:02 14027M 928M 839M 6335M 1M 4645M 0M 0M
00:20:02 14027M 963M 842M 6393M 1M 4641M 0M 0M
00:30:02 14027M 756M 873M 6617M 1M 4638M 0M 0M
00:40:02 14027M 576M 871M 6596M 3M 4634M 0M 0M
https://www.atoptool.nl/index.phpPersonally I consider two processes and 40MB of ram to be negligible for the benefits it brings.
You can indeed use it as a standalone top without either of these processes too. You're just giving up one of the main benefits (replayable logs) outside of the extra stats.
In short; you're moaning about what exactly?