I usually run 'w' first when troubleshooting unknown machines
rachelbythebay.com
rachelbythebay.com
One hour long component of a (now deprecated) SRE interview loop was for the candidate to SSH into a series of EC2 instances and debug issues which got progressively harder as the interview wore on.
I had wasted a substantial amount of time on the first and easiest problem by really overthinking it, and not trying the simplest of debugging techniques first. By the time I got to the final, hardest problem, I had just over 5 minutes remaining. The interviewer gave me a pass to skip it, but I was having fun, and really wanted to take a crack at it.
The final problem was to try to figure out why logging into a particular machine with SSH was slow. While I sat waiting for a prompt, I had a number of thoughts. Is a reverse DNS lookup timing out? Is there a huge i/o load on the machine? Am I going to have to wire up strace to `sshd` and log in again?
When I finally get to a shell prompt, I instinctively run `w` and it just hangs. I hit ^C, `strace` it, and discover that it's blocking on:
fcntl64(5, F_SETLKW, {type=F_RDLCK, whence=SEEK_SET, start=0, len=0}) ...
I look up a bit, and discover that file descriptor 5 is /var/run/utmp. So `w` is trying to get an advisory lock on utmp and failing. Then it hits me, `sshd` is likely also trying to acquire a lock on utmp, failing, and then eventually timing out.A little bit later, I've found and killed the the rogue program that held the lock, and SSH logins were fast again. Solving that last problem so quickly really boosted my spirits, and gave me the energy to push through the harder interviews that came later in the day.
Thanks w!
uptime # uptime and CPU stress
w # or better yet:last |head # who is/has been in
netstat -tlpn # find server role
df -h # out of disk space?
grep kill /var/log/messages # out of memory?
ps auxf # what's running
htop # stressed? , look out for D (waiting on I/O typically) processes
history # what has changed recently
tail /var/log/application.log # anything interesting logged?
df -hi
The post argues that 'last' will not always give the desired information:
> Obviously, I could also use 'last' to see who's been on the box recently, but this isn't the whole story. It's totally possible to "ssh root@box /path/to/command" and never start a login shell, which then leaves no trace in the lastlog, but then goes on to break something on the box. The syslog is how you'd find this.
We have this and a few other elastic beats on most of our vms in amazon baked into the amis we use. So anything deployed by us starts sending lots of data to our logging cluster for metrics, auditing, internal application events, stacktraces from our docker infrastructure, syslogs, etc.
And I still run w first thing almost every time out of habit.
What am I looking for? A session I accidentally left open somewhere else? Unauthorized access? A friend? Dunno...
But in my examples, I never use the "hit-by-a-bus". Sysadmins tend to frown upon that comment.
I use the "win-the-lottery-go-to-Fiji-and-never-look-back". It always makes them smile :)
Death is a real thing that really happens to people, and from an organizational perspective it is valuable to keep in mind that no one in your company is immune to that.
After a little soul-searching, and deciding I didn't want to harbor anger, I helped him with the ones I remembered each time.
Had I actually been hit by a bus, that wouldn't have been possible at all.
I'm sure that if I'd hit the lottery and gone to Fiji, I'd have been even more likely to help him with those passwords.
In the end, not-burning-that-bridge did help me earn more money as he hired me back several times, and I demanded more money each time until I was asking almost as much per hour as he was getting from his customers and he simply couldn't afford me.
Aside from one of them being super annoying, over and over again, helping them out was the right thing to do.
It's not like it cost me anything other than a couple of minutes of time.
If someone already screwed you over, then... But otherwise, why not
https://en.wikipedia.org/wiki/Bus_factor
"The bus factor is a measurement of the risk resulting from information and capabilities not being shared among team members, from the phrase 'in case they get hit by a bus'. It is also known as the lottery factor, ..."
A good sysadmin has, or at least expresses traits of in their work, pessimism and realism. We have to constantly viscerally feel and know that any component can fail at any time. Remembering your own mortality goes along well with that.
Being easily disturbed by this is not a trait I would like to see in someone with these kinds of responsibilities.
I rather rely on beginner level sysop skills than on management understanding the implications of "this is my access to your mail server, don't handle with care, don't handle at all unless I am run over by a bus".
No. There aren't any ethics rules implicated in the way we do this.
No client files are stored on these servers.
echo ""
/usr/bin/netstat -f inet | /usr/bin/grep --color -C 100 "MYHOSTNAME.ssh"
echo ""
echo ""
w
echo ""
(note, the netstat command arguments are for FreeBSD ...)So, when I log in, I see all current network connections, and any current SSH connections are highlighted in red.
Then double carriage return and then the output of 'w'.
Another very handy command is
sudo netstat -atpn
which shows you the processes and owners of open TCP/UDP ports. The argument combination is as weird as "ps aux" that I just memorized it by heart. ss -atpn | column -t
Or pipe into `cat`Chan is a japanese name ending for children, various communities have created several Anime mascots over the years. Theyre generally suffixed with that -tan, as a cute misspronounciation of chan.
https://en.m.wikipedia.org/wiki/OS-tan
Btw, 4tan should be that 4chan mascot. Ive seen it previously on sankaku complex, but its been too many years ago. Cant find it right now.
/Edit: i probably should mention that you shouldnt visit sankaku complex at work. Its very... Questionable with nudity
-r everyone needs, but usually -R is a better default since it does not ignore symlinks
-h I practically never use since I'm almost always using grep to find what file I need to further investigate and letting the stuff down the | line deal with it
-o makes sense, kinda, I want to do that in the next stage of the pipe, maybe I have used it once? cant remember
but -w... I feel like I'm missing something there.
I think what you are doing here is search, but -o and -w shine at data extraction and data manipulation. Things you would use Excel for, except on multiple GB files.
If you are not using search, but are using a recursive (-r) grep, then you are feeding a bunch of data to your pipe but don't care what file it came from. I get that... I mirrored the SEC's ftp back when it was ftp... and it's interesting for correlation and name stuff... but that's not going to work because it's freeform, and it's millions of small files... so you are using (large) files with a known format? Something like spectral datasets where the fields are passed with the data prefixed with it's important tags? Maybe logs?
grep -whor '[A-Z][A-Za-z]*Exception' * | sort | uniq -c
for a histogram of exceptions inside logs downloaded from a bunch of nodes.As for -w, my favourite case is https://unix.stackexchange.com/questions/110645/select-lines...
As for -o,
grep -o ........-....-....-....-............
It may look like Morse code, but what it does is greps UUIDs out.My most common grep phrase is grep -iIRn - good for searching in code (when IDE can't find something). You can remove -i if you care about case, but I mostly used it for plsql code, and that's not case-sensitive.
lsof has a good case to join coreutils.
$ type ls
ls is aliased to `ls --color=auto'
On slow file systems (for instance, think of shared machines with large directories, NFS and heavy load) this can take ages to give output. Instead, if you just call /bin/ls, this call only calls http://man7.org/linux/man-pages/man3/readdir.3.html and thus gives output promptly.naw, that would be `netstat -lntup` (listening, nummeric, tcp, udp, programm)
your mentioned command shows essentially all currently active tcp network streams. https://explainshell.com/explain?cmd=sudo+netstat+-atpn
btw, glances is even more complete than htop
TCP, UDP, listening ports, user, PID and process name, extended detail.
What's the ultimate netstat?
:-)
htop is my trusted standby. Fell in love with netdata though. https://github.com/firehol/netdata - since we all do our best to stay away from getting shells on the servers, netdata is awesome
It seems to do something very similar to Cockpit, but netdata has way more stuff out of the box.
My experience of the last five years is so heavily weighted to (effectively) immutable infrastructure that checking to see who had been on a box hadn't event crossed my mind.
Ansible doesn't get you everywhere.
InfluxDB is a simple command to extract what I need and then work with it from there.
When I'm on-call and get an alarm on a server, checking that colleagues aren't logged in is normally the first thing you do. If a service fails it's more often than not someone doing maintenance, and they just forgot to tell you, or disable monitoring. And as someone else pointed out, also check disk usage.
Hacker News can be a little blind to the fact that most software projects are rather small and modern servers are really powerful. It's actually pretty rare that someone manages enough infrastructure for a single service that logging isn't a viable option.
It seems to me Logstash provides what I want, but I sure as hell won't run 3 JVMs and a Redis instance to aggregate logs.
tl;dr How to handle log aggregation within a small-scale cluster without losing your sanity?
In the past have created a server just with syslog. Adjust all your servers syslog configs to point to that one. Then log in and use grep.
Not fancy. But gets the job done with a single line config change in your syslog configs.
This re-enforces my idea that a purely terminal based business dashboard might be a cool product with a fairly large market. I've been eyeballing some of the go/ncurses work for this.
I have been running oklog for a customer successfully. You ought to deploy it like a regular (non-cloud-native) service though: When a server crashes, restore the oklog instance from backups, not spinning a new one.
I've worked with ELK before and I'm not a fan, the usability is very low compared to Splunk. It's much cheaper though. At a previous employer we switch from Splunk to ELK and the number of searches we did on a daily basis dropped to almost zero. Before that almost every support "ticket" would start with a Splunk search.
You could also try out https://www.humio.com. I only seen it deploy by one customer, but it looks nice.
See also https://lowendbox.com/ and https://hackernoon.com/10-things-i-learned-making-the-fastes...
http://h41379.www4.hpe.com/openvms/whitepapers/high_avail.ht...
Those few building blocks done in a highly-robust way on one of the stable, Linux platforms could probably achieve the same thing. Just RAID 0, a good filesystem, clustering itself, and backups would get far. I know there's Linux-oriented products out there but I don't know if their availability or failover time have caught up yet.
I put the bootstrap into an OS package and my server images point to my personal repo so I can configure run and build time dependencies as needed. At the end of the day, I can get a new server, run update, install my package and walk away and everything should be running, but it takes almost no additional overhead over doing it all by hand the first time.
I haven't seen human users in passwd, since virtualisation kicked in couple of years back.
I think there’s another level we need beyond treating servers as pets or cattle, which is treating servers as wild animals; after you’ve captured one and interacted with it, you’ve doomed it because now it has the scent of humans on it.
I also had a difficult time to explain that a prod deploy of 30min (image creation, deploy with blue green) is normal for this kind of inf... Did you face the same thing?
At the end of the day, if you're on continuous deployment, a commit should be rebuilding only what it touches. We have 4 min long deployments + 1.5 min tests and I definitely don't think we're optimizing aggressively.
Many places still keep producing very stateful software (sometimes even very much by choice) that is better off managed through Puppet / Chef rather than an immutable containerized approach. If your software needs to take an hour and a half to shutdown, for example, you have to get a bit creative with your deployment strategies.
But they aren't interested in investing in more infrastructure when a single LNMP server does the job.
For my startup, I often run bash on Heroku because I do migrations manually (again, it's a feature, I'm too inexperienced to have automated migrations that work everytime, I prefer to be already on it if it breaks). Sometime when something breaks I'll also poke around the filesystem (which is a copy, so no fear to break anything).
Basically, I'd say the smaller your team, your uptime requirements and your traffic is, the less you need automation, and the more you are susceptible to login directly to a box (I combine all of that: team of 1, no uptime requirement, not enough traffic to even max the most basic server).
(Also, when someone asks me "how do you configure X", I can just link them to the corresponding place in my system configuration repo on Github.)
The reasons for this are several, but the most pressing are that if one falls over I can spin it up again quickly somewhere else, and I can have a lot of commonality in the configurations between them, so they all have backups configured the same way, or setting up letsencrypt requires editing things in only one place.
It also means I don't have to dig through a bunch of config when I want to work out how stuff is set up three years later, I can just look in one version-controlled directory in a central location.
In web scale production, you mostly have to worry about your fellow sysadmins troubleshooting the problem, and they mostly won't be too mad about you clearing state.
But not everything is web scale production. Productionizing an app is a lot of work. Yes, yes, "Cloud is reliable!" and that's kinda true if you write your application to deal with any failure at any time. the reality is that the hardware can and will sometimes go away without warning, and you can't call up your oncall and tell them to head to the datacenter to fix it. all that recovery is done at the application layer; you had better hope you managed to make the data redundant enough. The whole idea behind the "cloud" is to take most of the lower level sysadmin work and make developers do it; and that's a fine way to some things, but... not so fine when it comes to other things.
(This is the value proposition of a VPS system rather than a "cloud" instance. when the hardware a VPS is on goes down, some poor bastard's pager goes off, and they are expected to wake up, drag themselves to the datacenter and fix it. The gamble is "is being down for a few hours now and again, when you can reasonably expect to be brought back up in the same configuration cheaper than writing your software such that you can always start over on a new node?" The VPS is for the former, the "Cloud" for the latter.)
Add to that, well, a lot of businesses need to run proprietary software. I personally, well, let us just say that for the last three years, the systems I am dealing with at work involve FlexLM.
so corp and smaller sites are littered with one-off systems... systems where if they break, a pager goes off, and someone fixes the problem. Sometimes that works better than the "web scale" stuff... like I've yet to see a "web scale" posix-ish filesystem that meets the user expectations the way NFS does, and I've never seen a nfs server that didn't require someone on pager.
* unique lines to file1
* unique lines to file2
* lines common to both files
You can pass it various options to suppress any of the three columns.I once had an interview at Facebook where one of the problems was easily solved by `comm`. I found it funny that the interviewer, or anyone who reviewed the interview question, had never heard of it. I was a good sport about it though, I ended up writing a janky Perl script that roughly implemented `comm` to solve the problem, which (modulo Perl) was what they wanted me to do.
If you specify that the files are way too big to fit in memory, well, that's a different story.
https://docs.oracle.com/cd/E19683-01/816-0211/6m6nc66m6/inde...
The Linux intro man pages are still pretty useless sadly.
I still do it instinctively on my own boxes.
Most of the time some database connection is laggy
Writes to replayable binary logfile with 10 minute system-state snapshots and uses process accounting to ensure it see's every process during that time.
Gives more metrics than any other top; including network, disk and all the counters that you have to check the man page to know what they refer to.
Its counterpart atopsar lets you replay the data for specific stats in an easily viewable format; i.e; - atopsar -m - this shows the memory stats for todays logfile in 10 minute increments.
It goes on every single server I manage without exception. With atop you can actually see why it died instead of guessing from old log entries.
Screenshot of atop : https://www.atoptool.nl/images/screenshots/genericw.png
Example output of atopsar:
# atopsar -m
*snipped* 2.6.32-896.16.1.lve1.4.51.el6.x86_64 #1 SMP Wed Jan 17 13:19:23 EST 2018 x86_64 2018/03/27
-------------------------- analysis date: 2018/03/27 --------------------------
00:00:02 memtotal memfree buffers cached dirty slabmem swptotal swpfree _mem_
00:10:02 14027M 928M 839M 6335M 1M 4645M 0M 0M
00:20:02 14027M 963M 842M 6393M 1M 4641M 0M 0M
00:30:02 14027M 756M 873M 6617M 1M 4638M 0M 0M
00:40:02 14027M 576M 871M 6596M 3M 4634M 0M 0M
https://www.atoptool.nl/index.phpPersonally I consider two processes and 40MB of ram to be negligible for the benefits it brings.
You can indeed use it as a standalone top without either of these processes too. You're just giving up one of the main benefits (replayable logs) outside of the extra stats.
In short; you're moaning about what exactly?
Their open source NCPA (Nagios Cross Platform agent) plugin for Windows and Linux hosts is pretty great, though we were stung when we found out it didn't natively support SPARC as of yet unless you compiled your own client from the source (I did). [2]
Plenty of other monitoring services equivalent to Nagios offer equivalent services. Nagios was just the flavour I suggested to my current employer because of my familiarity with the product over my last few employments.
If you don't want to fork out the licensing for the supported version (XI), their open source version is free and relatively easy to deploy using Ansible if you don't mind writing in PHP. [3]
Edit: words, grammar.
[1] https://www.nagios.com/products/nagios-xi/
[2] https://github.com/NagiosEnterprises/ncpa
[3] https://hobo.house/2016/06/24/automate-nagios-deployment-wit...
Deploying Zabbix roughly consists of setting up one (or more) Zabbix Master servers, and installing a 'zabbix-agent' process on each device you want to monitor. The agent process extracts a variety of statistics about the system and makes them available to the Monitoring server, either in push ("active") or pull ("passive") fashion.
The Master logs all of these statistics over time. You can then define 'Triggers' that apply logical tests to statistics, such as "in the last 5 minutes, was the free diskspace < 10GB". When this happens, it triggers an event.
Then elsewhere you have defined your Notification rules that act upon generated events, to send out an email or text message. https://www.zabbix.com/features#notification
This was obviously really simplified, and there is so much more you can do, but I hope this at least gave you a basic picture.
# node_exporter on the client to get metrics # prometheus on server to pull the metrics from node_exporter # alertmanager for raising alerts # grafana on top of it for pretty graphs(optional)
Still dont know what lesson I should learn from this story.
-m reserved-blocks-percentage Set the percentage of the filesystem which may only be allocated by privileged processes. Reserving some number of filesystem blocks for use by privi‐ leged processes is done to avoid filesystem fragmentation, and to allow system daemons, such as syslogd(8), to continue to function correctly after non- privileged processes are prevented from writing to the filesystem. Normally, the default percentage of reserved blocks is 5%.
-g group Set the group which can use the reserved filesystem blocks. The group parameter can be a numerical gid or a group name. If a group name is given, it is converted to a numerical gid before it is stored in the superblock.
> when a process running as root runs df (or any similar command or syscall), then it sees that capacity as available.
I have never seen that. I have only ever seen the “real” available space being shown by tune2fs, not df or anything similar, which have always shown the space available after subtracting the reserved space.
1. df -h; notice that it hangs.
2. Log in to a new shell, 'strace' the old process or a new one doing the same thing, see what path it was choking on.
3. If the breakage is on an external/network filesystem, reboot the host in almost every case. Unless it was happily completing day 364/365 of some incredibly important task elsewhere, it's just not worth my time to remount a dead share and go clean up everything that was broken trying to talk to the old one. I've had database servers lose some random NFS share that the DB process wasn't using, then crash months later due to PID exhaustion because some monitoring script in cron that kept trying to talk to a somehow-corrupted mountpoint and hanging forever. Yes, timeouts and client programs should be able to handle these failures perfectly in theory. Given my experience, I have very little faith in theory matching up with reality.
4. If it's on an internal drive, check dmesg/syslog (if I can) for any smoking guns. Reboot and see if the problem goes away. If it does, unless I can find something blindingly obvious indicating that the issue was transient and unlikely to reoccur, I'm probably reprovisioning the system after a hardware diagnostic. Even if the server isn't critical and just serves a cat blog or whatever, it's not worth my time and repeated head-scratching to deal with issues like this more than once per host.
5. If I need data off of the questionable filesystem, I'll get it exclusively via a recovery environment; not worth the risk in the case of a flaky/failing drive otherwise (this applies even if the server itself was virtualized). I hope the server had some sort of LOM console set up so I can do that, otherwise someone's getting travel expenses for my trip onsite.
Edits: grammar.
TBH I get rather claustrophobic when I can't `w` with aplomb.
df -h &Downside is you need to log in with a new shell and tread lightly because it's easy for anything to get stuck in that state. Checking the syslog for NFS errors is a good place to start, or inspecting the fstab to see what is supposed to be mounted.
thereyago.
Is the implication of "recently" that it should have been rebooted every couple of weeks or something?
My desktop is at 54 days (since I moved desk, I think) and picking a random Hadoop node:
> 10:26:47 up 303 days, 11:56, 1 user, load average: 44.12, 50.56, 47.48
It's private, so kernel updates aren't a security issue.
(Keeeping this on-topic, "history" tells me the most frequent thing I do on this node is "sudo iftop"; we've been doubting the accuracy of our monitoring system's network utilization graphs.)
usually there's a loud sob somewhere in between.
Ferrari does this. They do a lot of experimentation on their most expensive internal test equipment, and every now and then the whole box is re-imaged automatically, even if it's completely locked down from outside. It's their internal staff who is corrupting/improving the system. Only if it's a really good and well-tested improvement they will make it stick.
Sure there are tons of unknown commands on any OS, but a one-letter command you've never heard of somehow amplifies the amazement.
http://techblog.netflix.com/2015/11/linux-performance-analys...
Web servers for example should typically have minimal state for scalability and robustness but for most VPS setups this is not the case. If the VPS got wiped it could take days of work to get up and running again, and if you wanted to scale horizontally you'll have some reengineering to do. You could try and replicate something like Heroku yourself to minimise state on your web servers but it'll take you a lot of time and won't be as robust.
These are 'gstat' and 'top -m io' commands, both I/O related.
Often in top/htop there is not much happening but the server is very slow, a lot of the times its because of I/O operations.
The 'gstat' command will tell You (besides other useful statistics) how much the storage devices are loaded:
# gstat
dT: 1.001s w: 1.000s
L(q) ops/s r/s kBps ms/r w/s kBps ms/w %busy Name
1 9679 9614 4807 0.1 63 671 0.2 82.3| ada0
0 0 0 0 0.0 0 0 0.0 0.0| ada0p1
0 65 0 0 0.0 63 671 0.2 1.4| ada0p2
0 0 0 0 0.0 0 0 0.0 0.0| ada0p3
0 0 0 0 0.0 0 0 0.0 0.0| gpt/boot
0 65 0 0 0.0 63 671 0.2 1.5| gpt/sys
0 0 0 0 0.0 0 0 0.0 0.0| gpt/local
0 0 0 0 0.0 0 0 0.0 0.0| zvol/local/SWAP
The 'top -m io' show which processes does how much I/O: # top -m io -o total 10
last pid: 51154; load averages: 0.31, 0.31, 0.28 up 3+18:01:00 14:58:15
54 processes: 1 running, 53 sleeping
CPU: 2.4% user, 0.0% nice, 15.3% system, 5.3% interrupt, 77.1% idle
Mem: 345M Active, 1236M Inact, 153M Laundry, 2158M Wired, 3903M Free
ARC: 834M Total, 46M MFU, 295M MRU, 160K Anon, 5006K Header, 488M Other
67M Compressed, 274M Uncompressed, 4.08:1 Ratio
Swap: 4096M Total, 4096M Free
PID USERNAME VCSW IVCSW READ WRITE FAULT TOTAL PERCENT COMMAND
51021 vermaden 6005 16 6005 0 0 6005 99.92% dd
51154 vermaden 8 9 5 0 0 5 0.08% top
51036 vermaden 6 10 0 0 0 0 0.00% xterm
50907 vermaden 0 0 0 0 0 0 0.00% zsh
50905 vermaden 0 0 0 0 0 0 0.00% xterm
50815 vermaden 0 0 0 0 0 0 0.00% zsh
50813 vermaden 0 0 0 0 0 0 0.00% xterm
50755 vermaden 0 0 0 0 0 0 0.00% tail
41780 vermaden 0 0 0 0 0 0 0.00% leafpad
41255 vermaden 29 11 0 0 0 0 0.00% firefox
Another command that is very useful on FreeBSD is 'vmstat -i' which show how much interrupts are happening: # vmstat -i
interrupt total rate
irq1: atkbd0 75135 0
irq9: acpi0 2575929 8
irq12: psm0 135060 0
irq16: ehci0 1532065 5
irq23: ehci1 3265677 10
cpu0:timer 102772345 317
cpu1:timer 94199942 291
irq264: vgapci0 370466 1
irq265: em0 7017 0
irq266: hdac0 1904427 6
irq268: iwn0 148690342 459
irq270: sdhci_pci0 147 0
irq272: ahci0 20875039 64
Total 376403591 1161
I always suffer when I have to debug Linux systems because of lack of these commands.Regards, vermaden
While 'dstat' gives some information its usefulness (at least for me) is far from 'gstat' outout.
uptime
uname -a
w
who
df -h
free -m
cat /proc/cpuinfo
mount
top
ps aux
then check appropriate logs or dmesg etc
nethogs
nload -u K
- the author seems to call out specific individuals, by name, to their team leads, based on where they were ssh'd in, instead of bringing up the issue with that person first and asking for the logic of why they were on the box doing the thing (sounds like a lot of assumptions combined with finger pointing)
- it sounds like there's no well-configured monitoring or observability at the DC/rack/machine level involved, at all, which is surprising in a modern enterprise setup
I'm not sure why you would assume that. She specifically says "The next thing I'd do is to go get in touch with that person."
Perhaps you're keying off "I've been able to track down some well-meaning but ultimately flawed attempts at fixing things that then blew up and became something much bigger. The folks who I pinged about it were amazed that I somehow had managed to "guess" that a specific member of their team had been poking at a specific box"? But keep in mind, that's specifically events that became a large issue. Is it not appropriate to notify management as to the cause of the issue? Either it's a first time mistake or not something the person may necessarily have known to look out for, in which case management should be lenient, or it's the latest in a string of events and management should possible take some other action.
If nothing else, it allows management for that other team to say "hey, we don't need to be messing with this aspect of the server. Either contact the team whose responsibility it is and get them to do the work, or get them to sign off on it first."
They did and they figured it out. Outage resolved.
Why did you automatically assume I ratted them out to management? At no point does the story go there.
I’m really curious, since misunderstandings like this can really poison a working environment when people think you’re doing things you’re not. I want to know what sent you down the wrong path here.
Commenters tend to proclaim bad intentions when none is present when they either skimmed without reading or that they are lashing out to compensate for some weird insecurity, e.g. "I caused an outage once and I didn't want anyone to ask me and find out! How dare you want to know!?"
When all you do is wander around looking for broken stuff to help fix, imagine the above sequence repeating itself.
> "I've been able to track down some well-meaning but ultimately flawed attempts at fixing things that then blew up and became something much bigger. The folks who I pinged about it were amazed that I somehow had managed to "guess" that a specific member of their team had been poking at a specific box"?
read like "I didn't like what someone did on a machine I had to troubleshoot, and told their manager" to me.
I was also reading this prior to coffee, mea culpa.
It’s called Chef.
If you have a hardware failure how would you rebuild?
Of course this all takes time to implement and manage -- it's a tradeoff.
Mikrotik routers and switches can boot from dhcp (or bootp?), but yes a typical switch can't.
Not a network person, but I assume you can give fine grained enough control that you can't do a "copy running-config startup-config" on a cisco switch, so have a startup config that boots to a known basic state then tftps it's config from your HA dhcp server.
Why revert it in the first place? Why not deny the changes? Doesn't this indicate that the wrong people have root access?
This just makes you sure actually go back and change the deploy code, else it breaks in ~2 days.