:-)
Had a colleague who asked for SSH access to production machines to debug an issue. Ops team asked what he wanted to do, guy just wanted to look at which env vars were set. Ops team told him how to do his job - log the configuration - instead of give him access, because they had a mandate to ensure five or seven nines uptime. Can't risk it.
I had a lot more respect for the ops people there than I had for my fellow SWE's.
But yes, CLI can be dangerous, even for observability. E.g., people running strace(1) on production apps and causing outages due to the strace overhead (I wrote a prior post about that). You need to understand the risks and overheads of all tools. It's why I have a "pull no punches" policy when writing eBPF tool man pages [0]: If the overhead can be bad, it should say so clearly.
[0] https://github.com/iovisor/bcc/blob/master/CONTRIBUTING-SCRI...
That's a change, inherently much more risky than the "cat /.../env" the colleague wanted to do.
Also the change might well cause the problem to go away, and now you know nothing instead.
The principle is good, but it sounds like it has taken on a life on its own. The ops guy could also have executed the cat command on the spot with the right privileges. Sure, it's gatekeeping, but so is the four eyes principle, and a little of that gatekeeping can be necessary to keep those nines rolling.
"do your job" is a really rude way to refer to debugging by logging instead of debugging interactively.
(Another interpretation is that the author previously held Ops in higher regard than SWE and that this event did not change that.)
Also like others noted here , If one guy sshing into one machine can screw your seven nine uptimes.
Then you never had a seven nine uptime (you’re company is probably just gambling on that uptime metric until some component fails)
I’d still rather not have to debug arbitrary mutations to the env or file system in a production container though.
It seems to me that shelling into a prod container/VM is discouraged not because you might cause it to fail, but because you might produce undefined behavior while claiming it is healthy (more like Byzantine faults).
For example if you unset a single env var by mistake, then 1/N of your requests will potentially fail. Debugging this kind of issue is a nightmare.
Not to mention that developers can often run arbitrary SQL from a prod shell when the app is backed by a DB.
Or far worse, not fail.
There are always bugs that will happen only on the prod machines. Sure, they are rare, but they exist.
> just wanted to look at which env vars were set. Ops team told him how to do his job - log the configuration
Well, that's not risk free neither. That needs a code change and you risk exposing secrets, or logging more than you should. (Though I agree with the general idea of you shouldn't be SSHing if you can do it another way)
Dogma results in religious wars
So you can log in and debug, but once you're done, the VM is replaced with a clean one.