You should absolutely always be typing out huge, verbose commands like kubectl —context my-context -n my-namespace <some commands> ...
The opportunity for tragic error when you have implicit context or namespace settings applied is just too great.
You should absolutely always be typing out huge, verbose commands like kubectl —context my-context -n my-namespace <some commands> ...
The opportunity for tragic error when you have implicit context or namespace settings applied is just too great.
It’s always an available option for people who want it, but many many use cases are poorly served by it.
The RBAC thing is different. I am 100% for those RBAC limitation as long as I have total authority to make a devops admin do what I need them to do right when I need them to do it.
If my team is going to get alerted on build failures or service errors that require kubernetes mutation actions or exec to debug or resolve (99% of the time they do), then my team needs full admin access.
If rbac prevents access, then you better route the alerts to someone else or give us authority to compel instant triage responses.
In the 3 different companies I’ve worked in that use large kubernetes clusters, rbac has been a miserable problem across the board, because SREs don’t want to be responsible for triaging or resolving application team issues, and application teams will not honor alerts that are not actionable because of poorly conceived permissions issues where people mistakenly think they are setting useful access control policy but really they aren’t.
At Zalando everyone has read access (except to secrets). You can request write access with a command line tool another employee can then approve it. There is also an option to request access with an incident ticket, in that case you immediately get write access without approval by someone else. Write access expires after 1h.
Access to the underlying AWS account uses the same mechanism.
You might use use-context to set yourself to that restricted production context at the start of an incident. 30 minutes later you resolved the incident, but the access is still active and whoops you delete a deployment you thought you were deleting from a stage experiment or something.
All that does is add an extra hoop to jump through. People need to stop believing that adding extra bureaucratic hoops offers any type of safety - it doesn’t.
If you want to restrict access then you need to actually restrict it and convert on-call alert responsibilities to a central devops team that, like it or not, is responsible for solving everyone else’s problems.
If you don’t want a central devops team that can be paged and on the hook for everyone else’s systems, then you cannot have write access controls.
There’s just no way out of this dilemma. Temporary access that can be self-granted == no access control. Temporary access that requires admin approval == admin team is on pager duty for every other team.
You can also explicitly add namespaces to manifests, which help avoid applying to the wrong namespace.
So make sure to spam the return key before deleting stuff :D