> Except we don't learn the most important lesson: Irresponsible people need to be off the team.
This is false dichotomy.
If people can screw something up, sooner or later they will. It doesn't matter how responsible or irresponsible they are - the processes in place should prevent mistakes from being made regardless of that, or to mitigate their consequences if they're unavoidable. Anything less and you're not addressing the root cause of the issue. These same processes should ensure that no one can create changes to the state of the system before them first going through another set of eyes and being validated.
Most sane OSes at least prompt you before deleting a file - the very kind of safeguard that makes you double check whether what you're doing makes sense. Similarly, you'll notice that cars have seat belts and air bags, even if you're not going to crash daily. Get the irresponsible people off the team if you'd like, but if you can't put the appropriate processes in place to prevent mistakes, you probably need to rethink your priorities and provide the adequate pushback against "the business", when they expect you to SSH into prod and do anything.
What that looks like in my current environment:
- all server configuration is managed through Ansible and Git
- developers have read only access to the servers when needed, no one can change the state of the system there
- all changes are versioned and use merge requests and automated CI processes, need to be reviewed first and synced with the change management system, to know who is attempting to change something, why and also how (with the appropriate permissions)
- furthermore, the Ansible processes are scheduled to run every day as well, just to make sure that the server status is as expected
- of course, the changes to servers are also first done on development environments, then on accept testing environments, then on clients' testing environments (any number of them that's necessary) and then finally to prod
- there is Zabbix for monitoring the infrastructure with alerting in place, as well as Skywalking for application APM, as well as some lower level tools
- the apps are also shipped as containers, which have very similar processes in place, e.g. full CI run on every merge request, before it gets merged
- those containers run in clusters, so if any of them fail for whatever reason, the load will be balanced as necessary and/or restarts will take place, each container also having health checks
- lots of manual QA in there as well, to catch the things that aren't easily automated, as well as regression testing
- automated integration and load tests done every now and then, to check that there are no regressions in a new release
- testing the reproducibility of this can also be done by wiping any server and letting the CI processes restore a new one, complete with the appropriate access roles, firewall configuration, cluster membership, observation tools for developers to use etc.
- everything also has backups and rollback strategies as needed
- despite all of the procedures in place, delivering a new version can be as easy as clicking a button on a CI pipeline, the container images being delivered and becoming available on the clients' side in a few minutes for further processing
Of course, that's still not good enough in my eyes and there are further improvements that can be done, as well as problems in place that prevent truly safe and easy development (e.g. no adherence to 12 Factor App principles due to historical reasons). If the projects used TDD and there was 90% test coverage, as well as ALL of the functionality was to be covered by integration tests (Selenium), then we could actually start talking about
software engineering. Until then, i don't believe that the term is really apt, perhaps outside of the aerospace industry.
Disclaimer: of course, the amount of effort that goes into designing a system and the workflows around it depends on a variety of factors, so not all systems need that sort of focus on quality. For most CRUD apps out there, you can probably just wing it, even though then you must be ready for the eventual outcomes and downtime that it will cause.