> not have automated tests for their config files
They very likely have automated tests. However, what if bug only triggers 90% of the time and you hit the lucky 10% during automated tests? Of course you can run tests 100 times but... is this a common practice? Moreover, we have both code and anecdotal evidences that the bug may indeed happen randomly. Tavis Ormandy posted a rough analysis of the crash context: https://x.com/taviso/status/1814762302337654829. It looks like the crash is caused by first checking if an uninitialized pointer is NULL, and if not, dereferencing it. If the uninitialized leftover data just happened to be zero, no crash happens.
And anecdotally, we saw people reporting that repeatedly rebooting their machines for 15+ times fixed the problem for them - because eventually you got lucky and in a boot it didn't happen and CrowdStrike managed to update itself to not crash.
> not have gradual/staggered rollouts for their deployments
No idea. Maybe their poor reliability guy got overrided by another team, like "how dare you delaying our important definition update? we're racing with threat actors!". I hope they learned their lesson.