Finding zombies in our systems: A real-world story of CPU bottlenecks
medium.com
medium.com
Especially noted that the root cause here was pegged because somebody else had blogged about it. Our "digital commons" should be valued as other than an externality. Edit: by that I think that "the internet should be free" but that commercialization through derivative work should compensate authors. That's off-topic, so, nevermind.
I don't know anything about the Linux scheduler, but if one core out of 96 is busy chasing after those ghost allocations, why would a network thread not use one of the other 95 cores, and instead be stuck for seconds? Unless that one core is blocking the entire kernel, but then that would be a lot more visible instantly.
On NT, this style of problem would be slightly less likely. The job object, which is NT's memcg equivalent, is a handle based object that naturally goes away when the owning process dies.
I also bet systemd would have native mechanisms to clean this memcg state up if it was used for this purpose. Seems like a bug in the AWS ecs code.
63 Cores Blocked by Seven Instructions: https://randomascii.wordpress.com/2019/10/20/63-cores-blocke...
Looks like Linux isn't immune to these types of problems.
Heck, even if userspace is trying to use all 96 cores to their fullest, I would expect the ENI driver to be high enough priority to pre-empt userspace and still not fall over?