I'm currently working on solving a problem with an insanely over complicated setup (for the task at hand) that is build by another engineer who has since left the company.
It's a cluster of 3 virtualised machines running docker swarm where a RabbitMQ instance ties 40+ worker pods together. Once every 5 days or so the connection between RabbitMQ and (some of) the worker container stops working, causing the worker to crash and the queue message is lost.
We are talking like 5 layers of virtualization and/or abstraction. It's impossible to debug. I honestly don't know how to explain this to my customer.