It was first suspected to be certain brands of RAM, so I requested a RAM-swap which unfortunately didn't help. Then a BIOS update which also didn't help. Then someone figured out that nohz=off on the KCL fixed the problem and I had it running like this successfully for a few years. Long after at least one dist-upgrade I remembered that and removed the option again, and the server still ran stable.
There's no real morale to this story I guess, but at least the support is super responsive, and as the root cause wasn't clear at that point didn't hesitate to swap random stuff if you requested so. Also had a faulty HDD last Sunday in one server and requested a swap, which they did within 20 minutes of me opening the ticket.