How about the costs? Isn’t this a very expensive bandaid? How is it not a priority? :)
I've seen multiple issues solved like this after engineering teams have been cut to the bone.
If the cost of maintaining enough engineers to keep systems stable for more than 24 hours, is more than the cost of doubling the container count, then this is what happens
x = time it takes to switchover
y = length of the cycles
x/y = % increase in cost
For us, it's 15 minutes / 120 minutes = 12.5% increase, which was deemed acceptable enough for a small service.