At scale, every test is flaky.
DISCLAIMER: I've been around IT for probably the majority of y'all's lifetimes, so I'm not saying this happens often. But just because something is fundamentally wrong, doesn't mean that all fundamentally wrong things are the same. In my experience they differ more from each other than the possible good ways of doing the same thing. Don't conflate things without a good reason.
It's once you start doing tests in real environments with real databases and real networks that the unpredictability of the real world creeps in!