It's why people spend a decade getting a PhD, and many years more still to become a professor. You can't do good science by putting numbers in a spreadsheet and have the computer do all the thinking for you.
Perhaps you are thinking about a situation where the result of an AB test is accepted and incorporated into a product without reflection on the associated factors and context?
My point is that the process of designing experiments and interpreting their outcome is nothing other than science.
Science is hard and goes much deeper than the trivial stuff like significance and power. Even scientists, who know all these formulae inside and out struggle with constructing experiments and interpreting the outcome.
To properly think about experimental outcomes, you need to actually have a firm grasp of statistics. The moment you try to convert it into a message that doesn't require understanding statistics is the moment it stops being informative.
* The better the tooling in general increases the proportion of people doing it _wrong_, because when it was harder, this selected better for people who wanted to do it right. Making it easier means more people do it right, but more people who would otherwise not do it at all can now do it, and do it wrong.
* Making aspects impossible by design will impress and amaze you with how humans and groups figure out how to not only prevail over what you tried to make impossible, but do it even harder than if you did nothing at all.
The best I think we can hope for is making things by default easier to catch, or easier to find later, or less deceptive.
Hard things are hard and that's ok. I think we should spend more time channeling our empathy into aspiring for ourselves and others to be better and do hard things, and making the ability to learn and do hard things accessible to everyone who wants it, as opposed to trying to pretend hard things are easy.
I'm not confident that it makes a big difference, because in the face of anything "idiot-proof", nature will provide a better idiot.
It is really easy to A/B test having a small static button vs a large, flashing, jumping button with sound, and measure engagement as "ever clicked the button". It passes statistical power and significance since most of your users now click the button, even if none of them did before. The failure here is that the button is simply annoying, and clicking the button is not a legitimate engagement with the feature, hence the statistical power to answer the question "does a flashing button promote engagement with the feature" is actually zero.
A change in mindset can be brought can be encouraged and developed by a combination of process and community. This includes educational processes supported by public policy, market mechanisms, and communities.
I believe we need a revolution on how we educate ourselves. How we get there is non-obvious.
If sound statistical thinking was taught, nurtured, and synthesized across many domains, I think it is likely that humanity would be better off.