> > the studies didn’t achieve statistical significance, a controversial but commonly used threshold for publication
> why would you publish a non statistically significant result
There's actually two separate things going on here.
The first is that researches (at least in the biomedical sciences) often decide not to waste their time writing up results that are either weak, inconclusive, or negative. This is only mildly controversial - everyone involved recognizes the inherent time availability constraint, but knowing what didn't work can help to inform the field at large in it's own way.
The second is that the term "statistically significant" itself has become _highly_ controversial in recent years. The issue is that when you perform statistical tests, they generally (oversimplifying) spit out a number telling you what the odds of your results being correct are.
In concrete terms, say you have some graphs that show a slight improvement in some condition when a drug is used. Is the drug actually effective, or is the "improvement" actually just random measurement errors? So you run a statistical test on your data and it gives you a score of 80%. It's telling you that there's an 80% chance that your results were due to an actual improvement and a 20% chance that they were due to random chance.
If you had unlimited time and money you could just keep collecting results endlessly. Eventually, that score would either go towards 100% (it definitely works) or 0% (it definitely doesn't work). But this is the real world, where we don't have unlimited time and money (particularly academic researchers).
So when is your result worth publishing? How do you decide when to throw in the towel? And if you're reviewing papers for a journal, how low a score is acceptable before you vote to reject the paper on the grounds that the conclusions are unreliable and not worth looking at?
Enter the term "statistical significance". At some point, people started classifying scores that were above some arbitrary threshold as "significant" and those below it as "not significant". But there's an obvious problem here - some measure of probability flipping from (say) 89.999% to 90.000% doesn't do anything magical! Worse, what a future reader intends to do with the results will determine how important any given score is in that particular case. Clearly, results and their associated score need to be interpreted in context instead of blindly. Using a term such as "statistical significance" flies in the face of that by actively encouraging lazy thinking.
So the controversy being referred to in that specific sentence you quoted isn't the decision not to publish but rather the usage of the term itself. (Which is confusing, because the article at large is addressing the controversy surrounding not publishing.)