My quick read of the study suggests that they examined several outcomes, got null or low powered results on most of them, and chose to highlight the single result which did show a significantly positive effect on recovery; that result was based on a meta-analytic estimate of the pooled sample of just two studies. In particular, my understanding is that this study found a null on "percentage days abstinent", a null on "longest period of abstinence", a null on "drinks per drinking day", a null on "alcohol-related consequences". For alcohol addiction severity they report the results of one study without doing any original analysis. And for "rates of continuous abstinence" they found a positive effect. In most cases these are graded as "low certainty evidence". The positive effect is claimed to be "high certainty evidence", but the CI / p-value for the positive effect is not adjusted for multiple comparisons -- I don't think that's fatal, but it does speak to the fragility of the outcome.
The other concern I have is that they separate "manualized" treatments (e.g. AA treatments administered according to a protocol) from "non-manualized" (e.g. ad hoc administration of the AA treatment) treatments. This is fine -- we would expect to see that if something about the protocol was useful, the non-manualized treatments would display an attenuated version of the same effect. Instead we get a null on the original positive effect, and a positive effect on one of the original nulls
Finally, the authors seem a little loose with what the comparison groups are. They claim they recruited studies which compared AA to no treatment, but the results are all motivated as being AA versus CBT. What happened to the no treatment studies? Perhaps these are buried in the full text. The rate of spontaneous remission of addiction is believed to be fairly high and most of the past criticism of the efficacy of AA has been motivated by comparison to spontaneous remission.
It strikes me that there is significant distortion of the study findings from the Cochrane write-up to the Stanford press release, and from the Stanford press release to the NYT piece, and that in both cases the distortion is in favour of claiming the study has found affirmative evidence that AA works better, rather than just no evidence to claim it works worse. For instance, one of the thrusts of the NYT piece is that the past Cochrane review was based on a limited number of studies... but the operative finding in this review is based on an even more limited number of studies, even if the pool of studies from which they drew has become larger over the last 15 years.
All of this is from a quick read of the review -- I am off campus right now and don't feel like VPNing or pirating the full text to deep dive the analysis.