Are you really suggesting that most GWAS studies don't calculate genome-wide significance? This is wrong. If that's not what you're suggesting, I don't know what you're saying.
Are you really suggesting that most GWAS studies don't calculate genome-wide significance? This is wrong. If that's not what you're suggesting, I don't know what you're saying.
False discovery rates, false positive rates, and family-wise error rates and how best to control for them are ongoing areas of interest and research in GWAS. There are calls for p-value requirements in the 10^-8 range to help avoid this. There are calls to address stratification (which can affect both type I and type II errors). A lot of research and debate on this topic over the past several years. Who are you exactly to dismiss all of this scientific inquiry? It's great if you have expertise in another field, but it seems odd to dismiss scientific questions and ongoing research in this particular field.
You'll find plenty of information if you actually seek it out rather than simply making a knee-jerk, snarky comment, but here is one example article which articulates some of the issues that have been under consideration in recent years: http://m.ije.oxfordjournals.org/content/41/1/273.full
Note, there, how the level of significance is discussed. A p-value of 10^-7 to 10^-8 is suggested (as compared to this study's 10^-5 level of significance... and, believe me, many GWA studies have been published with much less significant p-values).
This is an ongoing and active area of discussion in the field. I'm not an expert, but some of my colleagues are, and it's a topic they sometimes discuss and brainstorm over lunch, etc.
Actually, even the wikipedia page on GWAS mentions some of this inquiry and debate, as well as the erroneous publication that has plagued this nascent field. It'll all be worked out over time and great discussions are happening here. Vast improvement in processes and standards has been made over the past few years in particular. But we don't move the ball forward by dismissing questions or incorrectly assuming all must simply be right and well.
I assume you are referring to
" If the seven associations that did not reach P ≤ 5 × 10−8 when additional data were considered are assumed to have been false-positives, the false-discovery rate for borderline associations is estimated to be 27% [95% confidence interval (CI) 12–48%]. For five associations, the current P-value is > 10−6 [corresponding false-discovery rate 19% (95% CI 7–39%)]."
That doesn't show anything relevant. Failure to replicate at 10-8 is a ludicrous way to define non-replication; to paraphrase Cohen, surely God loves the 10-7.99 almost as much as the 10-8... This paper needs to adjust for power, and ask how many hits one would expect to not replicate at 10-8 given the power of the replicating studies. If you do remember power, GWASes replicate fantastically, for example https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3681663/
"Replicability rates are high within Europeans, with 155 successful out of 181 attempts (85.6%), when only 9 positive replications (∼5%) would be expected under the null hypothesis of no association (binomial test, P<10−16). This excess was robust to the significance threshold (e.g. 122 observed vs. 0.18 expected if only replication attempts achieving P<0.001 are considered successful and 56 observed vs. 1.8×10−5 expected for a threshold of P<10−7, Table S5). Moreover, replicability rates within Europeans approach 100% when accounting for statistical power. For the 168 attempts for which we could calculate the power to replicate the original finding (Table S5), we observed 147 positive replications, which is almost identical to the expectation of 149.1 positive replications given that average power is 89.1% (see Materials and Methods). This is expected, since most GWAS already contain an internal replication phase [1], [24]."
> and, believe me, many GWA studies have been published with much less significant p-values
I don't think they have. Ever since Ioannidis and others demonstrated what a total debacle the early candidate-gene studies were around 2009-2011, using the first GWASes to demonstrate that, GWASes have been pretty standardly done at 10-8.
> I don't think they have. Ever since Ioannidis and others demonstrated what a total debacle the early candidate-gene studies were around 2009-2011
Yes, pre-Ioannidis is the period I was referring to. Things cleaned up a lot in 2012+. What you said does not conflict with what I said... the field, especially early on, has published some spurious results. It seems like you're simultaneously saying "I don't think they have [published low-quality results]" and then immediately admitting what a "debacle" early studies sometimes were.
> That doesn't show anything relevant. Failure to replicate at 10-8 is a ludicrous way to define non-replication
This is a silly statement. Doesn't the significance level being "ludicrous" depend on things like the degree of multiple testing happening? Obviously, yes. The reason why the field (not just this one paper) has pushed for significance in the 10^-7 to 10^-8 range is partially for this reason. So... are you questioning the entire field's movement over the past few years? If so, on what basis?
As high-throughput, low-cost full genome sequencing begins to replace SNP-based techniques, GWAS will have to wrestle with this issue even more.
I'm not sure what exactly you're debating me on here. I'm not saying anything controversial in the field. Again, even the wikipedia article cites well-known studies and quotations from respected sources in the literature, including "Particularly the statistical issue of multiple testing wherein it has been noted that "the GWA approach can be problematic because the massive number of statistical tests performed presents an unprecedented potential for false-positive results"... which is what I originally pointed out. This is an issue the field has struggled with from the get go. It's matured and is much better now (as I've noted), but the field still struggles with the issue. And there are still low-quality papers being published. I also gave this particular paper praise for holding to a higher standard than some other GWA studies.
A candidate-gene study != GWAS. It's particularly bizarre to criticize GWASes for the sins of candidate-gene study when GWASes were literally partially designed to avoid those problems. Don't equivocate. If you have criticisms of actual GWASes as they are run and good reason to doubt that the hits are noise and will not replicate in well-powered followups (contrary to what we actually see, in this GWAS and others...), give them; don't swap in criticisms of candidate-gene studies and pretend they're the same thing, because they're not.
> Doesn't the significance level being "ludicrous" depend on things like the degree of multiple testing happening?
It does. And that's exactly why holding replications to 10-8 is ludicrous. The error rate when testing 5 or 15 SNPs at 10-8 is much much much smaller than when testing 500,000 SNPs. Why would you hold a multiple-test of 5 tests to the same standard as 500,000+ tests?
> I'm not saying anything controversial in the field.
If you're insinuating that GWASes are as bad as candidate-gene studies were, or that a p-value threshold of 10-8 should be used for everything, or that we cannot have high confidence in any given hit at 10-8 (especially when replicated at 10-5) you certainly are saying controversial things.
> I also gave this particular paper praise for holding to a higher standard than some other GWA studies.
Such as?
> A candidate-gene study != GWAS. It's particularly bizarre to criticize GWASes for the sins of candidate-gene study when GWASes were literally partially designed to avoid those problems.
Obviously. I don't understand why you are referring to candidate-gene studies, as I haven't mentioned them at all. Are you saying that GWASs did not exist prior to 2012? Are you saying that GWASs somehow have not had to develop more rigorous analyses, change their standards for publication, or discussed and improved replication techniques in recent years? That somehow false positives, FDR, and FWER are of no concern in any published GWAS study today? That early GWA studies didn't struggle with effect sizes and power or that there aren't plenty of examples of published GWA studies that are underpowered or even which published spurious results? That there aren't fundamental questions about the assumptions underlying most GWAS analyses (eg potential epistasic effects vs common GWAS assumptions of SNP independence / additive effects... which is a big question in the field currently).
I'm talking about GWA studies. And I'm just pointing out well-known issues the field has wrestled with. For example, here are a couple of the early, fairly influential papers overviewing issues with GWAS (not candidate-gene studies):
2008.. http://jama.jamanetwork.com/article.aspx?articleid=181647 "GWA studies are an important advance in discovering genetic variants influencing disease but also have important limitations, including their potential for false-positive..." This and this paper, 2009: https://projecteuclid.org/download/pdfview_1/euclid.ss/12717... both helped move the field forward in more rigorously examining analysis and defining how replication could be achieved.
These are helpful papers which moved the field forward substantially and didn't just point out issues then-currently facing the field, but also pointed out solutions (many of which were subsequently adopted). It demonstrates that GWAS has struggled with such issues. I imagine you are familiar with some of this, as you mentioned Ioannidis. So... what are you even saying here? Power, effect sizes, false positives, FDR, FWER, etc have been major issues since GWAS' early days, and despite a great deal of progress over the past few years, it's still an ongoing area of discussion, research and debate in the field.
When I say caution is warranted and that we shouldn't over-generalize results, that's because the field has learned this the hard way.
> GWASes were literally partially designed to avoid those problems.
Yes, they were. And also, as I'm saying, they suffered from analytical and publishing issues for a number of years, despite being designed to address some of the early issues in single gene association studies. In the past few years, things have gotten much better. But there is still cause for concern and a lot of ongoing debate in the field.
> Don't equivocate. If you have criticisms of actual GWASes as they are run and good reason to doubt that the hits are noise and will not replicate in well-powered followups (contrary to what we actually see, in this GWAS and others...), give them; don't swap in criticisms of candidate-gene studies and pretend they're the same thing, because they're not.
I'm in no way equivocating, and you are the individual here discussing candidate-gene studies, not me. I have no idea why you're referencing them when I'm talking specifically about GWA studies.
GWA studies, currently, today, as of right now... still have issues with FDR and FWER at a base level, despite improvements in analytical rigor. With study power. With changing data collection techniques and quality control. With population sample bias. With concern for heterogeneity when applied to difficult-to-definitively-diagnose or subjectively defined traits (such as mental health diseases) in particular. With myriad issues.
It's still a relatively nascent, immature field and it is absolutely silly and unhelpful to dismiss a call for caution in interpretation, especially in the light of public press which is historically prone to jump to conclusions, under-report detail, under-qualify nuance, and overstate results.
If you want some recent examples of this debate in the field regarding GWAS, not candidate-gene studies, see a number of articles from even the past few years, including the following couple examples:
2013: http://www.nature.com/nrg/journal/v14/n7/pdf/nrg3457.pdf This is also in Nature.. and it was published because the editors in the field thought it held an important message for the field... despite what you seem to be saying here.
2016: http://www.ncbi.nlm.nih.gov/pmc/articles/PMC4756503/ "Identifying disease-associated SNPs facilitates the clinically relevant task of identifying higher-risk individuals. However, the large amount of reclassification that we demonstrated in individuals initially classified as Higher Risk but later as Average Risk or Lower Risk, suggests that caution is currently warranted in basing clinical decisions on common genetic variation for many complex diseases." [Emphasis added] This example is very on the nose in calling for exactly the same sort of caution that I'm calling for here, and which you for some reason seem to be disputing the merit of. This paper specifically addresses problems that arise when different numbers of SNPs are used in calculated genetic risk scores, as well as other issues common in the field today.
Now, I am not saying the GWAS are generally spurious or in any way bad science. I am saying we should treat them with caution, because it is a nascent field. Because there are a lot of underlying analytical assumptions and heavy analytical techniques involved in deriving results. Because the field has already seen a lot of flux over its short lifespan. Because there has been some incidence of spurious publication in the past. Because there are active and ongoing debates in the field about quality control, replication, analytical technique, etc. Because there are issues with population sampling. Because pathways aren't well-understood in many cases. And because, as I initially said, it's easy to get false positives here. It's hard to get appropriate power with a high degree of quality control. It's hard to get fully independent replication data of appropriate power. The field has challenges (against which brilliant and valiant efforts have and are being made), yet those challenges and qualifying statements rarely come to the fore in the public press regarding GWA studies.
Are you actually disagreeing that caution is duly warranted, or are you just picking at nits?
> here is one example article which articulates some of the issues that have been under consideration in recent years: http://m.ije.oxfordjournals.org/content/41/1/273.full
From the abstract of that article: "Currently, associations of common variants reaching P ≤ 5 × 10−8 are considered replicated. However, there is some ambiguity about the most suitable threshold for claiming genome-wide significance." So, people do calculate genome-wide significance, and there is some ambiguity over where exactly that line should be. This is in line with my understanding of the situation.
The exactly analogous statement can be made within particle physics ("there is some ambiguity about the most suitable threshold for claiming significance"), where it is often called the "look-elsewhere effect". But this prudent caution doesn't cause to people to say FUD like "particle physics studies should be treated with great caution" or "Most particle physics studies leave a lot to be desired". Such statements may absolutely in fact be true for GWAS studies, but you didn't give any good reasons for it.
Indeed the very article you link to ends this way: "Conclusion: A substantial proportion, but not all, of the associations with borderline genome-wide significance represent replicable, possibly genuine associations. Our empirical evaluation suggests a possible relaxation in the current GWS threshold."
How should one square that conclusion with your original comment?
You asked initially what exactly I was saying, and here again how to square what I'm saying with my original comment. So, let me try to explain, and I would hope we're not on different sides of this as I think what I'm saying is reasonable, given the context.
I originally said: "GWA studies general should be treated with great caution. The way they work generally is based on a simple p-value test of association among outcome (in this case, depression) and all genes based on SNPs. There is a high degree of mere chance association and false positives. Most GWA studies leave a lot to be desired."
For context: genome-association studies have had a history of being blown out of proportion in the press. And often for outcomes which greatly affect people's lives. Depression is one such issue and it'd be a shame if people were led into thinking there is necessarily a great breakthrough here in understanding possible gene-linkages to depression outcomes. It'd also be a shame if the result was ignored. I tried to provide some praise to this paper for being fairly rigorous, but also note that GWAS studies should be treated with caution generally.
Why should GWAS studies be treated with caution generally? (Besides that results of any study should be treated with some degree of caution.)
Well, firstly, GWAS is a fairly nascent field. Unlike physics or even particle physics, it hasn't had that much time to mature. This is doubly so when applying GWAS to mental health. I'm sure the paper covers some of these risks (or at least it hopefully does)... but relying on self-reporting introduces potential selection bias in the population sample, as does using people seeking help vs a general population study. Quickly reading the paper, it looks like it relied on self-reports and analyzed only people who'd been diagnosed with major depression (meaning they'd sought help). We should be cautious in over-generalizing based on this.
Secondly, GWAS has had a rough history of overstating results and misapplying analyses. It's much better today than it was even, say, 5 years ago. Ioannidis and others made some heroic efforts to convince the field to clean up it's act, in effect starting ~7-8 years ago.
Thirdly, there is a historical pattern of results in this field being overhyped in the press.
Finally, there are active and sometimes heated debates in the field about how best to do GWAS. This is getting worse as a high-throughput, low-cost full-genome sequencing comes online to a greater and greater degree and SNP-based data sets fall by the wayside. Some question taking a frequentist approach at all in the face of such a huge degree of multiple testing. Others call for much, much higher requirements for holdout data sets, cross-validation, and replication before a study is published or considered final, especially when dealing with things like mental health (and their likely application to the field of pharmacology).
This is serious stuff that could end up affecting people's mental health treatments and lives, so caution is warranted. Especially given the field's relative nascence, self-admitted history of publishing low-quality results, the rapidly changing techniques, and the fact that there are ongoing debates within the field of how best to do GWAS analysis and how to effectively replicate results.
Look at the language there. Scientists "pinpointed 15 locations in our DNA that are associated with depression..." [emphases added]. There is no sense of nuance or any caution in conclusion here. Even reading the whole article, which perhaps most people won't even bother with, you don't get much sense that there's any degree of uncertainty here. There is no indication that the field at large is still wrestling with how best to analyze SNP studies at all, despite publishing studies "practically ever week". There is little qualification around the data set here (it only briefly mentions that 23andme's data is based on saliva, not blood or other cells... which may be problematic. But also it should be noted that disease outcomes in this study were self-reported and from a self-selected subset of the general population who sought professional help... and there are many, many other nuances to consider). Instead, we get the fairly flat-out impression that this is a definitive discovery and, not only that, but it is of huge significance in the field as well.
It may well be. But caution is warranted until time and replication hopefully do their work. Unfortunately, this is a pattern in GWAS studies and their public relations. And while the underlying methodology in the field has improved tremendously over the past 5 years, there is still a lot of debate within the field about how much more improvement needs to be made (on data sources, on analytic techniques, on review standards), and how to adapt to forthcoming changes in available study data.
I think this study appears to be more robust than many others. So, I'm not picking on it. In fact I gave it particular praise for being relatively robust compared to some other GWASs. But I do think generally people should be more cautious about interpreting GWA studies or over-generalizing them.
And, yes indeed, false discovery rate has been a huge issue which the field has highlighted... that is my point. Caution has been warranted here historically, and it continues to be warranted even today.