A lot of the problem is that people pay attention to the hype of news articles rather than the particulars of the papers.
In the one you link, they don't describe what they mean by "hold up." It turns out that the only definition that those authors give is in the caption of Table 1
http://www.nature.com/nature/journal/v483/n7391/full/483531a...
>The term 'non-reproduced' was assigned on the basis of findings not being sufficiently robust to drive a drug-development programme.
Now, this is a completely different thing from reproducing in terms of getting the same science. Driving a drug program is not biology, it means does this phenomena that you found in one thing apply to the vast majority, so that a pharma like Amgen can run a trial, put all stage 4 patients on it, and have it work on enough people that the trial has good enough stats.
However, that's really not how cancer works. Cancer is extremely diverse. If somebody is publishing a "landmark" paper (whatever their definition is of that), then they are publishing something new and unexpected most likely. That means that the landmark finding is almost certainly not going to generalize to enough that a large pharma can run an uninformed clinical trial and throw it all patients without pre-selecting for patients where it will most likely work.
Further, that silly thing is just an editorial. It's a plea for science journals to only publish things when it's found in the majority of cancers, or something. If we did that, we'd know far less about cancer, because it's treating something that's actually diverse as a monolithic entity.
Are you going to get a particular cell line to have exactly the same gene expression as another lab? Probably not, if it's a tricky to grow cell line. Are you going to get cancer drug sensitivities to be all the same? Probably not, if you're testing when the cell lines are growing at different densities.
People really should distrust that any individual data point from a particular model biological model is going to generalize, especially for difficult biological models. That doesn't mean we should stop publishing, because if we do, we're not going to be able to figure out what's actually driving those differences.