Not being mean to you, just showing how typically the goal posts are moved.
To give you an example from physics, if you find just one experiment that goes against your model, you immediately invalidate the model. You don’t just make grand claims that the model in general works.
Usually isn’t it looking for an experiment that proves it and is repeatable?
If I discover a new element in one experiment, the results are published.
After publication, many labs will try to repeat and its not taken away if one can’t do it. Only if all can’t and it casts doubt on whether I did it in the first place.
Example of a test that invalidated our old theory of gravity and validated Einsteins claims:
https://en.wikipedia.org/wiki/Eddington_experiment
This is how science is done. But apparently not data science.
Maybe very far in the future there will be models of human biology that are as robust as classical physics, but right now there is such a large amount that is not understood, it's simply not feasible. A drug could work for one person and not another for reasons beyond the realistic scope of the original development hypothesis. It requires a probabilistic view to make any sort of statement about the efficacy then.
I suppose you could argue these models are just wrong and thus trivially disproven, but I don't think that's a productive framing. I doubt any biologist or doctor would claim they have anywhere near a complete model of how their specialty works. That doesn't mean a particular model isn't useful or isn't the best we currently have to work with.
Plus maybe the third best model will actually turn out to explain a separate puzzle piece in an eventual better model. Mechanistic models in biology aren't always well done in practice, but it's certainly not binary either.
Pierre Duhem would like to have a word with you:
https://plato.stanford.edu/entries/scientific-underdetermina...
> Holist underdetermination ensures, Duhem argues, that there cannot be any such thing as a “crucial experiment”: a single experiment whose outcome is predicted differently by two competing theories and which therefore serves to definitively confirm one and refute the other.
It’s like when people got upset about Trump winning when 538 only gave him a 30% chance. That one event tells us nothing. But if all predictions 538 says have a 30% chance of occurring happen 30% of the time, then they are spot on. That’s not apparent with a single event though.
The problem is that most managers, companies, and people (including yourself, apparently) are statistically illiterate enough to not understand this, and jump head first into data science initiatives expecting immediate results, which is usually doomed to fail, at which point they blame others and not their poorly formed expectations.
There’s plenty of bad data science out there, but most failed data science initiatives are doomed before anyone every builds a model or analyzes any data.
Who do you think you are?