Dispelling Myths about Randomisation
bps.org.uk
bps.org.uk
I do not think this sort of word-play is useful. If your random samples are small (and even if statistically adequate) the chance of confounders in one or both groups of an A/B can be relatively high, even though the selection procedure for treatment is random. So, "No – not least because confounders do not exist in experimental studies!" is misleading if it is expected on the basis that randomisation of treatment allocation somehow makes confounding impossible.
That the possibility of confounding is equally likely in both branches remains true for all sample sizes of an A/B where #A=#B and allocation is random. So, in my opinion, not a myth.
The idea that even thousands a data points in subgroups are going to be 'well mixed' relies on extremely strong assumptions about the distribution of those traits.
All these things can and do happen in randomised experiments, but it is still orders of magnitude more interpretable than what can happen in observational studies.
If on the other hand you just redo the sampling until it seems balanced then you’ve violated the assumptions behind the standard statistical tooling.
In English, a confounder is any factor that distorts an observation. (My dictionary defines it as throwing into confusion or disarray.)
In causal inference, a confounder is a factor that is correlated with both treatment and outcome. If the treatment is randomly assigned, by construction it is independent of all other factors. This, there can be no confounders.
Your example is about observed occurrences of imbalance, but the technical definition is about probabilities. Observed imbalances can still skew inference, but that causes high variance (or low precision). It doesn't cause bias (or affect accuracy).
Adjusting for observed imbalances can reduce variance, but in some circumstances can actually cause bias.
Many do think of confounders in an experimental context as just those effects which correlate with both outcome and treatment. The non-sequitur — barring a specific definition — is concluding that since nothing can correlate with random allocation, confounders are impossible by construction.
Why impossible? Because we are talking about the probability of allocation, not the actual allocation, and confounding does not refer to the result. We’d instead say there are imbalanced covariates, but that’s ok because randomisation converts “imbalance into error”. Yet, the covariates may be unknown, and without taking measurements prior to the treatment, how are we supposed to know whether the treatment itself or just membership of the treatment group explains the group differences?
Had we not tested the samples prior to treatment, the result would be what many would call “confounded” by the differences in the samples prior to treatment.
From https://en.wikipedia.org/wiki/Confounding#Decreasing_the_pot..., please note the use of the word:
The best available defense against the possibility of spurious results due to confounding is often to dispense with efforts at stratification and instead conduct a randomized study of a sufficiently large sample taken as a whole, such that all potential confounding variables (known and unknown) will be distributed by chance across all study groups and hence will be uncorrelated with the binary variable for inclusion/exclusion in any group.
I always thought that this expression used the word "imply" in its mathematical sense, i.e. "correlation is not a sufficient condition for causation". Did I get that wrong, or did the author?
Just false. It's kinda sad and funny, I guess, that this article exists on a psychology site.
The outcome can be caused by uncontrolled factors, and randomisation is a very limited technique for causal control.
Eg., suppose traits T1...Tn are distributed throughout a population in a powerlaw fashion (eg., Zipf), so that Ti is had by (1/2^i)%. Suppose each has an effect on the relevant outcome of some non-trivial %.
Then even for tens of thousands of people, you can still get "random" divisions of a population where, say, one half has all the T9 traits -- and thereby the "effect" is just due to this random segmentation.
In the hard sciences causes can actually be controlled, by eg., literally placing your hand on some part of the experiment to stop it moving (or equivalent). This is the "third option" which is missing from the intro: actual science.
When causes cannot be controlled, all inferences are highly provisional (, defeasible, suspect, ...). And i'd prefer we used a whole different set of methods, terminology, etc. in this case -- "speculative science" or some such
Science becomes a half-empirical speculative activity in all other cases. So yes, a lot of genetics (but not all), and so on, is a speculative science. Eg., very rarely some finite number of known genes can be given a known causal mechanism (etc.) -- in this cases, you have a Science.
When geneticists are studying eg., single-gene-to-single-disease relationships then this is science. When they're studying possible trajectories of 1000s of genes on down-stream phenomneon with an amazing number of uncontrollable causes.. then i'd be inclined to call this pseduoscience.
The line isn't arbitrary at all, but based on hard facts about the nature of reality and our ability to measure it.
Modelling the weather is a science if it is done based on extremely recent history, and the models are explanatory and accurate to within a relevant time horizon. If you use the same models to predict the weather next year, that's pseduoscience.
Pseudoscience is often deeply apparently scientific -- resuing all the same statsitical formulae, etc. -- this is a charade. It is reality which decides that these techniques are broken, not the techniques themselves.
So we're in a very very bad situation in these cases. It isnt that mere speculation is invovled as an input, its that mere speculation is the output.
Sure, and you’re welcome to have personal preferences. My point is just that “hard science” sounds like an objective term but it’s neither well-defined nor objective. Calling it “specialties which mjburgess likes” would be just meaningful.
I guess, there is rather little hard science done in the world then/a lot of science is speculative in your view (maybe only most of mathematics would survive if it were a science).
Mathematics isnt a science at all, since its a study of number not of cause: there is no inferential gap between measures and their causes in mathematics, because there are no measures.
> operationalizable theory of causal mechanisms
A theory whose terms refer, at some point, to measurable variables that can in some contexts be controlled. So, eg., you might have a theory of light which tells you necessarily how to adjust for some lighting effect in some experiment, even if you cannot control it in that experiment.
The key relation that Science has is Necessity, since it is a study of cause. If your "science" has no means of obtaining necessary relata between physical objects, then it's not a science.
This is perfectly fine. And it does make science in some minimal literal sense "speculative".
My issue is with an extreme depature from this method. Where you open a psychology textbook and there are no causal theories, no quantification of the number and nature of causes, and hence no way of designing experiments with controls, and so on.
This kinds of activity may, at some great distance, resemble each other -- but one is capable of reliably producing knowldge (even if in small quantity), and the other is not. That latter mechanism produces, imv, at least as much wholesale nonesense.
That's not true. For example in researching prime numbers, mathematicians carry out simulations and produce observational results from those simulations, without having a fully-proven formal theory to support the results.
In the example you gave, a test is going to have very low power because of the important factor with huge variance. If that factor is observed, you can create pairs of units with that factor identical within the pair, then randomly assign treatment to one unit in each pair.
In the scenario you're describing, this other factor drowns out any influence the treatment has on the outcome. You'll struggle to get a statistically significant result (low power) and the confidence interval on the treatment effect will include 0. This too can be a valuable finding: sometimes the answer is that the treatment is not particularly effective.
What a snide mischaracterization of a word that essentially means "methodology"
Given there's a tremendous amount of reputational damage that has been done to science by those who have omit from their practice of science, science, I don't have much patience for this omission.
If one wants to educate an informed reader on the scientific method, you ought begin with a setup of the "problem of science" (that of causes, effects and their controls) that makes it clear that these far less reliable methods are indeed, far less reliable.
What this article does, instead, is claim the opposite. It omits the ideal case where science is possible, then proceeds to claim a status for randomisation (as a method) far above what it's capable of --
Just because you prefer "hard science" because it's easier to control variables doesn't give you license to push your pet definition (notably, not provided) and value-judgements about the word onto other people. (Or at least—taking this license destroys your own credibility.) Doing so does just as much reputational damage to the aforementioned institutions, processes, methodologies, and cultures than people who try to draw too much certainty from poorly controlled variables.
What ever happened to nuance and understanding? C'mon! I believe you're capable of better. This kind of rancid tone has no place in serious discussion.
You have chosen a very high-variance distribution for the trait, so in this experiment also the sampling error would also have very high variance - big enough to capture this effect you are talking about. The article mentions this. Random sampling does not guarantee we don't get a false positive, but it lets us quantify the probability of a false positive, and pick an acceptable risk of false positive.
This probability depends on the variance of the underlying trait: a high-variance trait takes an impossibly large sample to get the same discriminatory strength as a sample of a low-variance trait.
> In the hard sciences causes can actually be controlled, by eg., literally placing your hand on some part of the experiment to stop it moving (or equivalent).
Things can be controlled in softer sciences also, it's just that it's sometimes hard to agree on which the meaningful causal antecedents are to get results that replicate out of sample. We can run into that case in hard sciences also, and we do from time to time – the typical example is how insufficient control of unknown factors caused Newton to believe relativistic effects didn't exist.
I fail to see how your "third option" is fundamentally different from the two options already specified in the article.
You don't see a difference between actually necessarily controlling a cause by causal intervention, vs., "somehow, hopefully, on average" possible relevant causes are controlled for?
One way is with modeling assumptions. Recall that the T test was invented to be used in controlled beer fermentation experiments, which I think should fall into your "hard" science category. But the T test has strict distributional assumptions, without which the test is not valid and its results are meaningless.
Another way is that we can improve our modeling assumptions, or at least our interpretation of modeled results, by doing basic descriptive data analysis before experimenting. Unless you are extremely starved for data, you will very quickly notice if your data follows something like a power law distribution. Then you can adjust your experimentation and modeling accordingly, even if it's just tempering your ability to draw conclusions for a randomized controlled trial.
Yet another way is with replication of experiments. We are not talking about observational data here, so these experiments can be replicated. A power-law distribution, again, would be noticeable here, and we'd see a high variance in results across individual experiments.
> You don't see a difference between actually necessarily controlling a cause by causal intervention, vs., "somehow, hopefully, on average" possible relevant causes are controlled for?
I don't, because "actually necessarily controlling a cause by causal intervention" does not exist in real life. It's just a matter of how much the data varies around the average case, which tends to be greater in the social sciences than in the natural sciences... on average, with plenty of variation around that average.
This complaint reveals a lack of understanding about statistics and probability. The point is that in a well-controlled experiment, the random variation is uncorrelated with the treatment. Nobody ever claimed, or ever would claim, that control removes random variation entirely, or that it's impossible to randomly obtain a data sample that looks like a causal effect.
Trying to deal with this problem is literally why the entire field of probability and statistics exists. It is not at all a valid criticism of this article or the causal reasoning behind randomized controlled trials. The article even talks about this (indirectly and clumsily) in the section about "balancing covariates".
Recall the definition of a p-value: the (estimated) probability that a result as extreme as the one we observed could arise simply due to random variation in sampling, measurement noise, etc?
> In the hard sciences causes can actually be controlled, by eg., literally placing your hand on some part of the experiment to stop it moving (or equivalent). This is the "third option" which is missing from the intro: actual science.
No, you still need to do statistical analysis of "hard science" lab experiments too, for exactly the same reason. Experimental control is never perfect, and literally every observed data point ever is a random draw from a noisy data-generating process.
> speculative science
There is a lot to criticize in how science, especially (but not exclusively) social science, handles statistical rigor and the communication of results that depend on statistical analysis. But this is not that.
The issue is that we have T1..Tn in an individual, so there's a very large number of ways you can get one group to have confounders.
The role of the powerlaw is to imply that the generative process which distributes these traits isnt "nice", so that one group can easily get a T9 that the other group doesnt have, and so on, for all T1...Tn
So you have this, let's say adversarial, background generative process which is giving you these confounding traits but never enough of each that you get nice mixtures.
You could see it as a problem of uniform sampling across many powerlaw factors to deliver uniform distributions of those factors. I havent written a simulation, but I don't see why this wouldnt be a serious problem for randomisation.