3,775 karma · joined June 17, 2021
Some causes for accelerated aging seem relatively direct and plausible with causal models that have supporting literature, i.e. air quality.
On the societal level, it's much more complicated. For example, there will be an immense number of paths how education affect aging, some positive and some negative.
I wonder how much of those effects boil down to a few highly influential (unobserved?) covariates, such as physical activity, drug consumption and crime rate.
[edit] by the way, look at the replication materials - I've rarely seen such clean code. Kudos!
https://github.com/euroladbrainlat/Biobehavioral-age-gaps/bl...
To solve the question of whether or not these harms can/will actually materialize, we would need causal attribution, something that is really hard to do — in particular with all involved actors actively monitoring society and reacting to new research.
Personally, I think that transparency measures and tools that help civic society (and researchers) better understand what's going on are the most promising tool here.
Bayesianism requires you to assume / formalize your prior belief about the subject under investigation and updates it given some data, resulting in a posterior belief distribution. It thus does not have the clear distinctions of frequentism, but that can also be considered an advantage.
[1] https://web.mit.edu/hackl/www/lab/turkshop/readings/gigerenz...
That said, I think your take is also empirically supported. There is this [1] very interesting study which comes to the same conclusion. It uses broadcast range of radio towers to do a quantitative analysis on the potential effects and finds few. Interestingly enough, I have seen other studies with similar designs that do show persistent effects of exposure to broadcasts, so I’m favorable to the idea that this one really is a valid null finding.
[1] https://www.ushmm.org/m/pdfs/20100423-atrauss-rtlm-radio-hat...
Linear regression, for all its faults, forces you to be very selective about parameters that you believe to be meaningful, and offers trivial tools to validate the fit (i.e. even residuals, or posterior predictive simulations if you want to be fancy).
ML and beyond, on the other hand, throws you in a whirl of hyperparameters that you no longer understand and which traps even clever people in overfitting that they don't understand.
Obligatory xkcd: https://xkcd.com/1838/
So a better critique, in my view, would be something that the JW Tukey wrote in his famous 1962 paper: (paraphrasing because I'm lazy):
"better to have an approximate answer to a precise question rather than an answer to an approximate question, which can always be made arbitrarily precise".
So our problem is not the tools, it's that we fool ourselves by applying the tools to the wrong problems because they are easier.
If you have a fraud model, just show the model and the data and the validation - everything else is marketing fluff.
Given they have enough data, at some point it's perfectly reasonable to have cars with 5+ crashes per 12 months - just because of chance.
This is exactly why statistics was invented, damnit!
Do you have some sources for Sociology being hit harder than, say, Psychology, in the replication crisis? I personally don't even know of a many labs replication attempt from them (Sociology). So we probably don't know.
That said, Sociology is the attempt to explain the most complex thing in the universe (collective behavior arising out of sentient individuals), so if you think there's a better way, I'd be extremely excited :)
Personally I find this the easiest simple mental model for reflexive modernity:
"Modernity linearly accelerated change. Reflexive modernity loops back on itself and accelerates the acceleration". So a bit of "singularity is near".
The whole theory itself has quite a bit more baggage, but again I agree, the wikipedia article does not help at all.
Sewing? The course will explain and show you, after which you gain experience by practice. Making financial decisions? This rests a lot on knowing facts about how financials work, and a course is absolutely one correct way of learning those.
For experience, I'd rather think of things like - how do you cope with the death of a loved one? How do you decide who to be or what to do? How do you manage your emotions, streghts and weaknesses?
- bayesian methods give you posterior distributions rather than point estimates and SEs
- bayesian methods natively offer prior and posterior predictive checks
- with bayesian methods, it's evidently easier to combine knowledge from multiple sources, which null-hypothesis testing struggles with (best way is probably still meta-analyses)
- No colliders have been included in the analysis, which would introduce appearance of causality that does not exist
There is also this great book on causality in ML, but it's a much heavier read:
Chernozhukov, V., Hansen, C., Kallus, N., Spindler, M., & Syrgkanis, V. (2025). Causal Inference with ML and AI.
If you have a DAG based on wrong assumptions, it doesn't matter whether you get a point estimate based on null hypothesis thinking or whether you get a posterior distribution based on some prior. The problem is that the way in which you combine variables is wrong, and bayesian analysis will just be more detailed and precise in being wrong.
Causality is a largely orthogonal problem to frequentist/bayesian - it makes everything harder, not just one of those!
Most things we learn about DAGs and causality are frustrating, but simulating a DAG (e.g. with lavaan in R) is a technique that actually helps in understanding when and how those assumptions make sense. That's (to me) a key part of making causality productive.
For some areas of research, truly understanding causality is essentially impossible - if well-controlled experiments are impossible and the list of possible colliders and confounders is unknowable.
The key problem is that any causal relation can be an illusion caused by some other, unobserved relation!
This means that in order to show fully valid causal effect estimates, we need to
- measure precisely
- measure all relevant variables
- actively NOT measure all harmful (i.e. falsely correlated) variables
I heartily recommend the book of why [1] by Pearl and Mackenzie for a deeper reading and the "haunted DAG" in McElreath's wonderful Statistical Rethinking.