Common statistical tests are linear models
lindeloev.github.io
lindeloev.github.io
Sadly enough this was true for so many other parts of my education.
Another thing that I wish was more widely known is that a linear model is linear in its parameters, not the data. You can apply arbitrary transformations to the data and still have a linear model as long as what you're fitting is of the form Ey = \beta_0 + \beta_1 f_1(X) + \beta_2 f_2(X) + ... + \beta_p f_p(X).
Now if we train a linear model on the new expanded set of features, it's linear in those features. It's not linear in the original data though, because of the new features that we introduced: x_{n+1} and x_{n+2}.
https://www.amazon.com/Cartoon-Guide-Statistics-Larry-Gonick...
In the back of Statistical Inference by Casella and Berger, there is a great chart making similar connections between the most common statistical distributions.
There is a lot of rhyming in statistics.
This must be what it felt to learn multiplication with Roman numerals.
Regress with a dummy means find a set of coefficients which minimize the sum of squared errors between a column (or columns) of a dependent variable from linear combinations of independent variables. The "dummy" is a 1/0 flag identifying whether the observation row belongs to a category, e.g. gender (assuming a binary classification) with Female mapped to 1. The coefficient that results from this algorithm will be the effect difference for a Female, and the intercept will represent the Male group. Simple model: dependent variable is wage, independent variable is years worked, and dummy represents female. The model is then
Wage = b + m * years worked + c * 1(gender = F).
The intercept term b (0 years worked) will be the average male wage with 0 years worked. b+c, the coefficient for female, will be the average female wage for 0 years worked.
With the regression, you get some nice things out like confidence intervals around your coefficient estimates and more. So you can test hypotheses like whether their is a wage gap (c=0) or how large the wage gap might be (c=any number).
Obviously a canned example, but this helps to introduce all but your last question. You can also test whether there is a slope effect. That is called an interaction term: you add an independent variable to the model that multiplies the gender coded variable to years worked. This would then show, if significant, that you can rule out a hypothesis that there is no wage differential between males and females for years worked (i.e. not just two parallel lines, but different slopes entirely).
Bayesian methods should be preferred as the default.
You know that it's an underestimate because you have prior knowledge about cancer incidence. Bayesian methods let you incorporate that knowledge into estimation process, pulling the estimate up towards a more realistic value.
Undefined variance isn’t nice.
Bayesian treat parameter as a distribution.
The point estimate is base on sample space where as the parameter distribution is base the parameter space.
I think learning both is good and people who pit those two school of statistic against each other are a bit too zealot. They're both tools and use them as needed and when one is easier than the other.
It's much more constructive (and compelling) to call attention to specific problems with how frequentist methods get used in practice, and talk about how Bayesian methods can help with that problem.
For example, here's Andrew Gelman talking about multiple comparison bias: https://statmodeling.stat.columbia.edu/2016/08/22/bayesian-i...
To transition the next generation of scientists to Bayesian will require universities/stats depts/other depts to do so. Have you seen this to be the case right now? Because in my experience of a few different research settings, almost no one I know, of around 50+ researchers, except those passionate about statistics (very rare person indeed) is using Bayesian methodology.
How to improve Bayesian knowledge in the world? This has its own set of challenges, as Bayesian thinking is arguably more mathematically challenging for most numerophobic people. IMO, it will take multiple generations of effort to transition everyone over.
Many "frequentist" methods can be rephrased as heavily-simplified special cases of the Bayesian approach. Most easily, by assuming a flat prior distribution for the relevant parameters (which is basically an artifact of the parameterization you choose anyway) you can assert that any MLE-based approach will yield a correct Bayesian posterior mode, which in turn is an optimal Bayes-estimator, assuming a constant loss function.
The limits of this whole approach are fairly clear, of course - for one thing, Bayesian stats is generally based on working with the entire posterior distribution, not merely a point estimate of it - and there are good reasons for this. But the basic point stands, and many "tweaks" on the basic frequentist approach can in turn be justified in Bayesian terms. This is not to say that frequentist statistics is all that we'll ever need, but merely pointing out that this whole argument of "we should work with what we have, and focus on making sure that frequentist approaches are put to good use" actually sounds rather vacuous. It's correct in a very limited sense, and hardly something that Bayes proponents are unaware of!
Anecdotally, I can definitely say that the vast majority of American physical scientists I've talked to about this also fell prey to this. With the inverse probability fallacy, one attaches a Bayesian interpretation to P-values, so I don't think training scientists in Bayesian stats would be harder than getting them to use frequentist stats correctly (which is a minefield in comparison).
[0] https://www.researchgate.net/publication/280580018_Interpret...
Do you have any evidence for this? My gut feeling is representativeness is a much bigger problem. Like you do a study on people of the same age race class during the same zeitgeist and then generalize to an eternal law.
The next year fashion has changed and the replication comes out all different.
- An automated test suite verifying the (near) equivalences.
- Might be fun for someone to provide parallel implementations in other linear modeling frameworks (scikit-learn, julia, etc).
I absolutely loved this article, and am happy to see it get airtime on HN.
https://news.ycombinator.com/newsfaq.html
This is what's going on. The previous submission didn't get any attention (votes, discussion) so the repost was allowed after a certain time threshold (about 8 hours).