What are the most important statistical ideas of the past 50 years?
tandfonline.com
tandfonline.com
When I was a student in the 1990s, I was taught about hypothesis testing (and all the hassle of p-fishing etc.), and about Bayesian inference (which is lovely, until you have to invent priors over the model space -- e.g. a prior over neural network architectures). These are both systems that tie themselves in epistemological knots when trying to answer the simple question "What model shall I use?"
Holdout set validation is such a clean simple idea, and so easy to use (as long as you have big data), and it does away with all the frequentist and Bayesian tangle, which is why it's so widespread in ML nowadays.
It also aligns statistical inference with Popper's idea of scientific falsifiability -- scientists test their models against a new experimental data, data scientists can test their model against qualitatively different holdout sets. (Just make sure you don't get your holdout set by shuffling, since that's not what Popper would call a "genuine risky validation".)
The article mentions Breiman's "alternative view of the foundations of statistics based on prediction rather than modeling". Breiman does make a big deal of evaluation on holdout sets; but his "prediction" idea isn't general enough, since it doesn't accommodate generative modelling (e.g. GPT, GANs). I think it's better to frame ML in terms of "evaluating model fit on a holdout set", since that accommodates both predictive and generative modelling.
Maybe we could talk about two cardinalities of "big" data. The first is when you can afford not to use all of your data for training. The second is when you can usefully fit highly overparameterized models.
If we know exactly the domain where our model will be used, and it matches exactly the disposition of the training data, then shuffling is fine. But if we want our model to work well in new domains, then we're like scientists looking for generalizable laws, and the best way to test this is by testing in unseen domains. It's often hard to get hold of data with this domain-diversity, which is why a lot of ML models are fragile.
For text, if you take a random sample of shuffled sentences then you'll get a different outcome than if you shuffle whole documents, because it'll be "cheating" as your test sentences will always have some relevant context in your training set, which won't be the case for real unseen data which will sometimes introduce totally new things. And if you train something on data gathered 2015-2020 and evaluate on data from 2021, then you'll likely get worse (but more informative!) measurements than training on a random sample chosen from the whole 2015-2021 range, simply because there are major world events and 'distribution shifts' over time.
IMHO a true test for generalization of many approaches would be to train stuff on data before 2020 and test on data from 2020 - to see how well it generalizes given a major event like Covid pandemic that changes all kinds of aspects everywhere in society that generates the new data.
I think the best way is to choose a minimum timespan you want to be able to predict on, train on that and predict on the future, then retrain, including that future and predict on the next future.
If your time series allow it (e.g. contains a bunch of independent target groups) you can train/test on a subset and leave a another subset for validation
Tied 1st place:
• Markov chain Monte Carlo (MCMC) and the Metropolis–Hastings algorithm
• Hidden Markov Models and the Viterbi algorithm for most probable sequence in linear time
• Vapnik–Chervonenkis theory of statistical learning (Vladimir Naumovich Vapnik & Alexey Chervonenkis) and SVMs
4th place:
• Edwin Jaynes: maximum entropy for constructing priors (borderline: 1957)
Honorable mentions:
• Breiman et al.'s CART (Classification and Regression Trees) algorithm (and Quinlan's C5.0 extension)
• Box–Jenkins method (autoregressive moving average (ARMA) / autoregressive integrated moving average (ARIMA) to find the best fit of a time-series model to past values of a time series)
(The beginning of the 20th century was much more fertile in comparison - Kolmogorov, Fisher, Gosset, Aitken, Cox, de Finetti, Kullback, the Pearsons, Spearman etc.)
Today, computation is easy and cheap. I wonder if we could learn stats in a different way, perhaps starting by just playing with data, and simulated random numbers, graphing, and so forth. Can something like null hypothesis testing be taught primarily through bootstrapping, with the formulas introduced as an aside?
Yes I know that statistics is a formal branch of math, with theorems and proofs. I took that class. But the students who take "stats for scientists" don't ever see that side of it. Understanding the formulas without seeing the proofs is like trying to learn freshman physics without calculus.
The book is called Regression and Other Stories.
Statistics+optimization+dynamic systems+linear algebra
Kernel methods are linear methods on data projected into very-high-dimensional spaces, and you get basically all the benefits of linear methods (convexity, access to analytical techniques/manipulations, etc.) while being much more computationally tractable and data-efficient than a naive approach. Maximum mean discrepancy (MMD) is a particularly shiny result from the last few years.
The tradeoff is that you must use an adequate kernel for whatever procedure you intend, and these can sometimes have sneaky pitfalls. A crass example would be the relative failure of tSNE and similar kernel-based visualization tools: in the case of tSNE the Cauchy kernel's tails are extremely fat which ends up degrading the representation of intra- vs inter-cluster distances.
You have never used Gaussian processes?
[2] Chris Molnar's interpretable learning book has chapters on Shapeley values and SHAP. If you'd prefer text instead of video.
[1] https://www.youtube.com/watch?v=B-c8tIgchu0
[2] https://christophm.github.io/interpretable-ml-book/shapley.h...
If you have a relatively small number of features and your features have relatively stable but not necessarily linear partial dependence, then shap can unpack them pretty well.
E.g. value of a house it might not necessarily increase linearly in squarefoot, distance from city centre and other factors so a random-forest/gbt might do better than a linear model without careful feature engineering. Shap is a brilliant tool to unpack the dependencies this model has learned.
There will be other cases where a linear model is adequate, so you can get the explainability just from the coefficients, or where there are very large numbers of features with complicated nested structure. In either case shap is not too useful.
But the 'fit lightgbm and explain it with shap' is a decently pragmatic approach for lots of tabular data machine-learning problems.
If you really need to unpack it, cluster your shaps. To see what kinds of predictions are being made
It is not always a good idea to do that. Always try different methods, there is no ultimate method. At the very least OLS should be tried and some other fully explainable methods, even a simple CART like method.
2) The fact that “substitute the data for the population distribution” both works and is sometimes provably better than other more sensible approaches is a little mind blowing.
Most things called the bootstrap feel like cheating, ie “this part seems hard, let’s do the easiest thing possible instead and hope it works.”
I agree "bootstrap" has expanded a bit in meaning but I think it's basically the same idea.
I used to think it was cheating but have realized there is a cost to it, which is replication. For many things it's just impractical in terms of computation time. So although it's great, it requires a lot.
Sidney Siegel, {\it Nonparametric Statistics for the Behavioral Sciences,\/} McGraw-Hill, New York, 1956.\ \
Right, there was already a good book on such tests over 50 years ago.
Can also justify it with an independence, identically distributed assumption. But a weaker assumption of exchangeability can also work -- I published a paper with that.
The broad idea of such a statistical hypothesis test is to decide on the null hypothesis, null as in no effect (if looking for an effect, then want to reject the null hypothesis of no effect) and to make assumptions to permit calculating the probability of what you observe. If that probability is way too small then reject the null hypothesis and conclude that there was an effect. Right, it's fishy.
Siegel, S., & Castellan, N. J. (1988). Nonparametric statistics for the behavioral sciences (2nd ed.) New York: McGraw-Hill.
Counterfactual Causal Inference
The formalisation of casual inference by Pearl is in my opinion a development comparable to the invention of calculus in mathematics: Suddenly the problem space solvable with statistics became at least an order of magnitude larger and we're only starting to see the benefits in other sciences.
Hopefully one day I'll be able to make use of it.
Well, it's an heuristic ;)
The idea is that if you ever find a distribution that seems to be normal, you look for the finite variance variables that are lurking behind it. If you do not find these variables, then your distribution was not likely normal.
This strategy is possible because p-values are themselves stochastic and a researcher will find one significant p-value for every 20 models that they run (at least on average).
p-hacking could also refer to pushing a p-value close to the significant cut-off (usually 0.05) by modifying the statistical model slightly until the desired result is achieved. This process usually involves the inclusion of control variables that are not really related to the outcome but that will change the standard errors/p-values.
Another way to p-hack is to drop specific observations until the desired p-value is reached. This process usually involves removing participants from a sample for a seemingly legitimate reason until the desired p-value is achieved. Usually identifying and eliminating a few high leverage observations is enough to change the significance level of a point estimate.
Multiple strategies to address p-hacking have been proposed and discussed. One of the most popular ones is pre-registration of research designs and models. The idea here is that a researcher would publish their research design and models before conducting the experiment and they will report only the results from the pre-registered models. This process eliminates the "fishing expedition" nature of p-hacking.
Other strategies involve better research designs that are not sensitive to model respecification. These are usually experimental and quasi-experimental methods that leverage an external source of variation (external to both the researcher and the studied system, like random assignment to conditions) to isolate the relationship between two variables.
Moreover, the lab pushed me out and decided to use my work anyway. Specifically, they requested that anybody that used the software I wrote give them authorship on their papers.
I'm happily employed as a software engineer now.
(also makes you wonder what if 20 groups study this phenomena, each restiricting itself to one color)
It was a shock to the world when it was shown that least squares methods (a method used since Gauss) are actually not optimal in most cases.
What makes meta-analysis special within a multilevel framework is that you know the level 1 variance. This creates a special case of a generalized multilevel model where you leverage your knowledge of L1 mean and variance (from each individual study's results) to estimate the possible mean and variance of the population effect.
The population mean and variance is usually presented in funnel plots where you can see the expected distribution of effect sizes/point estimates given a sample size/standard error.
Researchers have also started to plot actual point estimates from published papers in this plot, showing that most of the published results are "inside the funnel", a result that is usually cited as evidence of publication bias. In other words, the missing studies end up in researchers' file drawers instead of being published somewhere.
"Due to reduced superstition, better education, and general awareness & progress, humans who are neither meticulous statistics experts, nor working in very constrained & repetitive circumstances, will understand and apply statistics more objectively and correctly than they generally have in the past."
Sadly, this idea is still wrong.
We barely have any concept of how education has penetrated our society, to a far greater degree, now than fifty years ago.
We as a population are FAR more educated, far less ignorant, and generally speaking extraordinarily better off in our daily lives than in the past.
It's only because the phenomenon of social media, which has crowd sourced ignorance and hurled it in front of our eyes, that we perceive we're in an unenlighted age/trajectory.
I've heard that Google and Baidu essentially started at the same time, with the same algorithm discovery (PageRank). Maybe someone can comment on if there was idea sharing or if both teams derived it independently.
Implementing it and catching edge cases isn’t trivial
(I do not want to link directly to the pdf shown in the search result). Section 2.1 deals with related work: "There has been a great deal of work on academic citation analysis [Gar95]. Go man [Gof71] has published an interesting theory of how information flow in a scienti c community is an epidemic process......" (and more)
I think that paper is worth a read.
> In 1996, while at IDD, Li created the Rankdex site-scoring algorithm for search engine page ranking, which was awarded a U.S. patent. It was the first search engine that used hyperlinks to measure the quality of websites it was indexing, predating the very similar algorithm patent filed by Google two years later in 1998.
The idea came first up in the 70's. https://www.sciencedirect.com/science/article/abs/pii/030645... and several times afterward before PageRank was developed.
What I find very interesting about PageRank is how you can trade accuracy for performance. The traditional way of calculating PageRank by means of squaring a matrix iteratively until it reaches convergence gives you correct results but is sloooooow. For a modestly sized graph it could take days. But if accuracy isn't that important you can use Monte Carlo simulation and get most of the PageRank correct in a fraction of the time of the iterative method. It's also easy to parallelize.
Jon M. Kleinberg, "Authoritative sources in a hyperlinked environment," 1998, Proc. Of the 9th Annual ACM-SIAM Symposium on Discrete Algorithms, pp. 668-677.
That being said, Page Rank is a more a stellar example of adapting an academic idea into practice, than a statistical idea in and of itself.
Afterall, it is 'merely' the stationary distribution for a random walk over an undirected graph. I say 'merely' with a lot of respect, because the best ideas often feel simple in hindsight. But, it is that simplicity that makes them even more impressive.
I don't think this means much. The history of science and technology is full of examples of results named after someone other than the first person to find them.
https://en.wikipedia.org/wiki/List_of_examples_of_Stigler%27....
In fact, based on the other comments in this thread, it seems that Pagerank being named after Larry Page is itself one of these examples.
Not attacking the mathematically field of statistics just pointing out that lots of people abuse statistics in an attempt to get people to behave as they would prefer.