Introduction to Modern Statistics
openintro-ims2.netlify.app
openintro-ims2.netlify.app
For the same approach in a slightly different context see [3].
[1]: https://openintro-ims2.netlify.app/11-foundations-randomizat...
[2]: https://en.wikipedia.org/wiki/Permutation_test
[3]: https://inferentialthinking.com/chapters/11/1/Assessing_a_Mo...
I think that still the best way to understand statistics is to start with the mathematical theory and to grind 1000+ textbook problems.
Are there any books you'd recommend for this approach?
I also liked "In all Likelihood" by Pawitan for a "likelihoodist" foundational approach.
Later I re-learned it in an applied way, and things made much more sense to me.
(That's fine if you liked learning it your way, but that's no reason to be skeptical about other approaches.)
If you study mathematical statistics, it is not taught as a cookbook. At the elementary level you learn probability theory and distribution theory, all the different distributions, hypothesis tests, regression, ANOVA and so on proceed from there. Meanwhile, I think research scientists are often taught statistics as a set of recipes because its usually a short course for a specific discipline. E.g. Statistics for biologists.
In R, and python::statsmodels you get the answer to (essentially) an ANOVA any time you run an LM or GLM; its the Z-statistic for your whole model.
I know there is more nuance to this, but teaching students that they can use regression for most of the problems they would have used seemingly arcane tests for is going to be much more useful for the students.
Here is a lovely page demonstrating how to do this in R: https://lindeloev.github.io/tests-as-linear/
It does assume you already have a solid understanding of calculus and combinatorics, though. Which I think is fair. Discrete statistics is arguably just applied combinatorics, and continuous statistics applied calculus, so if you have a strong foundation in those two subjects then you're already 90% of the way there. (And, if you don't, stop the cart and let the horse catch up.)
IMO, this is where Bayesian Statistics is far superior. There's a Curry-Howard isomorphism to logic which runs extremely deep, and it's possible to introduce using conjugate distributions with nice closed form analytical solutions. Anything more complex, well, that's what computers are for, and there are great ways (STAN) to run complex distributions that are far more intricate than frequentist methods.
Generative models, (implemented in e.g. Stan, PyMC, Pyro, Turing, etc.) split models from inference. So one can switch from maximum likelihood to variational inference or MCMC quite easily.
Generative models, beginning from regression, make a lot more sense to students and yield much more robust inference. Most people I know who publish research articles on a frequent basis do not know p-values are not a measure of effect sizes. This demonstrates current education has failed.
This is an odd way of putting it. I think it's better to say that, given some mostly uncontroversial assumptions, if one is willing to assign real number degrees of belief to uncertain claims, then Bayesian statistical inference is the only way of reasoning about those claims that's compatible with classical propositional logic.
I'm hoping to see, over time, a shift away from ad-hoc null hypothesis testing in favour of linear models (yes, in introductory courses, from the start-- see link below) and Bayesian-by-default approaches.
https://lindeloev.github.io/tests-as-linear/#:~:text=Most%20....
It's here: https://civil.colorado.edu/~balajir/CVEN6833/bayes-resources...
This is the third result from a query I ran:
https://github.com/Booleans/statistical-rethinking/blob/mast...
Main repo w/ some notebooks: https://github.com/Booleans/statistical-rethinking
[0] https://bookdown.org/ajkurz/Statistical_Rethinking_recoded/
It's less comfortable to use Bayesian methods because you have to be explicit about your assumptions as the user, which opens your assumptions up for easier inspection. There's also way less specific information implied by priors than most people think. Informative priors should try to make distinctions between something that's reasonable-ish and something that's essentially infinity (take pharmacokinetics for example, the diffusion velocity of a molecule in your blood stream shouldn't have a velocity near the speed of light in a vacuum should it?). They should not be forcing your model to achieve a particular result. Luckily, because of the need to explicitly state them in a Bayesian analysis, it's much easier to determine if they were properly set.
Prior specification is essentially problem domain-informed regularization where you can actually hope to understand if the hyperparameter is going to work or not.
Is there anything where I can start today, as a guinea pig? My statistics education is basically zero.
It goes over Bayes' Theorem early on, which I assume is Bayesian-by-default. I didn't realize this isn't universal.
I watched it because my professor seemed to be teaching from the same textbook, so it followed the same general course structure. The textbook we used was "Probability and Statistics for Engineering and the Sciences — 9th edition". You can find a PDF of the eighth edition just by googling the name.
You could probably follow along just by taking notes during the video lectures, but if you want to give yourself homework, the textbook provides a lot of practice problems.
Bayesian statistics is not. Most basic intro courses go the frequentist routes for historical reasons, but really both methods have their pros and cons.
But if you ever do a randomized test with a suitable linear model to estimate the efficacy of these two methods, do let us know, that would be 10/10 :)
[1]: https://lindeloev.github.io/tests-as-linear/#41_one_sample_t...
You don't need much beyond basic calculus. Most suffer from some mental block they got installed at a young age akin those that say "I'm bad at math" because their teacher sucked. Dive in and you won't regret it.
I feel stats is has a somewhat similar effect even among those with math education. Several friends who have a degree in math recoil at the first mention of stats concepts.
It is actually even fashionable in non-english countries. Declaring "I'm bad at [my native language], I only use english anyway" makes you a better person somehow. And it's not rare in other areas either – in post-truth world it's trendy not to know things.
Not that he was bad at German, necessarily, but that it was overly complex vs English, so he and his significant other even used English at home when they both didn't have to.
This is just one example, of course.
I was actually surprised to hear him say that (though it may have been to thwart my attempts at speaking some German when we had a meeting. I don't know any German,lol).
In my country it's especially fashionable not to know a native language for a people in tech. "It's impossible to talk about tech anyway in native languages, so we all should use english anyway in future" is very common. I tried to fight with it localizing/translating software for many years, but I've given up for now.
I can learn math fairly well if I have the right written material and the right direction. However, I do not retain math skills: without active practice, I revert back to "how do fractions work?"
For example, I did extremely well in a college algebra course that was partially online (combined with Khan Academy to catch up). I could do my tests perfectly in pen, much to the amusement of the assistants. I could make connections and see the implications and applications of the math. Roughly three to six months later, I was back to forgetting fractions.
I can't learn these things over time, but I can learn them all at once. I'm collecting resources for my next math adventure.
Would you be OK with elaborating on this a bit more?
The injury was diagnosed via SPECT scan. Two of them: one after mental activity, one after physical relaxation. This also revealed that my brain relaxes with mental activity, and "lights up" with physical relaxation.
I always knew there was "something" wrong with me, so I jumped at the opportunity to get a brain scan. The clinic was a little suspect in that they claimed to be able to diagnose mental health issues via the scan, but they combined the scan with more traditional evaluations.
To quote pg, many conversations on the Internet:
Person 1: ∃x P(x)
Person 2: -(∀x P(x))!
Something related: I often feel like I'm doing math when I work with information, especially when drawing out concepts and their relationships. Like 3D Tetris, but with recursion. There are also patterns in categorization. One of my purposes in learning math is to be able to quantify relationships between concepts, and create models, etc. What would I need to know about in this case?
However I will say that, as you learn more math, the intuition and big picture thinking is way more important than details. I normally forget the specifics of how something works, but I know what I’m trying to do and what thing I need to look up to do it. There is very little need to memorize things besides having enough “RAM” for the moment. You always have pen and paper to write stuff down!
Regarding modeling, the sky is kind of the limit. The more math you know, the more abstractions you learn and start to see in the world. Dynamical systems is a good field to look at but I’m biased because it’s my main topic. Despite its usual applications in physical modeling, an algorithm is basically a time-discrete dynamical system: you recursively apply a function to a state and that gives you a new state. I have certainly seen algorithms analyzed from that perspective before.
Another area you might find interesting is Algebra. This isn’t like the algebra you took in school but more about looking at a space of objects that interact via some operation and characterizing what you know about the space given that operation. A classic example is that rotation operations form an algebraic structure known as a “group” due to the fact that any two rotations gives you a third, rotations are invertible (you can cancel a rotation by rotating in the opposite manner), and they follow the associative property. There’s a book I’ve been meaning to read about how to use this kind of algebra to design safer and more intuitive APIs by considering the data structures of the APIs as objects and functions / methods as operators on those objects.
Dynamical systems is a great new keyword! Algorithms are a perfect example. I am exploring system dynamics from a metacybernetic point of view (via the viable system model). My focus is the dynamics of humans in complex systems, how systems change human behavior and vice versa.
Does big picture thinking in math involve understanding/intuition of the implications of formulas, and how they interact with other formulas (or other mathematical entities, I don't know if formula is a general enough term for "group of math actions")?
Graph dynamical systems looks promising, because I draw similar pictures when showing relationships over time. Fractals are always good, drawing fractals taught me to think recursively.
Is this the definition of group in the algebra you mentioned? https://mathworld.wolfram.com/Group.html
I am also interested in the way models create understanding, and how technical models can be visually altered to aid in understanding. I have a background in visual psychology, and see many mistakes that cloud the meaning of what is being communicated.
And in that context, there is a lot of agreement that this approach is fundamentally flawed and outdated. if you are interested, I can provide references when I get to the office. But off the top of my head consider Gigerenzer and Cummings.
[1]: https://pure.mpg.de/rest/items/item_2101336/component/file_2... [2]: Sample at https://tandfbis.s3.amazonaws.com/rt-media/pp/common/sample-...
As for a undergraduate text that 'teaches the difference', you can look at 'An Introduction to Statistics' by Carlson & Winquist.
* Statistical Inference by Casella and Berger. This book has a very good reputation for building statistics from first principles. I won't link to them, but you can find full PDF scans online with a simple search. Amazon reviews: https://www.amazon.com/Statistical-Inference-Roger-Berger/dp...
* Statistics by Freedman, Pisani, and Purves has similarly very good reviews and can be easily found online. Amazon reviews: https://www.amazon.com/Statistics-Fourth-David-Freedman-eboo...
* The majority of the Berkeley data science core curriculum books are online. This is not purely statistics but 1) is taught in a modern style that makes use of computation and randomization and 2) uses tools that may be useful to learn about.
1. https://inferentialthinking.com/chapters/intro.html (Data 8)
2. https://learningds.org/intro.html (Data 100)
3. http://prob140.org/textbook/content/README.html (Data 140)
4. https://data102.org/fa23/resources/#textbooks-from-previous-... (Data 102; this gets into machine learning and pure statistics)
The Berkeley curriculum is not the only one; there are tens, possibly hundreds, of online courses. The Berkeley curriculum is just 1) quite extensive and 2) the one I happened to read the most about when I was recently researching how data science is currently taught.
You could also look at Introduction to Probability by Joseph K. Blitzstein and Jessica Hwang (available for free here: http://probabilitybook.net (redirects to drive)).
Casella’s exercises are absolutely brutal.
Can second Statistical Rethinking though if you have the basics of stats and want to learn it again from a very different, more causal/bayesian point of view.
There can be a rather wide gap between a theoretical approach that you might encounter as taught by a statistician and an applied approach you might encounter in a business statistics or social science statistics course.
Depending on your math background and the area of intended application, in my opinion, it would sway recommendations for a first 'book' on statistics for self-learning.
https://www.thegreatcourses.com/courses/learning-statistics-...
Might be available for free via your local library, too.
https://pubs.aip.org/aip/cip/article-abstract/3/5/106/136800...
Another good book at a more advanced level is this one: "Regression and Other Stories" by Andrew Gelman and Jennifer Hill <http://www.stat.columbia.edu/~gelman/arm/>.
Another good book at an even more advanced level is this one: "Data Analysis Using Regression and Multilevel/Hierarchical Models" by Andrew Gelman and Jennifer Hill <http://www.stat.columbia.edu/~gelman/arm/>.
These books are all very polished productions. What make makes them special is that they emphasis teaching the statistical way of thinking.
The too big error was from openintro.com, where they have a "send to kindle" option which I was hoping was a proper epub
Otherwise you can try building a PDF from the very similar Data 8 book[1] using [2]
Using it truly feels like a 'fresh way' to do statistics. Its main website provides ample use cases, guides and tutorials, and I often return to the blog for the well documented deepdives into how traditional frequentist methods and their bayesian counterparts compare (the animated explainers are especially helpful, and I appreciate the devs reflecting on each release and future directions).
R Studio for the uncompromised R language integration.
Jamovi for the clean interface and simple to use analyses and workflows.
Stenci.la for the multiple language notebook paradigm.
KNIME for the node based interface and workflow productionisation.
PSPP for the 'same but different' SPSS interface and feature set.
Quarto for the 'analysis as publication' approach, reactive notebooks and custom integrations.
Myself I often opt for BigQuery + Javascript stat libraries, which is free insofar you remain within the sandbox mode.
Stats, imv, should be taught simulation-first: code up your hypotheses and see if they're even testable. Many many projects would immediately fail at the research stage.
Next, know that predictions are almost never a good goal. Almost everything is practically unpredictable -- with a near infinite number of relevant causes, uncontrollable.
At best, in ideal cases, you can use stats to model a distribution of predictions and then determine a risk/value across that range. Ie., the goal isnt to predict anything but to prescribe some action (or inference) according to a risk tolerance (risk of error, or financial risk, etc.).
It seems a generation of people have half-learned bits of stats, glued them together, and created widespread 'statistical cargo-cultism'.
The lesson of stats isnt hypothesis testing, but how almost no hypotheses are testable -- and then what do you do
How do you teach any of this to someone who hasn't already taken introductory statistics? How do you learn anything if you first have to learn the myriad ways something you don't even have a basic working knowledge of can fail before you learn it?
To teach this, from scratch, I think is fairly easy -- but there's few with any incentive to do it. Many in academia wouldnt know how, and if they did, would discover that much of their research can be shown a priori to not be worthwhile (rather than after a decade of 'debate').
All you really need is to start with establishing an intuitive understanding of randomness, how apparently highly patterned it is, and so on. Then ask: how easy is it to reproduce an observed pattern with (simulated) randomness?
That question alone, properly supported via basic programming simulations, will take you extremely far. Indeed, the answer to it is often obvious -- a trivial program.
That few ever write such programs shows how the whole edifice of stats education is geared towards confirmation bias.
Before computers, stats was either an extremely mathematical disipline seeking (empirically useless) formula for toy models; or using heuristic empirical formula that rarely applied.
Computers basically obviate all of that. Stats is mostly about counting things and making comparisons -- perfect tasks for machines. with only a few high-school mathematical formula most could derive most useful statistical techniques as simple computer programs.
Bingo. Cargo cult stats all the way down. It’s not just personal interest, it’s the entire field, it’s their colleagues, mentors, and students. Good luck getting somebody to see the light when not just their own income depends on not seeing it, their whole world depends on the “stat recipes” handed down from granny.
The better the alternatives, the more fierce the passion with which they will be rejected by the mainstream.
If every formula is producing bell curves then that’s a failure to educate people. 50d6 vs 50d6 + 1 is easy enough you can include 1d2 * 50 + 50d6 for a 2 tailed distribution, but also significantly different distributions which then fail various tests etc.
I’ve seen people correctly remember the formula for statistical tests from memory and then wildly misapply them. That seems like focusing on the wrong things in an age when such information is at everyone’s fingertips, but understanding of what that information means isn’t.
Sucks, as we seem to have taught everyone that statistical models are somehow unique models that can only be made to get a prediction. To the point that we seem to have hard delineations between "predictive" models and other "models.".
I suspect there are some decent ontologies there. But, at large, I regret that so many won't try to build a model.
Competent stakeholders and decision makers use the uncertainty around predictions, the chances of an outcome that is different from the point-predicted outcome, to come to a decision and the plan includes what the course of action should be should the outcome differ from the prediction.
What are you trying to say here? If there are two normal distributions, both with variance one, one having mean 0 and the other having mean 100, and I get a single sample from one of the distributions, I can guess which distribution it came from with very high confidence. Where did the number 30 come from?
Yeah, I've also heard 30 for normal distributions over and over in ~7 stats courses that I've taken.
This SE stats answer sounds reasonable enough: https://stats.stackexchange.com/a/2542
(Requires some very lax assumptions like finite variance on the underlying distribution)
I want to build intuitions for how these statistical methods even work, at a high level, before getting drowned in math about all the details. And like you say, I want to understand the boundaries: "when statistics will not work; and what it even means when it "works".
I imagine that different methodologies exist on a spectrum, where some give more reliable results, and others are more likely to be noise. I want to understand how to roughly tell the good from the bad, and how to spot common problems.
Just raw hypothesis is just too easy to juke by overwhelming it with trials. Lots of research papers have "statistically significant" results, but give no mention of how many experiments it took to get them, or any indiciation of negative results. Eventually, there will always be the analysis where you incorrectly reject the null hypothsis given enough effort.
Do you have anywhere I can read more about this? I would have assumed that a trillion data points would be sufficient to compare any two real-world distributions
I don't agree with this as a statement of fact (except in the obvious case of two power-law distributions with extremely close parameters). Supposing it was true, that would mean that you would almost never have to actually worry about the parameter, because unless your dataset is that large one power law is about as good as any other for describing your data.
I do genetic epidemiology (which is considerably more compute intensive than regular epidemiology), and R is still the most common language, with the most libraries and packages being used for it, compared to python for example.
I think maybe you should consider being less forthcoming with your opinions on topics which you are not well informed on.
On the other hand, R has too many obscure options for what I can find in scipy or sklearn. So I find it easier to just jump into sklearn, use the very nice unified interface "pipelines" to churn through a whole bunch of different estimators without having to do any munging on my data.
So I think it just depends on your field. But R seems to stick more with academia.
By contrast, from first downloading R to running my first R script took about 1 hour (the most difficult part was opening the 'script' pane in RStudio IDE, which doesn't open by default on new installations, for some reason).
There's huge demand out there for statistical software that's accessible to people whose primary pursuit is not programming/cs, but genetics, bioinformatics, economics, ecology and other disciplines that necessitate tooling much more powerful than excel, but with barriers to entry not much greater than excel. R is a fairly amazing fit for those folks.
Try installing an older version of a package without it pulling in the most recent incompatible dependencies, it's a whole adventure.
R sucks as a language but it excels at that specific application, just because of its tremendous ecosystem (putting even python to shame in some niche areas).
Not all non-typed languages are bad. Clojure, for example, is one if the most elegant languages I’ve worked with (despite my dislike of the JVM).