Statistical Computing for Scientists and Engineers
zabaras.com
zabaras.com
https://github.com/melling/MathAndScienceNotes/tree/master/s...
I haven’t listened to the Notre Dame course yet. How would others rate it?
I have gone through the UC Berkeley Irvine 131a class. I have high-level notes:
https://github.com/melling/MathAndScienceNotes/blob/master/s...
and detailed notes in PDF for the first four classes;
https://github.com/melling/MathAndScienceNotes/blob/master/s...
Actually, wrote this up in a blog yesterday:
https://h4labs.wordpress.com/2017/12/30/learning-probability...
The syllabus could almost be from an inverse theory class in my field (geology), albeit one with more focus on the underlying mathematics. I don't think it's trying to do too much, it's just not trying to be a intro course.
> It's not expected that this is the first time someone in the course has come across most of the concepts and methods.
It is just cramming at least 4-5 stats classes in there that's all.
Unless you get those right combination it'll be your first time anyway. There are multivariate stat, 2 courses of Bayesian ( 1 tradition and 1 nonparametric), 1 course of comp stat, and whatever the heck else I've haven't encounter.
Bayesian isn't even taught in my grad program at all. I had 2 months to learn Bayesian.
I disagree with the not expected to be first time and maaaaybe it's just to get your feet wet.
Considering the stuff the professor is going over it seem more sink or swim.
That seems impossibly out of step with the world today. How can they explain something as simple as a multi-armed bandit from that point of view?
At 3:40 - For this topic he actually has has 6 or 7 lectures, 100 pages of notes. But it's compressed to an introduction.
At 7:10 - He states HMM is 2-3 lectures, but he's going to compress it to half a lecture.
Honestly, I had courses like this in grad school. The were typically seminars and usually graded on a curve because people crammed stuff into their brains as quickly as possible and barely understood anything! They were meant to give you a broad coverage of a field, not comprehensively cover any particular set of topics.
Perhaps some topics are simply too big to really teach effectively through lectures. You need to go digging on your own to understand all the myriad details, and put in the hours alone with a textbook.
At universities that offer PhDs, being a professor forces you to develop a strong theoretical foundation, which makes it easier to grasp the theory of related subjects. Example: If you have a deep understanding of math/statistics theory, it's much easier to understand a paper on machine learning, even if you're not a computer science professor.
Professors at research oriented universities supervise PhD research students and each student tends to be a multiplier on their supervisor's knowledge since each PhD is a collection and extension of an existing research area. A good PhD student is basically a massive funnel of information.
I'll also say, don't ascribe more abilities or skills to a professor than what you see in the presentation. Sometimes having a strong theoretical orientation/experience comes at the expense of a practical one. Don't assume that having a strong theoretical foundation in Stats/ML automatically makes you a good Data Scientist or ML Engineer.
Edit: Let the down-voting begin
I can assure you they aren’t trying to make the math look scary - to be a bit blunt, a comment like that’s usually a sign that someone hasn’t made a serious attempt to engage with a field. That sort of bad faith regarding a core discipline of the mathematical sciences probably should be downvoted.
For example, in some areas of research, once you have the right instrument take measurements, you can just plot the results, and your measurements (or some transformation of them) will show up as a linear function of whatever you manipulated.
But really, I'm not totally sure what you mean, since in any situation where there's uncertainty, it makes sense to me that you'd want to try to capture that uncertainty in your analysis, and that's statistics. Make those analyses more complex (e.g. taking measurements from a field with obvious spatial dependence between measures), and the models become more complex too.
After n tries (as the data from the experiments were spread semi-randomly) he drew a line roughly in the middle of the diagram and called the whole thing a "shotgun diagram" (or something similar).
When he got back to class, he was surprised by the results of his mates, that more or less led to a neat line.
Then the Professor gave him the maximum vote, as the experiment was intentionally leading to "senseless" results, and he was the only one in the class to have actually honestly reported the results, while all the ohers had evidently faked or invented them.
This is on a similar note:
Some particular problems:
1) Scary math and pointlessly obscure terminology are indeed a problem. For unnecessarily scary math, the early literature on Dirichlet process mixture models is a good example - almost like they were designed to be incomprehensible to most people who do actually have enough background to use the results. At a lower level, there are pointlessly obscure and misleading terms like "score function", "coefficient of determination", and worst of all, "standard error of the estimate" for the estimated standard deviation of residuals in a regression model (worst since it is not in fact a "standard error" by the general definition of that term).
2) Introductory statistics is generally taught from a naive "frequentist" perspective, because that's been the tradition for the last century or so. The justifications offered in such courses for using p-values and confidence intervals are not defensible - they just sound plausible if you don't know better. There is no good solution, since more sophisticated frequentist arguments will be beyond the level of the course, and shifting to a Bayesian perspective cuts the students off from the scientific literature with p-values, etc. that they will need to be able to read.
3) Outsiders coming into the field often have strange ideas. You might think that physicists capable of building a billion dollar accelerator would be able to recognize when a statistical method they think of is nonsense, but you'd be wrong. There's a tendency for anyone who learns information theory before statistics to think that information theory is tremendously relevant - but no, rephrasing maximum likelihood or Bayesian methods in information theory terms may sometimes be slightly helpful in thinking about them, but doesn't really add anything fundamental. And no, there's nothing particularly special or interesting about distributions that maximize entropy subject to some (generally arbitrarily selected) constraint.
4) There's a tendency to want more than you can get. There is no one "objectively correct" model/prior/analysis for a data set. Subjective assessments are unavoidable. But a lot of people don't want to accept this fact, and devote great efforts to ways of trying to pretend otherwise.
However, if you think statistics is just a simple matter of running a curve through a cloud of points, you're very wrong. Even running a curve through a could of points is a complicated and subtle enough task that deep issues arise, and these issues become much more obvious if you're trying to fit a function of hundreds of variables rather than just one. And if you're trying to not just "fit" data but come to valid conclusions about cause and effect, or about underlying latent variables that will provide useful information in new contexts, then you really do need to know a lot.
And no, there's nothing particularly special or interesting about distributions that maximize entropy subject to some (generally arbitrarily selected) constraint.
Wow, some two heavyweight opinions there. Care to elaborate?
The maximum entropy idea is just wrong (in general), in that there is no good argument for doing it. Actually, it's "not even wrong", since maximizing the entropy subject to the observed values of some expectations is just not possible, since we do not observe expectations, but rather particular finite data sets.
However, there is some truth in this regarding the application by scientists. I did my PhD in the field of computational statistics applied to population genetics. There you are usually trying to infer past history, and therefore have no hope of experimental verification. Although some important advances in computational statistics originate in that field, there were also lots of papers just throwing increasingly fancy Bayesian MCMC software at datasets without the honesty of experimental physics about what is knowable and what sadly might not be.
Also, many scientists, away from physics, simply do not have the mathematical maturity to understand most of statistics.
You are half correct. Something is intentionally hidden, but it's not the lack of real content.
The real stuff is called 'Mathematical Statistics' and what you learn in school is 'Statistic's for Scientists and Engineers' or 'Applied Statistics'.
Statistics loos to you random because it's useful everywhere and it must be taught to so many people in other fields. You get a basic set of per-selected tools that are useful and the real stuff is hidden.
If you really want to spend the time and effort to see it as fundamental from ground up, you have to learn mathematic background like set theory, Borel sets, Lebesque measures and integrals, Fourier integrals, etc. Then you start working up from fundamental axioms of random variables. The payday for practically oriented mind comes later and learning it may not be as motivating.
There is two ways to learn statistics. Statistical intuition vs. learning the core concepts. When probability is formulated using measure theory things really click together with the rest of the mathematics.
> That is math, not statistics.
Sometimes in the university level statistics department teaches applied statistics and you need to go mathematical department to learn mathematical statistics. What you think is 'real' statistics for you is matter of opinion, but non-applied statistics is pure abstract mathematics.
Statistic is a field to represent the noise and uncertainty by quantifying it.
It is unclear to you or lacking in brutal honesty but the honest truth is the world is not perfect. If Kobe can shoot 100% of the free throw then you can perfectly model it without statistic but sometime he misses and because of that we need statistic to take into account the fact that while we expect him to make freethrows sometime he miss his shot. Without statistic you cannot take into account the sometime Kobe miss shot.
It is also the field that define in definite what significant means. Statistically significant is the difference between this med really works versus it's a placebo/not sig enough.
Another concrete example of not perfect is data. Either collection of data (missing data, measurement, etc..) is not perfect and statistic helps. You can do ordinary least square without caring about statistic but finding outliers in the data, fixing the missing data, etc.. require statistic iirc.
Nope, you are making at least two errors here. You can search for "statistical significance vs practical significance" and "research hypothesis vs statistical hypothesis" to learn more.
That is a confusing topic though, it would be best done away with altogether like this course appears to do.
"It is a course intended for those that value the role of Bayesian inference and machine learning on [sic] their research."
Basically the value of statistics lies in the fact that there's some human-tractable set of noise processes and correlations which appear universally in a wide spectrum of real phenomena.
Knowing the real mechanism which generates the noise is always better, but not always tractable.
It violates the HN guidelines to include this sort of baity noise in comments here. We'd appreciate it if you'd read and follow the rules. They're at https://news.ycombinator.com/newsguidelines.html.