292 karma · joined February 12, 2013
Programmers have unit tests, statisticians have model checking. Be careful! A disregard for a model's underlying assumptions can invalidate the conclusion.
Email: christopherBaden at gmail.
Known Languages: C (used infrequently), Python (used somewhat regularly), R (used daily), SAS (thankfully, used infrequently), SQL (used daily), Profanity (used daily).
Statistical Interests: Computational statistics, Bayesian inference, "embarrassingly parallel" statistical algorithms.
Applied Interests: Fraud, Data Science, Experimental Design, Interpretable Machine Learning
Favorite Reading Topics: The Hadleyverse (and most other R topics), parallel/concurrent programming, machine learning, applied statistical analyses
Religion: Bayesian
The targeting is sometimes only _so_ good. While it works in aggregate well, sometimes there are laughable targetings, and it also depends on what audiences the advertiser wants to hit.
I get some pretty strange targeted ads through Facebook that don't seem at all relevant, and the "Why Am I Seeing This Ad" dropdown has very nebulous explanations ("Targeting Men between 25-35 in San Francisco"). I would imagine this problem to be similar with targeted ads on Live TV unless they keep the ads pretty generic or have vastly better targeting than FB or Google.
This is pretty close to Survivorship Bias (https://en.wikipedia.org/wiki/Survivorship_bias). Outcomes of midwifery of yore is no different than present midwifery if you ignore the fatalities and only focus on successful deliveries.
Until you get into high dimensions, it probably doesn't hurt too much to visualize the data. Additionally, it can be helpful to understand what signal has been left in the residuals (ex: you fit a linear model, but failed to include a quadratic term), which is something hypothesis tests aren't as good at telling you.
> OLS works fine in classification problems. And it has advantages.
Do you have more explanation of these advantages? I read through the link you sent, and a bit more about linear probability models. Such things were never discussed in my statistics curriculum (BS, MS, PhD), except for motivating why logistic regression was necessary. I'm not sure I understand the economist's arguments in favor of LPM. Both the interpretation and the distribution of the test statistics will be totally different with OLS versus Logistic Regression, and the overall probability of a defunct project is pretty small ( \hat{P(y=0)} = .07 )--enough where there would be pretty big differences. To be clear, my reservation is with the p-values in the OLS model, not the predictions it generates. While the models agree on the direction of the covariates, the magnitudes are quite different, even when you convert logit/probit to be on the same scale as LPM.
> Multicollinearity refers to perfect multicollinearity.
Perfect multicollinearity will definitely mess up the estimation, but even if Score and Comments are not perfectly collinear, it's difficult to talk about each one's effect on the probability individually, as is the interpretation of coefficients in a (logistic) regression. What does the VIFs look like for Score and Comments, in particular?
> I mention R² as one measure of predictive power.
But the outcome is binary, so you'll have a similar issue as Minimaxir's first point about OLS. If you wanted to talk about prediction accuracy, what about a confusion matrix, misclassification rate, or specificity/sensitivity/F1? Granted, you'll not want to predict on the same tagged examples that you trained the model on, but maybe you could split it 80-20? Or tag another 20-50? There are also R²-like measures you can use when the dependent variable is binary (a whole class of pseudo-R² measures).
I would be curious to see the relationship between these predictors and the response. It's usually been my experience that linearity is a strong assumption to make, and that I'd expect for something like comments or score that once it reached a certain threshold, there was no extra value added by getting more comments/score. Are the log-score and log-comments linear over their entire support?
While there is (almost?) always a way to do a SQL query using idiomatic R, I have to admit that sometimes my brain thinks up a solution in SQL faster (a product of upbringing).
Is this problem similar enough to math that we can draw on data about mathematics education in schools? You can make a lot more money in industry with a math degree than you can teaching math to grade schoolers.
Certainly, we hear tons of stories about unqualified math and science teachers, and all the harm they do to a desire to learn math, but I'm sure some students get exposure to math that wouldn't normally have much desire to immerse themselves in it.
The same thing happened to me with Mavericks, so I chalked it up to Apple not optimizing the early betas for battery life.
When I install the scientific python stack on a new machine, I almost always just use Anaconda. I already know how to install the stack (I'll always value the weekend I spent in undergrad fighting with a customized BLAS in R!), but sometimes I have more pressing/fun things to do.
With Anaconda, I just tell my students to download a quick installer, slap on iPython or PyCharm, and it's ready to go. It's one less thing to worry about! The installation is dead-simple, and is almost exactly the same whether on Mac, Windows, or Linux. When I do data science, I don't want to have to be doing IT, too!
Part of the problem is that data science doesn't have nearly the same formalism in its definition that statistics does. What's the difference between BI's, Data Miners, Data Analysts, Data Scientists, etc? The tools used to arrive at conclusions (R vs. Python vs. SAS vs. Tableau/Excel/SPSS) doesn't seem like a good way of differentiating the roles.
A more useful discriminator would be the application of statistics (BI vs. Biostatistician, for instance), the depth and complexity of the statistical algorithms used, and whether the main use is stat inference or prediction (machine learning doesn't seem to focus on inference a whole lot, for example).
I've written something up where I used tenure instead of the ranks. http://christopheraden.github.io/SickTime.html.
Here’s my concern with the original model, having slept on it: What does it actually buy you? In order to use the model, you must know the rank of sick time of the employee, relative to all other employees. In order to calculate this, you have to have the raw sick time numbers. At which point, what’s the point in making a regression with it—just work with the raw sick times to begin with!
Wouldn't sick time taken per year be a more interesting measure? Over time, wouldn't you expect that an employee of 30 years would've taken more sick time than an employee of six months? Aren't you really concerned about whether in a particular year older employees are not taking as much vacation as newer employees?
Is there a picture associating the sick-time and the tenure? You mention it increases the R^2, but don't have an associated graph. Maybe you could post a scatterplot of Sick-Time and Tenure and show that logarithmic relationship?
Would you be able to provide anonymized versions of the data, or something like a table with (ID, Sick Time, Years of Tenure)? I'd love to play with it, and try a QQ plot. My email's in my profile.
The graph he plots looks like the data fits the Exponential Distribution: http://en.wikipedia.org/wiki/Exponential_distribution
This is why I wish meta-analysis was introduced much earlier than it is in statistics education. There are sensible ways of combining the information from the five studies, weighting them according to their sample size (provided the studies are similar in design and cohorts). =)