HNHacker News
TopNewBestAskShowJobs

christopheraden

292 karma · joined February 12, 2013

Data Scientist in Silicon Valley.

Programmers have unit tests, statisticians have model checking. Be careful! A disregard for a model's underlying assumptions can invalidate the conclusion.

Email: christopherBaden at gmail.

Known Languages: C (used infrequently), Python (used somewhat regularly), R (used daily), SAS (thankfully, used infrequently), SQL (used daily), Profanity (used daily).

Statistical Interests: Computational statistics, Bayesian inference, "embarrassingly parallel" statistical algorithms.

Applied Interests: Fraud, Data Science, Experimental Design, Interpretable Machine Learning

Favorite Reading Topics: The Hadleyverse (and most other R topics), parallel/concurrent programming, machine learning, applied statistical analyses

Religion: Bayesian

submissionscomments
christopheraden··on Zappos 2012 data breach settlement
The Ashley Madison Breach comes to mind. If the core demographic cares about not wanting their data on the platform to get out, they will vote with their feet. That said, I think this example is not the norm, and most people probably won't care for most applications. https://en.wikipedia.org/wiki/Ashley_Madison_data_breach
christopheraden··on Targeted Ads in Live TV
> Have fun explaining to your spouse why your household's TV is showing more dating site ads than that of their friends.

The targeting is sometimes only _so_ good. While it works in aggregate well, sometimes there are laughable targetings, and it also depends on what audiences the advertiser wants to hit.

I get some pretty strange targeted ads through Facebook that don't seem at all relevant, and the "Why Am I Seeing This Ad" dropdown has very nebulous explanations ("Targeting Men between 25-35 in San Francisco"). I would imagine this problem to be similar with targeted ads on Live TV unless they keep the ads pretty generic or have vastly better targeting than FB or Google.

christopheraden··on Study Reveals U.S. Consumers and Economy Lose Billions to Occupational Licensing
> I feel like this argument should be a class of fallacy. Lots of things have "existed" for a long time, but that doesn't mean previous iterations were effective.

This is pretty close to Survivorship Bias (https://en.wikipedia.org/wiki/Survivorship_bias). Outcomes of midwifery of yore is no different than present midwifery if you ignore the fatalities and only focus on successful deliveries.

christopheraden··on 1Password Travel Mode: Protect your data when crossing borders
While it is a policy issue at its core, changing law and policy moves at a glacial pace, and it's not even a certainty that it'll get changed at all (the "nothing to hide" defense gets brought up a lot on these matters, and it's a pretty persuasive argument to those that can't recognize the fallacy). Technical solutions have the benefit of being much quicker to enact, albeit in a flawed way that is, if highly successful, a band-aid on a bullet wound.
christopheraden··on Same Stats, Different Graphs: Datasets with Varied Appearance and Identical Stats
Sure, using hypothesis tests could pick out some of the structured examples in the Datasaurus, but in practice, things are often more subtle. Goodness of Fit tests to check for normality, in particular, are a little bit thorny, lacking power in small sample sizes, and rejecting normality for slight departures in higher sample sizes. My experience has been with assumption checking that by the time a hypothesis test has sufficient evidence to reject an assumption, you'd usually be able to see it visually.

Until you get into high dimensions, it probably doesn't hurt too much to visualize the data. Additionally, it can be helpful to understand what signal has been left in the residuals (ex: you fit a linear model, but failed to include a quadratic term), which is something hypothesis tests aren't as good at telling you.

christopheraden··on Analyzing 10,000 Show HN Submissions
A couple questions.

> OLS works fine in classification problems. And it has advantages.

Do you have more explanation of these advantages? I read through the link you sent, and a bit more about linear probability models. Such things were never discussed in my statistics curriculum (BS, MS, PhD), except for motivating why logistic regression was necessary. I'm not sure I understand the economist's arguments in favor of LPM. Both the interpretation and the distribution of the test statistics will be totally different with OLS versus Logistic Regression, and the overall probability of a defunct project is pretty small ( \hat{P(y=0)} = .07 )--enough where there would be pretty big differences. To be clear, my reservation is with the p-values in the OLS model, not the predictions it generates. While the models agree on the direction of the covariates, the magnitudes are quite different, even when you convert logit/probit to be on the same scale as LPM.

> Multicollinearity refers to perfect multicollinearity.

Perfect multicollinearity will definitely mess up the estimation, but even if Score and Comments are not perfectly collinear, it's difficult to talk about each one's effect on the probability individually, as is the interpretation of coefficients in a (logistic) regression. What does the VIFs look like for Score and Comments, in particular?

> I mention R² as one measure of predictive power.

But the outcome is binary, so you'll have a similar issue as Minimaxir's first point about OLS. If you wanted to talk about prediction accuracy, what about a confusion matrix, misclassification rate, or specificity/sensitivity/F1? Granted, you'll not want to predict on the same tagged examples that you trained the model on, but maybe you could split it 80-20? Or tag another 20-50? There are also R²-like measures you can use when the dependent variable is binary (a whole class of pseudo-R² measures).

I would be curious to see the relationship between these predictors and the response. It's usually been my experience that linearity is a strong assumption to make, and that I'd expect for something like comments or score that once it reached a certain threshold, there was no extra value added by getting more comments/score. Are the log-score and log-comments linear over their entire support?

christopheraden··on What's new in TeX
What about something ala Neovim? I've only ever looked at Tex from the perspective of a user (I don't program my own macros too often), so I don't know how hard it'd be, but why not a language overhaul?
christopheraden··on Comparison – R vs. Python: head to head data analysis
I use SAS professionally at my job, and R in all my academic/hobby work. R has a couple packages that give similar functionality as PROC SQL (about 95% of my SAS workflow, since it's far nicer than data steps for a lot of things). There's an ODBC package (RODBC), as well as SQLDF, which allows you to use SQL queries to manipulate data frames in R.

While there is (almost?) always a way to do a SQL query using idiomatic R, I have to admit that sometimes my brain thinks up a solution in SQL faster (a product of upbringing).

christopheraden··on New York Mayor De Blasio to Require Computer Science in Schools
>What kind of negative effect might result from a bunch of unqualified high school teachers teaching CS poorly? Is some exposure better than none regardless of teaching quality?

Is this problem similar enough to math that we can draw on data about mathematics education in schools? You can make a lot more money in industry with a math degree than you can teaching math to grade schoolers.

Certainly, we hear tons of stories about unqualified math and science teachers, and all the harm they do to a desire to learn math, but I'm sure some students get exposure to math that wouldn't normally have much desire to immerse themselves in it.

christopheraden··on Moebio Framework – A JavaScript toolkit for data analysis and visualizations
I agree it's a barrier to have such crazy prices, but there are free resources available, especially on a topic as popular as Grammar of Graphics. Hadley Wickham (in my mind, synonymous with the concept, since he implemented Wilkinson's ideas in R's ggplot package), for instance, has numerous materials on it, including a short primer (http://vita.had.co.nz/papers/layered-grammar.pdf). It might not be as exhaustive as the Wilkinson text, but surely there's enough material out there to implement GG in JS, especially considering there's successful implementations in other languages?
christopheraden··on OS X El Capitan GM Candidate available
Excellent! I guess now is the time for me to finally make the move over to El Capitan, since your usage seems pretty similar to mine. Thanks!
christopheraden··on OS X El Capitan GM Candidate available
Can anyone comment if the GM fixes the battery problems Beta 1 had? I was on the first beta awhile back, and my 2012 rMBP got about an hour battery life (I usually get closer to ~3-4) and was always hot to the touch.

The same thing happened to me with Mavericks, so I chalked it up to Apple not optimizing the early betas for battery life.

christopheraden··on “I have already used the name for my programming language” (2009)
On twitter, the common way to refer to R is rstats, which seems to work okay. See: https://twitter.com/hashtag/rstats
christopheraden··on Tufte CSS
Tufte-Latex, mentioned in the article, is a really nice template that produces some gorgeous Latex handouts with very little effort (I've used it a couple times when I wanted something to not look like the standard Latex article class). I thought it worth linking here for those that didn't read the link: https://tufte-latex.github.io/tufte-latex/
christopheraden··on Continuum Analytics Raises $24M Series A for Anaconda Python and PyData
Fixxer, I totally agree that it's great that one be capable of doing these things, but sometimes it's not as important as other things that could be taught. Like acbart, sometimes I want to teach why/when to use a statistical algorithm and not teach them how to grab all the dependencies, troubleshoot whether they have gfortran installed, etc. This problem is horribly compounded teaching undergrads when you have Windows, Linux, and Mac users in your class, where the procedures for getting a working scientific stack vary, and the errors are often not the same across platforms.

When I install the scientific python stack on a new machine, I almost always just use Anaconda. I already know how to install the stack (I'll always value the weekend I spent in undergrad fighting with a customized BLAS in R!), but sometimes I have more pressing/fun things to do.

christopheraden··on Continuum Analytics Raises $24M Series A for Anaconda Python and PyData
Very well deserved!

With Anaconda, I just tell my students to download a quick installer, slap on iPython or PyCharm, and it's ready to go. It's one less thing to worry about! The installation is dead-simple, and is almost exactly the same whether on Mac, Windows, or Linux. When I do data science, I don't want to have to be doing IT, too!

christopheraden··on The Risky Eclipse of Statisticians
Definitely agree.

Part of the problem is that data science doesn't have nearly the same formalism in its definition that statistics does. What's the difference between BI's, Data Miners, Data Analysts, Data Scientists, etc? The tools used to arrive at conclusions (R vs. Python vs. SAS vs. Tableau/Excel/SPSS) doesn't seem like a good way of differentiating the roles.

A more useful discriminator would be the application of statistics (BI vs. Biostatistician, for instance), the depth and complexity of the statistical algorithms used, and whether the main use is stat inference or prediction (machine learning doesn't seem to focus on inference a whole lot, for example).

christopheraden··on Potato paradox
I guess mine is sort of an anti-food-related math concept, then! https://en.wikipedia.org/wiki/No_free_lunch_theorem
christopheraden··on Designing Machine Learning Models: A Tale of Precision and Recall
I come from the world of biostatistics, where diagnostic tests are usually measured in terms of Sensitivity (probability of Predicting Evil, given actually Evil, same as Precision) and Specificity (probability of predicting Not Evil, given actually Not Evil). What is the reason for choosing to balance Precision/Recall, versus Sensitivity/Specificity? They are definitely similar, so why prefer one tradeoff versus the other?
christopheraden··on Designing Machine Learning Models: A Tale of Precision and Recall
They hint at withholding data ("The same logic applies when it comes time to splitting our data into training and validation sets"), though they don't outright mention cross-validation. Seems like a bit of an oversight for an article about building ML models, I agree, but the article seems like it's mostly surface-level.
christopheraden··on Data Science from Scratch: First Principles with Python
Most likely http://en.wikipedia.org/wiki/Design_of_experiments
christopheraden··on Sick time is logarithmic
Thank you for posting the data--it makes it easier for us to follow along at home.

I've written something up where I used tenure instead of the ranks. http://christopheraden.github.io/SickTime.html.

Here’s my concern with the original model, having slept on it: What does it actually buy you? In order to use the model, you must know the rank of sick time of the employee, relative to all other employees. In order to calculate this, you have to have the raw sick time numbers. At which point, what’s the point in making a regression with it—just work with the raw sick times to begin with!

Wouldn't sick time taken per year be a more interesting measure? Over time, wouldn't you expect that an employee of 30 years would've taken more sick time than an employee of six months? Aren't you really concerned about whether in a particular year older employees are not taking as much vacation as newer employees?

christopheraden··on Sick time is logarithmic
Depends on whether you represent sick-time as continuous or discrete. Zipf's Law takes support on the integers, whereas the Exponential takes support on all positive reals.
christopheraden··on Sick time is logarithmic
This is an interesting application of statistics! I'm surprised how well the sick-time ranking correlates with the sick times. Intuitively, I'd imagine we'd expect correlation between the ranks and the raw values (your X axis is just an ordinal version of the Y axis), but this is still quite high. I imagine that's a function of the underlying distribution. Your data could be conveyed effectively with a histogram, and I think that would paint an interesting picture.

Is there a picture associating the sick-time and the tenure? You mention it increases the R^2, but don't have an associated graph. Maybe you could post a scatterplot of Sick-Time and Tenure and show that logarithmic relationship?

Would you be able to provide anonymized versions of the data, or something like a table with (ID, Sick Time, Years of Tenure)? I'd love to play with it, and try a QQ plot. My email's in my profile.

christopheraden··on Sick time is logarithmic
Try graphing y = -1 * log(x) and imposing a limit on the upper bound of x and you'll get close to what he has. Perhaps that's the angle he's coming from. He provided the fitted equation further down in the featured article, and the log term does have a negative coefficient, plus an intercept term.

The graph he plots looks like the data fits the Exponential Distribution: http://en.wikipedia.org/wiki/Exponential_distribution

christopheraden··on Normally distributed and uncorrelated does not imply independent
But then the requirement that the rv's be jointly normal is violated. The "jointly normal + uncorrelated" combination is special. There aren't too many other named distributions that have the property that uncorrelated implies independence.
christopheraden··on Normally distributed and uncorrelated does not imply independent
Could you clarify this statement? Independence is defined in terms of distributions (the joint distribution can be split up into a product of marginals), so I'm not sure how "the way a set of data is distributed" and "can never imply independence" jive.
christopheraden··on The harm done by tests of significance (2003) [pdf]
>But if you look at the effect sizes, you see the five studies found nearly the same answers -- its just two of them didn't quite cross the threshold for significance.

This is why I wish meta-analysis was introduced much earlier than it is in statistics education. There are sensible ways of combining the information from the five studies, weighting them according to their sample size (provided the studies are similar in design and cohorts). =)

christopheraden··on A Faster Way to Try Many Drugs on Many Cancers
You are correct about the tissue samples. CCR mostly collects treatment, demographic, and disease information. I didn't get the feeling from their website that FlatIron was collecting tissue samples, either, though.
christopheraden··on A Faster Way to Try Many Drugs on Many Cancers
It seems like they are doing more than that, though. Cancer data has been collected (by law) and stored for decades. I'd be curious to see if FlatIron was using historical records from Cancer Registries (California's Cancer Registry has millions of cases, for example), or just "EMR and billing systems".
Page 1 of 4Next →