Well-known paradox of R-squared is still buggin me
statmodeling.stat.columbia.edu
statmodeling.stat.columbia.edu
If he doesn't like that he should just use R by itself, which would turn that 0.01 into 0.1, and it'd turn that 0.16 into 0.4. R is the Pearson correlation coefficient of a univariable linear regression.
> if X,Y,Z are "generic" random variables such that ρ(X,Y)=a and ρ(Y,Z)=b, we should on average expect that ρ(X,Z)=ab.
https://www.lesswrong.com/posts/vfb5Seaaqzk5kzChb/when-is-co...
c) The R^2 used with a linear model requires a constant term, in this case the constant term or bias explains a lot about preferences (almost 50/50) so there is less information available for the slope term.
Hope this helps.
This explains the paradox, basically. When you take the null model "preference = 50%" (i.e. intercept-only model), there simply isn't much residual variance left for the linear model to explain.
That's why you get an R^2 = 1 if you use the "R^2 = rho(state, preference)^2" formula (you are ignoring the role of the intercept in explaining most of the variance, and exploiting the translation-invariance of the Pearson correlation) vs. you getting an R^2 = 0.01 when you use the (more correct) "R^2 = explained variance / total variance" formula.
TL;DR: It makes sense to get a very low R^2 when it is the intercept and not the predictor that is explaining most of the variance.
I'd say that there is still quite a lot of residual variance to explain. You need a baseline - the worst choice would be to predict 0 (or 1) and the mean squared error would be 0.5. Using 0.5 as baseline halves the mean squared error to 0.25.
> data <- data.frame(state = c(0, 1), pref = c(0.45, 0.55))
> sum(residuals(lm(pref ~ 0, data = data))^2) # null model
[1] 0.505
> sum(residuals(lm(pref ~ 1, data = data))^2) # intercept-only model
[1] 0.005
> sum(residuals(lm(pref ~ state + 0, data = data))^2) # predictor-only model
[1] 0.2025
So, it seems clear that you only get a "perfect" prediction with the full (intercept + predictor) model mostly because of the intercept (which explains (0.505-0.005)/0.505 = 0.99 = 99% of the variance).
Thus, it makes sense that the predictor is only explaining the rest (i.e. 1%) of the variance... hence, the R^2 = 0.01
Maybe I'm completely missing your point but the calculations in the blog post are, adapting your code (I think you meant mean where you wrote sum):
> data <- data.frame(state = rep(c(0, 1), each=20), pref = c(rep(0, 11), rep(1, 9), rep(0, 9), rep(1, 11)))
> mean(residuals(lm(pref ~ 0, data = data))^2) # null model [NOT IN THE BLOG POST]
[1] 0.5
> mean(residuals(lm(pref ~ 1, data = data))^2) # BASELINE intercept-only model
[1] 0.25
> mean(residuals(lm(pref ~ state + 0, data = data))^2) # predictor-only model [NOT IN THE BLOG POST]
[1] 0.34875
> mean(residuals(lm(pref ~ state, data = data))^2) # MODEL
[1] 0.2475
> summary(lm(pref ~ state, data = data))$r.squared # MODEL
0.01
The blog post is about what you call "intercept-only" model (MSE 0.25) and the full model (MSE 0.2475), the R² is (0.25-0.2475)/0.25=0.01. His calculation is slightly different: instead of 0.25-0.2475 he calculates directly 0.05^2 which is the variance of the predictions (in this case the total variance 0.25 can be decomposed as the variance of the errors 0.2475 plus the variance of the predictions 0.0025).Either way, the point stands... the improvement in using a full linear model (that predicts 0.45 or 0.55, depending on state) is marginal compared to the baseline model that always predicts 0.50, as you demonstrate with your code.
To me, this doesn't seem paradoxical... the predictor is indeed providing little information over the "let's flip a coin to predict someone's voting preference" null/baseline predictor, since people's preferences (in aggregate) are almost equivalent to "flipping a coin".
note: I meant "sum", but it's the same, since the ratio between sums of squares is equivalent to the ratio between mean squares
Yes, I think we don't disagree. I was just puzzled by the "little variance left to explain" remark.
> note: I meant "sum", but it's the same, since the ratio between sums of squares is equivalent to the ratio between mean squares
You're right, sum of squares made sense if it was just for the ratio.
It's not that there is "little variance left to explain", but actually that (no matter what) there will always be too much variance left to be explained, when the response is Bernoulli-distributed and the parameter is not too far from 0.5 (i.e., the data generating process is like flipping a slightly loaded coin).
If you use the expected value to predict the Bernoulli variable, you will always be somewhat wrong (0.45 and 0.55 are both far from 0 and from 1, which are the only possible responses).
If you use a binary response to predict, you will quite often be very wrong, even if you are right on average, and even if your prediction is to generate Bernoulli-distributed samples from the exact same distribution (i.e., you know exactly how the coin is loaded/biased and you can exactly replicate its data generation process).
So... yeah... no "paradox" ;)
But, anyway, seems like I interpreted things incorrectly.
The reason we do that is because we are assuming the errors are normally distributed and finding the slope that gets the best possible R2 is equivalent to getting figuring out how to fit the line with the maximum likelihood estimator of the error (aka mean of the distribution).
So ultimately it’s about curves. If you wanted to get a sense for why this is strongly desirable you should try to fit a linear regression using absolute error instead of squared error.
A perception engineer took one look at the controller and saw linear error terms for distance left + distance right (distance to walls), changed it to distance left^2 + distance right^2, and the whole thing magically worked beautifully. Exercise for the reader: What position in the hallway minimizes sum of distances squared, vs what position(s) in the hallway minimize sum of distances without square.
This is essentially the same problem you pose.
Edit: no, derp, just walked into the same problem lol. Maximize the min should work though
Isn't that around the top of the statistics no-nos list? Probability theory in general? Predicting individual result based on the whole sample base? It was long ago and I am not in this field at all but my recollection tells me that it was mentioned in the beginning of the first class of probability theory 101.
But it requires uniform sampling.
My opinion is LLMs are just applied statistics and you see people losing their minds thinking the models have "come to life". Most people just really have no intuition for stats.
Are they? Sounds like they're both swing states, pretty close to 50-50, so which state you're from doesn't have a big effect on what your politics are likely to be. Which is exactly what the R^2 tells us. Where's the paradox?
Also, traditional R squared with binary variables, or maybe categorical variables, never made much sense to me. The "meaning" of the variance I don't think is quite the same. You generally have to nonlinearly transform (e.g., logit) any linear model quantities into something else to put it on an observed variable scale.
In other countries 45/55 could easily be a swing state.
For example the 2014 Scottish independence referendum was decided 45/55 and that is usually held to be a fairly close result, at least close enough that it hasn’t ended the question.
So the real story here is voter polarization. The reason that Texas isn’t a swing state is because there’s a big hard core of immoveable voters at either end. So the actual population of swing voters is small.
I genuinely don't think that outside of a small minority who drank the Trump kool-aid the political views of the country at large have really changed. The biggest shift is that my friends who were and still are run-of-the-mill "Republicans" don't feel particularly represented right now, and I get it.
With the murmurs of Biden maybe stepping down I think if the Democrats nominated a Moderate Republican they would win in a landslide. It's nuts how much I didn't appreciate having two genuinely good candidates in 2008/2012.
No, it isn't. Your "ten point gap" is something like a 5% delta. That's below the confidence level of some polls.
Exactly. I propose that the paradox is in first-past-the-post voting, a 5% swing leads to a 100% change in representation. How can that be?
> The most commonly-used example of a ranked-choice system is the familiar plurality voting rule, which gives one "point" (vote) to the candidate ranked first, and zero points to all others (making additional marks unnecessary).
Additionally, rated choice voting systems are not subject to Arrow's theorem.
If a state had an election with millions of voters and got a 55-45 result, it would be a decisive landslide victory; If a elementary school classroom had an election 20 voters had got a 55-45 split, it's be the narrowest possible margin of victory.
Most would likely say that the effect in the former 'feels' much larger even though proportions are identical, which suggests that under the hood we're factoring in sample size to our intuition about effect sizes (probably something chi-square-ish).
The result is that the framing of the problem can change our sense of how big the effects are. When we hear that these are state-level elections, we think it's a huge effect and feel that we should be able to do reverse inference. If it was reframed as an election on a much smaller sample, the paradox disappears and you'd say "of course you wouldn't be able to reverse that inference"
It's very common to confuse the two ideas.
In particular, in an election with many millions of votes and a 55-45 margin, it's common to describe the winner as receiving a mandate to rule, because it was so easy to determine who the winner was, despite the fact that they appear to be extremely unpopular. That's not a mandate in any ordinary sense.
The more accurate terms to describe this is whether an effect is significant (i.e., "do we have enough information to be able to claim it is different from zero?") vs. whether an effect is relevant (i.e., "is [our estimate of] the size of the effect meaningfully different from zero?").
Statistics can only address the first issue (effect significance); the second issue (effect relevance) requires domain knowledge beyond statistics (in some contexts, a difference of 0.01 units can be irrelevant, while it may be relevant in other contexts).
Guessing 0.5 will have you wrong wrong by 0.5 100% of the time. SST is 25 for a 100 sample example.
Guessing 0.55 for the 0.55 state will have you wrong by 0.45 55% of the time and 0.55 45% of the time for the other. SSE is 24.75
1- 24.75 / 25 = 0.01
Looking at it this way it’s not too hard to see why the R2 is bad. It barely explains any more difference in the individual behavior than the basic guess.
R2 is not a great metric for percentages or classification problems like this.
Right. R² is 1% because the prediction is bad - only marginally better than the basic guess.
> R2 is not a great metric for percentages or classification problems like this.
Using a different metric won't improve the prediction.
Is the Brier score a great metric for problems like this?
The Brier score for the model is 0.2475.
The Brier score for the "basic guess" is 0.25.
The improvement in the Brier score for the model relative to the basic guess is 1%.
R2 is not the correct measure to use.
This article is a perfect example of the principle that simply doing math and getting results is not necessarily meaningful.
Does another measure give substantially different results?
As the user fskfsk.... says in another comment, here the constant term explains a lot of the variance so that the slope terms contains less information, that is not available using your definition or idea of R^2
Different from what?
According to wikipedia:
The most general definition of the coefficient of determination is R^2 = 1 - SS_res / SS_tot ( = 1 - 0.2475 / 0.25 = 0.01 in this case)
Edit to clarify the definition above:
SS_res is the sum of squares of residuals (also called the residual sum of squares) ∑( y_i - predicted_i )^2
SS_tot is the total sum of squares (proportional to the variance of the data) ∑( y_i - ∑y_i/N )^2
The population variance is the sum of the Between Group Variance and the Within Group Variance weighted by the number of elements in each group.
> data <- data.frame(state = rep(c(0, 1), each=20), pref = c(rep(0, 11), rep(1, 9), rep(0, 9), rep(1, 11)))
> summary(lm(pref ~ state, data = data))$r.squared
0.01I may be reading too much from your comment, but it seems that you relate R^2 to the reduction in the prediction error in each state, so it seems you are thinking about the formula of computing the R^2 as the (average variance in each state)/(total variance), that I think is not correct in general since at least it should require the total variance to be the sum of the variances in each state. If you based your ideas in that formula then your intuition is not correct, that is my point. When I apply R^2 I am thinking in a multivariable linear model with continuous variables, and this is not the case. I should measure this problem by how the entropy change when we apply the information about the state, something like the cross entropy using the total distribution and the distribution by states.
The mean squared error of the baseline model which doesn't include the state as a regressor is 0.25 (it predicts always 0.5 - it's off by 0.5 in every case).
The mean squared error of the model which includes the state as a regressor is 0.2475 (it predicts 0.45 or 0.55 depending on the state - in both cases it's off by 0.45 with 55% probability and it's off by 0.55 with 45% probability).
The mean squared error is directly related to variance when the predictor is unbiased. The ratio of the sum of squares is the same as the ratio of the mean square errors.
Edit: http://brenocon.com/rsquared_is_mse_rescaled.pdf
"R2 can be thought of as a rescaling of MSE, comparing it to the variance of the outcome response."
https://dabruro.medium.com/you-mention-the-average-squared-e...
"Also it is worth mentioning that R-squared (coeff. of determination) is a rescaled version of MSE such that 100% is perfection and 0% implies the same MSE that you would get by simply always predicting the overall mean of the dataset."
There must be a formula to compute R^2 from variances both among states and inside states but anyway, when the variances inside any state are bigger that the total variance that should imply that the feature that divides the population in groups is of little value for prediction so it should have a small R^2 value.
I was replying to someone who claimed that "R2 is not the correct measure to use. This article is a perfect example of the principle that simply doing math and getting results is not necessarily meaningful." I've not seen any comment from anyone getting "different results" with a different measure.
Edit: You used var(...) which includes a factor N/N-1 and doesn't give exactly the total sum of squares.
The example dataframe contains 40 observations (20 per state) and you get higher variance estimate for the subsamples than for the aggregate sample but if you put toghether a few copies of the data (for example doing "data <- rbind(data, data, data, data, data)") even the adjusted (unbiased) estimator of the variance is lower for the states.
You can calculate the "exact" values yourself doing (x-mean(x))^2 or undoing the adjustment:
> var(data$pref)*39/40
[1] 0.25
> var(data[data$state==0, "pref"])*19/20
[1] 0.2475
> var(data[data$state==1, "pref"])*19/20
[1] 0.2475
> when the variances inside any state are bigger that the total varianceThey are not. But you're right in that a small difference shows that dividing the population in groups is of little value for prediction and that's why the R^2 value is small.
I just added another comment that relates analysis of variance to this post to show that there is no real paradox here.
Finally, the formula for the total variance above is related to my intuition that having some information (having the data for each state) should make the means of the variances in each group smaller that the total variance, because variance is related to lack of information. But analysis of variance suggests (see other comment of mine) that the state factor is not representative because the high variance in each group (each state) and the low difference between the groups means and the total mean.
You would also want to be predicting 0.45 and 0.55 not 1 and 0 because we solve for squared error.
The blog post was rebutted by Seth in the comments 2 weeks ago, same as Colin Percival's HN rebuttal, and Andrew didn't reply. It seems like a weird goof. Andrew was "buggin".
Because classification is a regression problem.
Think about it for a second. You want to put together a tool to tell which class an input belongs to. You have training data you can use to build your tool around. Your training data is already divided into sets that belong to a specific classm Your goal is to put together a model that can tell you what's the closest class your input belongs to by comparing with how close your input is to elements of the training data belonging to a specific class.
What's your strategy?
Well, one of the textbook strategie starts by specifying how you measure the distance between elements of your training set, and from that point you work on putting together a function that not only minimizes the distance between elements of your training set but also, when used to evaluate elements of a training set, works well in telling the type of elements of the training set that are closest to them. Then you assume the class of your input element is the same as the class of the elements of the training set that are closest to them.
In the example above, the minimization step is... Yes, regression. You use regression to fit your model to your training data so that it is able how close your input element is to elements of a certain class, and then outputs how close it is to each of the classes.
You're simply wrong. Regression is a tried and true classification technique. Posting personal and baseless assertions don't refute that. I mean, pick up pretty much any textbook on supervised and unsupervised learning. You always end up with an approach which boils down to having training data, put together a trial function, apply a minimizer to fit trial functions to training data, and evaluate the resulting model by running trial data through it. Fitting trial functions to training data has a name: regression. Minimum squares has a very precise interpretation both in linear models and in probability. There is no way around it.
What I said was that regression metrics are not good for evaluating usefulness in classification problems. An R2 of 0.01. The fact that there are mitigating circumstances to justify why this might not be the case is not evidence that R2 is still a good thing to use. It’s actually evidence of the opposite.
With classification problems we are concerned more often with the ordinality of estimated probabilities vs outcomes and/or the calibration of a model.
> Posting personal and baseless assertions don't refute that.
Your mom didn’t refute me either.
You're voicing very opinionated takes while showing considerable ignorance on the topic.
Classification problems are solved with trial functions adjusted to the training set through regression. These trial functions,once fitted, represent membership functions. They are essentially interpolation functions that, say, converge to 1 when close to elements of the training set of a specific class, and 0 for members of all other classes. In grey areas where elements of the training set are sparsely distributed, these membership functions can output values in the middle, because they are interpolating.
That's literally data mining 101.
> What I said was that regression metrics are not good for evaluating usefulness in classification problems.
That's simply not true. It's like saying models that do not fit the data are good approximations of the data. Another way to put it is praising the accuracy of a broken clock because it's spot on two times a day.
> The fact that there are mitigating circumstances to justify why this might not be the case is not evidence that R2 is still a good thing to use. It’s actually evidence of the opposite.
I don't think you have an adequate grasp on the subject, neither classification problems nor linear models. Therefore, your personal baseless assertions don't mean anything nor bring any value to the discussion.
You’ve twice now tried to explain how fitting a model works in response to me stating that regression “metrics” are not well suited to describing classification “efficacy”. If anyone is making a basic mistake here it seems to be you failing to understand the difference between a “regression model” like a linear regression, and a “regression metric” like R2.
I’ve stated my semantic meaning of regression vs classification. You can Google this to see it is not a fringe view. It’s been standard for over a decade. Eg https://math.stackexchange.com/questions/141381/what-is-the-...
> Another way to put it is praising the accuracy of a broken clock because it's spot on two times a day.
Accuracy is a classification metric kiddo
Here’s an introductory Wikipedia article if you would like to learn more about what metrics are appropriate for binary classification evaluations:
https://en.m.wikipedia.org/wiki/Evaluation_of_binary_classif...
The normal R^2 formula can't be applied to a logit/probit model; instead you use an alternative such as McFadden's or Cox & Snell pseudo R-squared. I'd be interested to see what value they take for this example.
Linear models are sometimes used even in models with many independent variables since it can be shown that the coefficients in a linear model are unbiased estimators for the average partial effects of any non-linear binary response model.
1. if you insist on using the r-squared (i.e., a linear regression measure), then properly center and normalize your data, and model what you actually predict: the difference between the baseline (0.5) and the probability to vote for party 0 or party 1. If you model the outcomes as 0/1 without this, then you are using a model made for gaussian variables on what should be a logistic regression 2. if you can live with something that more accurately captures the idea of "explanatory power", you can use a GLM (logistic link function), do a logistic regression, and then use the log odds or another measure.
In both cases, the variance explained by the state that you are in is 1, because of course it is, that's how the thought experiment is constructed - p(vote for party 1)=0.5+ \delta(state).
"Paradoxes" like this are often interesting in the sense that they point to the math being the wrong math or you using it wrong, but instead people tend to assume that they are obviously understanding things correctly so it must be some weird property of the world (which then sometimes is used to construct some faulty conclusions as in some of the cited papers)
The above is in the context of analysis of variance. In our example the means in each state are 0.55 and 0.45 and the total mean is 0.50 so first summand is small but the variances in the red and blue states are both 0.247, large summand, so the variations we observe are just sampling variations. Hence the state factor is not important and that explains the low R^2 value. Note that in each state the predicted value for the model is the group mean of that group. So analysis of variance explains that the OP result is not a paradox or something strange.
In a generic setup, imagine you have a binary classifier that outputs probabilities in the .45-.55 range - likely it won't be a really strong classifier. You would ideally like polarized predictions, not values around .5.
Come to think of it, could this be an issue of non-ergodicity too ( hope I'm using the term right)? i.e. state level prior is not that informative wrt individual vote?
People who try to correct for “unbalanced classes” and contort their model to give polarizing predictions are frankly being pretty dumb.
The correct answer is to take your well calibrated probabilities and use you brain on what to do with them.
For that you need to threshold your predictions. Ideally you'd like your model to generate a bimodal distribution so that you can threshold without many false positives etc.
You should instead be asking, what are the odds that that X voters could change their vote.
Removing cookies for the domain doesn't help, because (doh) I've never visited it before.