Because classification is a regression problem.
Think about it for a second. You want to put together a tool to tell which class an input belongs to. You have training data you can use to build your tool around. Your training data is already divided into sets that belong to a specific classm Your goal is to put together a model that can tell you what's the closest class your input belongs to by comparing with how close your input is to elements of the training data belonging to a specific class.
What's your strategy?
Well, one of the textbook strategie starts by specifying how you measure the distance between elements of your training set, and from that point you work on putting together a function that not only minimizes the distance between elements of your training set but also, when used to evaluate elements of a training set, works well in telling the type of elements of the training set that are closest to them. Then you assume the class of your input element is the same as the class of the elements of the training set that are closest to them.
In the example above, the minimization step is... Yes, regression. You use regression to fit your model to your training data so that it is able how close your input element is to elements of a certain class, and then outputs how close it is to each of the classes.
You're simply wrong. Regression is a tried and true classification technique. Posting personal and baseless assertions don't refute that. I mean, pick up pretty much any textbook on supervised and unsupervised learning. You always end up with an approach which boils down to having training data, put together a trial function, apply a minimizer to fit trial functions to training data, and evaluate the resulting model by running trial data through it. Fitting trial functions to training data has a name: regression. Minimum squares has a very precise interpretation both in linear models and in probability. There is no way around it.
What I said was that regression metrics are not good for evaluating usefulness in classification problems. An R2 of 0.01. The fact that there are mitigating circumstances to justify why this might not be the case is not evidence that R2 is still a good thing to use. It’s actually evidence of the opposite.
With classification problems we are concerned more often with the ordinality of estimated probabilities vs outcomes and/or the calibration of a model.
> Posting personal and baseless assertions don't refute that.
Your mom didn’t refute me either.
You're voicing very opinionated takes while showing considerable ignorance on the topic.
Classification problems are solved with trial functions adjusted to the training set through regression. These trial functions,once fitted, represent membership functions. They are essentially interpolation functions that, say, converge to 1 when close to elements of the training set of a specific class, and 0 for members of all other classes. In grey areas where elements of the training set are sparsely distributed, these membership functions can output values in the middle, because they are interpolating.
That's literally data mining 101.
> What I said was that regression metrics are not good for evaluating usefulness in classification problems.
That's simply not true. It's like saying models that do not fit the data are good approximations of the data. Another way to put it is praising the accuracy of a broken clock because it's spot on two times a day.
> The fact that there are mitigating circumstances to justify why this might not be the case is not evidence that R2 is still a good thing to use. It’s actually evidence of the opposite.
I don't think you have an adequate grasp on the subject, neither classification problems nor linear models. Therefore, your personal baseless assertions don't mean anything nor bring any value to the discussion.
You’ve twice now tried to explain how fitting a model works in response to me stating that regression “metrics” are not well suited to describing classification “efficacy”. If anyone is making a basic mistake here it seems to be you failing to understand the difference between a “regression model” like a linear regression, and a “regression metric” like R2.
I’ve stated my semantic meaning of regression vs classification. You can Google this to see it is not a fringe view. It’s been standard for over a decade. Eg https://math.stackexchange.com/questions/141381/what-is-the-...
> Another way to put it is praising the accuracy of a broken clock because it's spot on two times a day.
Accuracy is a classification metric kiddo
Here’s an introductory Wikipedia article if you would like to learn more about what metrics are appropriate for binary classification evaluations:
https://en.m.wikipedia.org/wiki/Evaluation_of_binary_classif...
The normal R^2 formula can't be applied to a logit/probit model; instead you use an alternative such as McFadden's or Cox & Snell pseudo R-squared. I'd be interested to see what value they take for this example.
Linear models are sometimes used even in models with many independent variables since it can be shown that the coefficients in a linear model are unbiased estimators for the average partial effects of any non-linear binary response model.
The blog post was rebutted by Seth in the comments 2 weeks ago, same as Colin Percival's HN rebuttal, and Andrew didn't reply. It seems like a weird goof. Andrew was "buggin".