There is so much wrong with the api design of sklearn (how can one think "predict_proba" is a good function name?). I can understand this, since most of it was probably written by PhD students without the time and expertise to come up with a proper api; many of them without a CS background. Compare this to e.g. the API of google/guava.
For example https://www.reddit.com/r/statistics/comments/8de54s/is_r_bet...
Case in point, sklearn doesn't have a bootstrap crossvalidator despite the bootstrap being one of the most
important statistical tools of the last two decades. In fact, they used to, but it was removed.
Weird right?
...
> We don't remove the sklearn.cross_validation.Bootstrap class because few people are using it,
> but because too many people are using something that is non-standard (I made it up) and very very
> likely not what they expect if they just read its name.
> At best it is causing confusion when our users read the docstring and/or its source code.
> At worse it causes silent modeling errors in our users code base.
...
Oh man, I thought of another great example. I bet you had no idea that
sklearn.linear_model.LogisticRegression is L2 penalized by default.
"But if that's the case, why didn't they make this explicit by calling it RidgeClassifier instead?"
Maybe because sklearn has a Ridge object already, but it exclusively performs regression?
Who knows (also... why L2 instead of L1? Yeesh). Anyway, if you want to just do unpenalized
logistic regression, you have to set the C argument to an arbitrarily high value,
which can cause problems. Is this discussed in the documentation?
Nope, not at all. Just on stackoverflow and github.
Is this opaque and unnecessarily convoluted for such a basic and crucial technique? Yup.
Or the following: https://www.reddit.com/r/haskell/comments/7brsuu/machine_lea...