Data mining Reddit posts reveals how to ask for a favor and get it
technologyreview.com
technologyreview.com
From the graph, it looks like only about 27% of requests are fulfilled in the best case (jobs). In which case I can do better than 70% just by constantly predicting "no".
(I assume that this is just bad reporting. I haven't read the reference.)
Edit: skimmed the referenced article. Average success rate is 24.6%. The 70% they give (well, 67.2%) is "the probability that a classifier will rank a randomly chosen positive instance over a randomly chosen negative one".
Even so, I still dislike the use of this figure, as ROC AUC overstates a poor score: it hints at the model being correct 67% of the time.
In fact, the model is only knowledgeably correct 33% of the time, and just guessing the rest (taking its binary-choice 'score' up to 67%).
What the algorithm performs poorly at is determining whether any single arbitrary request is accepted or rejected; that's the test that would require around 80% success rate.
Can't quite recall the paper that gives the details. Think it might be this one http://www.hpl.hp.com/techreports/2003/HPL-2003-4.pdf
"Since the AUC is a portion of the area of the unit square, its value will always be between 0 and 1.0. However, because random guessing produces the diagonal line between (0, 0) and (1, 1), which has an area of 0.5, no realistic classifier should have an AUC less than 0.5."
So it appears that not 0.8 but 0.5 is the randomness threshold; therefore 0.7 is not so bad.
This is model-generated data so they could put an infinite of markers at arbitrary locations. The use of markers implies to the viewer that is a data-generated figure, which it is not.
There could be a common cause to these two factors (a certain way of writing, for instance), that could explain the correlation. I can't imagine someone basing a judgement on a stat reported on someone's user page.
That doesn't mean score-viewing doesn't happen, but it isn't likely to be the only mechanism at play.
So, it's totally possible that, like others have said, people deciding which request to fulfill take karma/account age into consideration to avoid gaming the system with multiple accounts.
The tone of the article, I think, was suggestive of a causal relation, which surely doesn't (necessarily or plausibly) hold.
Makes "craving" seem to perform much worse than it actually does at a status of 0.
What exactly is a "standard machine learning algorithm"?
I'm sure that probably means that they used something from scikit-learn but "comb[ing] through correlations" isn't as simple as clicking the Go button.
The rest of the article does start to get into labeling, holding out a test set, and some of the data cleanup (the real combing working).
I guess I was just hoping for more detail of how it worked and not that it worked. I get that this wasn't meant to be a PhD thesis on supervised machine learning, but the mechanics of data analysis are really interesting as a process of discovery.
Curious to know what others think. How did the balance of how vs that work for you?
"Combing through the correlations", it seems, literally means calculating the (Pearson) correlation between two variables (success vs. an input feature) and adding some interpretation, as they do on p6 of the arXiv paper. For test/training data, it looks like they just used a 30/70% split rather than k-fold cross-validation and holdout, but I'm sure it makes no difference either way and in this case (as often) is trivial to design and implement. Presumably their AUROC could be increased just by dropping in an SVM or a Random Forest in place of the logistic regression.
From what I've skimmed of the paper you're over-estimating the complexity of the study.
And it has a live demo!
Internet access costs less per month than a takeaway pizza in my country. Libraries and other community centres provide free access to computers with broadband internet.
One of the surprises for me when, some years ago, I found myself in a developing nation (with unstable power supplies and non-potable tap water) was that every street corner in the town I was staying in had a (hand painted) advert for an internet cafe. It wasn't cheap compared to food prices - but since then globally food prices seem to have gone up and internet prices gone down considerably.
But I decided not to phrase it this way because this is obviously a very simplistic view and the people on the internet will tell you that in full detail even if you are aware of it. They might have free access to internet. It might be a temporary situation and not be economical to sell and later buy back the computer. Corn is much cheaper than rice. You have to cook rice and that requires energy. Rice alone is not healthy. Maybe they just do this to get in contact with others not because of having no money for food. They might need the computer for work. They just got robbed late at night, no money left but they still have a phone.