I would like to see more about how they are formulating the problem, I know there is work lately in "sequential recommendation" that is focused on generating a sequence rather than scoring content items, I'd like to learn more about that.
I would like to see more about how they are formulating the problem, I know there is work lately in "sequential recommendation" that is focused on generating a sequence rather than scoring content items, I'd like to learn more about that.
Right. If your AUC is in the 90's you usually either have a bug, or a predictive task that isn't very challenging.
For some other like “is this an astrophysics or materials science paper?”, getting 0.95 or better was straightforward in 2005.
If you had to pick one number for judging classifiers that is the one.
"Accuracy" is a really bad one because it is not meaningful w/o careful thought about the problem. For instance if you have a diagnostic system for a disease that has a .01 prevalence, you get 99% accuracy if you say nobody has a disease.
Practically a person might want to operate a classifier with different thresholds depending on how bothered they are by false positives vs false negatives, the ROC shows what choices are have available with a particular and the AUC of the ROC gives you one number that characterizes the quality of the classifier overall. A system with an AUC under say 0.55 hasn't really learned anything at all, at 0.6 it is showing signs of life, looking at those curves I'd say Tik Tok is beating me solidly. In this case it is limited by the problem being fuzzy (something I thumbs up today might get a thumbs down tomorrow), something up in the 0.95 level could be attained for something like "is this article about the outcome of a sports game?"