Learning from Imbalanced Classes
svds.com
svds.com
If you want to work with really imbalanced data, try working with data from the LHC. There were on order of 1000 Higgs in a year, while there are around 600 million proton-proton collisions a second.
[0]: https://soundcloud.com/linear-digressions/whats-the-biggest-...
Check out slides 21 and 22 from [1]. There are parts that will process 4 PB/s (!)
[1] http://www.slideshare.net/SparkSummit/distributed-data-proce...
I had two research projects on finding low quality - deleted and closed - questions on Stack Overflow [0, 1]. Since, these question classes always suffered from imbalanced class problems (2% closed and 8% deleted) - I decoded to under sampling of the majority class. Later, I used an ensemble - Gradient Decision Boosting Tree, Random Forrest - for classification. However, rather than taking bootstrap samples - I took various random samples from the majority class (equivalent to the minority class) and built many ensemble classifiers for each random sample. In the end, I checked the variation on the final classification results. It just seemed intuitive to me. I had no idea about Wallace et al.!
[0] Denzil Correa and Ashish Sureka. 2014. Chaff from the wheat: characterization and modeling of deleted questions on stack overflow. In Proceedings of the 23rd international conference on World wide web (WWW '14). ACM, New York, NY, USA, 631-642. DOI: http://dx.doi.org/10.1145/2566486.2568036
[1] Denzil Correa and Ashish Sureka. 2013. Fit or unfit: analysis and prediction of 'closed questions' on stack overflow. In Proceedings of the first ACM conference on Online social networks (COSN '13). ACM, New York, NY, USA, 201-212. DOI=http://dx.doi.org/10.1145/2512938.2512954
Are you suggesting that: 2 percent of samples are positive, drawn from p(x|y=1)
98 percent of samples are drawn from a distribution p(x), but may be either positive or negative?
The setting that you described above is called "positive and unlabeled (PU)" learning. This paper: http://cseweb.ucsd.edu/~elkan/posonly.pdf is one of the seminal articles on the topic (although the equation on the bottom of page 214 contains a statement that may not necessarily hold true). There are quite a lot more recent papers on this topic.