I am familiar with logistic regression (studying to be an actuary) but the problem is that my documents are unlabeled: I literally have 2000+ unlabeled documents and a list of positive and negative words.
I'm willing to label 10% (~200) documents but how should I 'score' them? On a scale of [-1,1]? Just {-1,0,1} for negative, neutral, positive? How do I create a training set?
Also can you point me in the direction of some of the implementation details e.g. how do you translate text into logistic regression model?
I would also like to implement POS tagging and 2-grams (e.g. "not bad" != "bad") - any advice on incorporating this into the system?
Thank you for your input!