For a particular term I get the following words having "positive" sentiments: boycott, punishment, greed, backlash, protesters, damaging, upset.
How in hell is that positive?
For a particular term I get the following words having "positive" sentiments: boycott, punishment, greed, backlash, protesters, damaging, upset.
How in hell is that positive?
Also consider the lack of labeled data for HN and Reddit messages: I had to use Twitter messages to train the classifiers.
This is the reason why I tried to play with BERT to see if I could get a model to generalize well from only Twitter messages. From my experiments, if you activate BERT (which makes the app much slower), you should be able to get 60~70% accuracy.
It's not perfect, but not too bad as well if you are getting averages over a large amount of messages.
Overall it's still a work in progress, I expect to greatly improve the accuracy over the following weeks!
For example, you can type in non-brand words as well. I typed in "houses" and the word "homeless" came up in green!
With a brand, facebook, I got this word "amiriteguyze" in red and clicking on it
Negative 11/19/2019, 12:13:31 PM
facebook is bad amiriteguyze?!?!?!?
Why is that even a word that would show up in the word cloud? I can't imagine it was entered a bunch of times. I can't intuit any correlation between the colors, sizes, or words themselves that show up in the clouds.
To prevent those words from appearing, I was thinking to implement some dictionary-check to only allow for meaningful words. However this approach also have drawback as you restrict people's words and can miss important new concepts.
Thanks for the feedback.