Show HN: Visualize how HN/Reddit talk about your company and products
brandimage.io
brandimage.io
- I typed "Chevrolet" and it says "mpg" and "venerable" have negative sentiment, while "underachieving", "supposedly", "kilometers", and "cars" have positive sentiment. "Stunning" and "awkward" are both neutral.
- I tried "Mazda" and it was only slightly better. "Praise" is neutral, "regret" is positive, and "touring" is negative.
- I tried "Petzl" and only shows 16 words, all neutral, and that includes stop words like "had" and "and" and "etc".
Those are at least companies with unique names. For companies with less common names, you might be lucky to get keywords related to the right company at all.
I think there can be good uses for word clouds, but they are few and far between, and this isn't one. Just make 3 lists, side by side, and title them "positive", "neutral", and "negative". Instead of font size, put the more common words higher on their respective list. The only reason I can see to use a word cloud here is to hide how bad the analysis is.
For example, in the case of Mazda where you say that "regret" is classified as positive, if you look into which message it comes from you can see the original sentence: "Buy a Mazda, you won't regret it :)"
I agree with you that the word cloud is not useful on its own, and this is why you can click on a word to see the actual messages. Think of the word cloud as merely an entry point into a more detailed analysis by a human.
Thanks for the feedback.
One interesting issue, you might need to read the context a bit more to gauge if the post is actually talking about the Brand. For example Target is a common english word, and if you look at the results, only 1 usage was actually referring to the store Target.
One nice bonus that would be useful, I'd like to be able to compare sentiment by Subreddit, if i'm marketing on Reddit I would not target "Reddit" as a whole.
As for the subreddit, it's already on my next features list :)
For a particular term I get the following words having "positive" sentiments: boycott, punishment, greed, backlash, protesters, damaging, upset.
How in hell is that positive?
Also consider the lack of labeled data for HN and Reddit messages: I had to use Twitter messages to train the classifiers.
This is the reason why I tried to play with BERT to see if I could get a model to generalize well from only Twitter messages. From my experiments, if you activate BERT (which makes the app much slower), you should be able to get 60~70% accuracy.
It's not perfect, but not too bad as well if you are getting averages over a large amount of messages.
Overall it's still a work in progress, I expect to greatly improve the accuracy over the following weeks!
For example, you can type in non-brand words as well. I typed in "houses" and the word "homeless" came up in green!
With a brand, facebook, I got this word "amiriteguyze" in red and clicking on it
Negative 11/19/2019, 12:13:31 PM
facebook is bad amiriteguyze?!?!?!?
Why is that even a word that would show up in the word cloud? I can't imagine it was entered a bunch of times. I can't intuit any correlation between the colors, sizes, or words themselves that show up in the clouds.
To prevent those words from appearing, I was thinking to implement some dictionary-check to only allow for meaningful words. However this approach also have drawback as you restrict people's words and can miss important new concepts.
Thanks for the feedback.
Over Quota
This application is temporarily over its serving quota. Please try again later.
I found an issue with "+" symbols in the brand name. Check out this URL - https://brandimage.io/insight/c++?source=hn
If "+" symbol is present, the UI won't show anything.
For large companies there are some proprietary solutions for this. Example:
Disclaimer: family member works there so that’s the reason I’m aware that this niche exists.
It's one thing if companies want to find out what people are saying about them so they can improve. However, I'd be surprised if that's how it is used.
This feels more like a way to measure the effectiveness of their astroturfing or low-key marketing efforts.
The back-end is just Python/Flask and I use the free Algolia and Pushshift.io APIs to source the messages from HN and Reddit (big thanks to them!)
The easiest library to do that would probably be scikit-learn with their ComplementNB class: https://scikit-learn.org/stable/modules/generated/sklearn.na...
For the data you can use the SemEval 2017 Task4-A dataset (around ~10K labeled tweets): https://github.com/cbaziotis/datastories-semeval2017-task4/t...
positive: "closes", "binary", "btw", "disclaimer", "bullshit"
The word cloud seems fairly useless.