Sentiment Analysis – Improving Bayesian Methods
github.com
github.com
Also, I would like to see a better metric than simple accuracy. Maybe harmonic mean of precision and recall?
Otherwise, great to see more stuff like that in more languages (Kotlin, here).
Edit: Oh, hang on, I missed this bit:
>> more developed tokenizers that understand language
"understand language"? Really? What about "can handle"?
I would definitely qualify this as "picking at nits" haha :)
When I run the code I see something like: 38/100 = 42.0% - 38.0% accuracy, 90.47619% accuracy of rated data.
90% is nice, but it looks to me like a tradeoff between precision and recall, where we just label those samples with the highest probability of pos/neg...
The primary take away was meant to be that a clustered Bayesian classifier (each trained on random samples of training data) offers significant improvements over traditional single Bayesian classifiers. Nothing revolutionary, just wanted to share a practical example that out performs most Bayesian implementations I have came across. I also wanted to share the results of model pruning and that tokenization matters.
I originally coded this without much intent on sharing, so the post write-up could probably use a lot of work. :)
Neat idea wrt. using a Naive Bayes classifier here!
This idea originally came to me after explaining Random Forests to someone as a way to improve their Decision Tree. On a whim, I built out a prototype and it offered improved results. :)
How does it compare to the state of the art methods that used those datasets? Is it better or comparable?
If you set a reasonable confidence threshold, I expect mid 80s to low 90s.
In both cases, the clustered Bayesian classifier outperforms the single classifier.
I'll run a quick test to demonstrate it's absolute accuracy. Though it definitely performs more than well enough to use in production, and is relatively little effort. It especially performs well in aggregate stats (>99% accuracy)
The main purpose is to demonstrate techniques to gain improvements over traditional/common Bayesian methods.