Mute uninteresting log noise with machine learning
blog.machinebox.io
blog.machinebox.io
I've been loving Kibana for filtering and reporting on log data in flexible and insightful ways, including automatically generated charts for certain data sources.
Let's say there's a service failure and I want to know what the service has done prior to the failure. I wouldn't want a classifier to filter the logs in that case, so that use case is out of the picture. What other use cases than filtering are there for this? Maybe as a way to provide feedback to developers to fix the log messages, as in "this thing that we log all the time can be determined to never affect the process of trouble-shooting our services, and the classifier thinks it's noise, so we'll remove it".
cat output.log | tr -d '[0-9]' | sort | uniq -c | sort -n
This is a fairly useful way of removing relatively useless information such as timestamps and line numbers when you're looking for rare or unique events. The alternative, I think, is to do a bunch of awk or sed magic, which isn't really fun for anybody. It's especially useful in a time crunch when there's an ongoing outage.But don't take my word for it. Try it yourself!
Another application would be a security camera that detects unusual events without having to train it on actual burglars.
That’s what we do here anyway, it’s worked well for us:
https://github.com/Morgan-Stanley/hobbes/blob/master/README....
What if I never see a critical log because the trained model decided that it is unimportant? How is such a situation generally solved in the industry?