Announcing Graylog v1.0 GA
graylog.org
graylog.org
https://web.archive.org/web/20130302051347/http://graylog2.o...
Sad to see it's gone all "professional" now :(
The web service had some hilarious bits too. Does it still say "enraging gorillas... | mounting party hats" etc when you log in?
This statement from the page makes me conclude that others were indeed able to crash the system, perhaps without even trying too hard.
My guess is that the statement should be re-worded.
The meaning should be that with v1.0 it is finally very hard to crash the system by overloading it. Previous versions were easier to crash by filling memory in load spike situations for example.
After waiting for a 500MB something docker image download and running the container halted the host OSX machine while looping with errors (failed API connections, failed to load SIGAR, ...)
The screenshots looked great but the first steps experience were a deal breaker.
It would be awesome if you could help us improve the Docker image by creating an issue at https://github.com/Graylog2/graylog2-images/issues with the error messages you've encountered when starting the container.
Seriously, the Docker image and the VM images are there to help people getting started and to quickly try out Graylog. We also offer regular OS packages in DEB and RPM formats (supporting Debian 7, Ubuntu 12.04 and 14.04, and CentOS 6) and a quick setup application (for a simple demo setup of Graylog and its dependencies). So choose your poison and be happy!
See also: http://www.infoworld.com/article/2885752/log-analysis/open-s...
That being said, we're aware that many people don't like the dependency on MongoDB and we'll work on that.
The graylog-web-interface connects to the graylog-server REST APIs and that is it. You can manage and monitor the whole system from the graylog-web-interface.
Both components are always released together.
(The documentation may need a "History" page)
Their search interface and field extraction let you perform complex queries on your logs. We used the following query to identify users that were having a bad time on the site, by counting slow queries by user and adding them up:
"slow query" source="/var/log/application.log" | rex field=_raw "in (?<time>\\d+)ms" | search time>2000 | rex field=_raw "User: (?<login>[^\\s]+) " | top login
You can do much more complex queries--it's better to think of Splunk as a temporal database than a log aggregator.
Does anyone know if Graylog or other logging solutions can do the same thing? Splunk is amazing, but it's annoying to manage the infrastructure and it's crazy expensive.
All open source.
So I can maybe tell easily where users had problems with my app, or I can even use it as some kind of analytics program for my app?
source:your-app AND http_status_code:>=500
This would give you a list of all HTTP 500s of "your-app".Or:
user_id:12346 AND http_status_code:>=500
... to see all errors a user caused after customer care called you, reporting an error the user got. Stacktrace is there immediately without having to find the right log file on the right web server.Its like splunk but not as poweful
I was initially drawn to the fact that it took five and a half years to get to 1.0. God knows how hard it must've been to work on it to perfect it all these years, heck spending five months on an app is tough enough. I would've given up reading about it otherwise if it hadn't been for the five years thing. I just think that, despite all it's capable of, if the descriptions were more in layman's terms it could reach wider audience.
It actually turns out that this is exactly what I was looking for since I'm about to have my app out, and if I hadn't asked the question, I would've just ignored it all together.
While I'm very thankful for the additional explanations, I still think that nice illustration or a simple video of telling how it can solve real world problems will help people understand it better.
- We have a server that runs dozens of websites. When the load spikes, we can quickly get a count of recent log entries for all the sites. The site with the anomalously high number of entries is where we start troubleshooting. This could also be automated as "anomaly detection" that sends us alerts, but we haven't configured that yet--happens rarely.
- One of our servers got hacked. Running log searches helped us pinpoint when it happened, which site was "patient zero," and how the bad guys got in.
- We launched a new site and forgot the Google Analytics code, which we didn't catch right away. We were able to run a report from the server logs to approximate the traffic data that GA missed.
Having all the logs feed into a centralized service made it easier and faster to find the information we needed across a bunch of websites, as opposed to working directly with Apache log files.
We looked at using ELK (Elasticsearch, Logstash, Kibana) to do the same thing "for free," but decided we did not want to manage a complex software stack to help manage a complex software stack. :-) We'll take a look a Graylog I'm sure, but there is something to be said for paying for this as a service--one less thing to worry about.
Sentry is great for monitoring all sorts of log emissions.
I suggest you try out Graylog and let us know about your findings! I'll make sure to look at Sentry again.
Thanks for the work!
We do not recommend it running on extremely small platforms like a raspberry pi, but a very small VM (we have OVAs and other virtual appliances ready) is able to process a lot of messages already.
Only if your load balancer supports UDP, which most don't. You'll most likely need to use DNS load balancing in this case unless you're sending GELF messages with TCP.
Does anybody have some experience using it in production?
The solution seems neat. It'd be nice to know who is using it in production, how big is their environment, etc. I couldn't find much information online.
We are using NXLog on our Windows servers to send event log messages in GELF format. This allows us to truly delve in and search event logs so much easier than what we normally would be able to in Windows.
It's really helpful being able to pull up all requests for a single user across the cluster. Or track the status of all components of a Sidekiq job. Or see which API requests are most popular. Etc.
Honestly, I have no idea what we did before we started using it. It's an absolutely essential tool.