First Monthly Challenge: Elasticsearch
engineroom.trackmaven.com
engineroom.trackmaven.com
Another dataset with similar problem is Google Mail takeout. Theoretically in an mbox form, there is apparently enough quirks in there to not be parsable by the standalone 3rd party libraries. Somebody said python might be able to do it, but I haven't seen the confirmation specifically for Google Mail.
If you manage to do either and document the steps/share the code, I'll personally sing your praises at the search engine workshops/meetup/lectures I do (usually Solr rather ES).
http://googleappsdeveloper.blogspot.com/2011/11/parsing-expo...
https://docs.python.org/3/library/mailbox.html?highlight=mai...
https://gist.github.com/ptwobrussell/8791064
http://www.kryogenix.org/days/2013/12/19/searching-my-email/
http://batleth.sapienti-sat.org/projects/mb2md/ (if going from mbox to mailbox is easier)
More info w/ docs: https://www.inboxapp.com/
Open-source version: https://github.com/inboxapp/inbox
http://www.elasticsearch.org/blog/kibana-4-beta-1-released/
It exposes quite a bit of the new aggregation functionality added to ES 1.x.
Example: Say you've got a dashboard showing hits on a web service. You've got a pie chart showing HTTP return codes, a bar chart showing response times, and another few graphs and charts detailing various data out of the requests themselves.
You could click on, say, the "500" in your return code pie chart, and then every visualization on the page would redraw and show you stats for just requests that that were 500s. (What's unique about the requests that return server errors?)
Or turn it around - click on the section of a chart that denotes requests that took longer than 100ms to process, and now you see info about those requests only. (What makes these long-time requests so special?)
This was a jaw-droppingly awesome troubleshooting tool, and now it's gone. I hope they return it before Kibana4 gets out of beta!