Logging: Unsexy, Important, and now Usable.
roadtofailure.com
roadtofailure.com
I disagree. Large software companies already exist in this space: Splunk, LogLogic, Arcsight, etc. The author Sounds like they reinvented Splunk in particular. (disclaimer: I work for Splunk and can see everything in your datacenter with a few finger presses. Call me Geordi LaForge.)
From my perspective the main hurdle to log aggregation/correlation is not scalability. If splunk doesn't cut it for your performance needs, you have probably hit the price point to where you can afford a loglogic or similar appliance.
Instead the barrier to entry is in the number of applications supported by a particular log archival product, and the ability to correlate across the different applications.
As I'm sure you know at this point, adding support for log types is a painstaking task. Most vendors punt on this and tell customers to do it themselves.
If there is a niche available to you as a startup I would think that it would be in offering a very low turnaround time in supporting new log types. For example: give us some log sources and we'll support and categorize your logs with our service.
As for running in the cloud on large datasets, I think you'll find that most customers are not going to want to double or triple their outgoing bandwidth -- In addition to concerns from a security compliance standpoint.
That being said, good luck in your venture. Logging is a mess, and could certainly use some clean up. :)
Um... what? Have the authors gotten stuck in a time vortex and been dropped off before, y'know, awk? Much less perl, or any of those new-fangled toys.
I mean: writing scripts to do log analysis is a pretty fundamental problem for server-side development, and lots of very smart people have spend the last two^H^H^Hthree decades working on tools to address the issue.
I don't even see how this (indexing the entries across a Hadoop cluster) is all that useful. In general, you don't do log analysis by asking "give me all the entries that match this pattern", you do it by walking them in order and extracting one or two fields from each line and building some kind of result data structure. This thing would be fine if you were asking for all the logs messages that mentioned "coffee", I guess. But what if you wanted a histogram of hit counts per page per day-of-week?
For analytics, you're right, search is only part of the equation. That's why we make MapReduce easy to use on a cluster. You can write Pig or Hive scr
We also have templates for common data formats (and ways to roll your own) so you can turn unstructured log text into structured data, so that a histogram of hit counts per day-of-week is just a few lines of a script (or maybe even a search).
If it's sensitive data, I'd recommend just spinning up your own cluster and installing the tool.
My sense has been that at least the low end of the market is increasingly preferring a search-oriented log management architecture versus a database-backed, query-heavy architecture because organizations are familiar with the search metaphor, the overhead of managing a search index is less than the cost of database administration, and the unsophisticated use case (i.e., free-text search rather than advanced query syntax) is increasingly common. Smaller organizations also rarely have the maturity to deal with logging systematically since it requires a pretty systematic approach to infrastructure. That said, the tradeoffs we made may or may not apply in a terabytes-per-day environment, I unfortunately can't speak to it directly.
Unfortunately, I don't work in a TB/day environment (though that would be fun), so I can't comment on that directly, but our experience with Splunk has been positive at this scale and I see no reason so far why a distributed installation with multiple 100GB/day nodes would change that. The only customer complaints we have are basically that long-term searches don't return instantly, and that's primarily a factor of I/O speed and can be resolved by setting up summary indexing (which requires forethought about what data you're interested in... therefore, it very rarely happens in my organization).
All that said, good luck in your endeavors!
I just wish Yahoo would open source Everest (their multi-PB column store DB based on PostgreSQL) -- this would be ideal for building an open source Splunk competitor.
Re. an open source log indexer: Agree, this is a space that will eventually become dominated by open source tools, used particularly by startups and small businesses. I think most people ignore this use of a MapReduce-like framework because they conceptually understand how it could be used, but 99% of all work is in the implementation, not the idea. And as of yet, I don't believe there has been a specific implementation beyond what companies like Shopify are doing where they add nice GUI tools on top of awk and grep (which admittedly is probably good enough for most people / business on this forum).
If you have huge datasets, we'd love to hear from you. If even to chat for a few minutes about what your data looks like. Ping me at info@drawntoscale.com -- maybe we can help!