Grafana Loki – Like Prometheus, but for logs
github.com
github.com
Grafana move really fast. Very practical. It first if I remember use ElasticSearch/InfluxDB to store its config itself. Then move to its own SQLite. Then MySQL/Postgres. Eventually add alerting. Add multiple data source.
Nowsday my team even use its with SQL for a cheap version of Periscope/Mode.
Then now they figure out this log thing. ElasticSearch is great but try to run it yourself. On a 50 nodes k8s, I bet your first ES config will down after 10mins the first time FluentD come up and send all the log under the sun since the cluster come up
Given that Grafan guys know how to deal with visualization, I trust them to deliver a great experience again for log.
Look at their history, I'm going to bet hard on this.
Tell me about it. I have not one, but multiple k8s clusters (some over 50 nodes), sending data to a single Elasticsearch cluster. A lot of tuning was done, and it's not yet perfect.
Elasticsearch is amazing, but it requires external support, and the tools available to do that are not on par. Curator specifically – it's dumb as a rock and is very limited on what it can do. Even a hot/warm architecture is a challenge to do on curator alone, you should ideally have custom scripts to manage it. Which is a shame as this is one of the "reference" architectures by Elastic. Whatever Curator does should really be part of ES itself.
The Kibana + ELK combo won't give you things like log tailing (logtrail is hackish and hard on the servers). The log forwarders are horrible to work with (be it logstash, filebeat – or worse: fluentd).
Some days it just works and you are happy. Some other days, you wonder why we moved away from syslog senders and text files...
A thing many overlook are restart it. You cannot just go and `systemctl restart` it, once that occurs, constant rebalance happenning because it puts the load to the rest of system and eventually bring them down, all around :(. That make it harder to operator in K8S when it needs time to detach/attach volume, plus the overhead of overlay network.
So I myself didn't know about a perfect tunning. But the log search on Kibana is super helpful though.
That's why I kind of sold on this Loki thing. I have a good feeling that software in Go tend to require less config/tunning compare with Java. Eg, InfluxDB.
> Some other days, you wonder why we moved away from syslog senders and text files...
I wonder the same. I actually like to grep log with tail more than a crappy web ui like Kibana, Sumologic etc.
Tailing log in a webui is clunky.
I have this same problem with a 3 node cluster, no Kubernetes. This was using Elasticsearch 2.3 and it worked for a year or two and then all of a sudden, any network interruptions would cause the thing to get split brain and all indices would go red. It would take 8+ hours to go back to green and usually required a reboot of all servers or else it would never recover.
It was happening often enough that the decision was just run a single ES node since the total data was only 25GB :(
Just as an FYI, Index Lifecycle Management (ILM) is landing in 6.6 and provides some of the features that Curator supports: https://www.elastic.co/guide/en/elasticsearch/reference/6.x/...
Basically, it allows you to define things like hot/warm/cold architecture, rollover, retention, etc at the index-level, inside Elasticsearch. Will be landing as beta
It is becoming part of ES. Index Life Cycle management UI is coming up in Kibana.
> The Kibana + ELK combo won't give you things like log tailing (logtrail is hackish and hard on the servers)
Recently, in 6.5.0 log tailing was added. Now one can look at streaming logs live.
> The log forwarders are horrible to work with (be it logstash, filebeat – or worse: fluentd)
I'd like to know the problem you faced.
Why are they horrible to work with? Why is fluentd worse than the other two?
Put a buffer in front of it, pace the initial ingestion and Bob's your uncle. If that was the main problem operating an Elasticsearch cluster it would be all roses.
I like the E?K stack a lot and I admit that operating ES is not exactly easy, but the initial ingestion issue the easy to solve.
But taking a step back, thinking about how many components you have in the system just to make it work. Compare with Telegraf/InfluxDB/Grafana stack, which is for metrics, operate it is seamless.
That's my high hope for Loki. I hope eventually Loki can just as easy to operate as InfluxDB.
- Can I use Loki without Prometheus? I'd like to feed raw logs to it, with custom-generated metadata. I don't want to have to use Prometheus, nor InfluxDB.
- Can I edit (modify) metadata for some old log line after it was already inserted? Specifically, I need to be able to rebuild metadata later, if I add some new "filters" to my logs (I want to be able to apply them retroactively).
- Can I run aggregate queries on the metadata (sum, avg, min/max)? If not, what can I do with the metadata? Can I graph the metadata on normal Grafana graphs?
- Is the text of the logs compressed? If yes, what compression algorithm is used? If not, why?
- Where can I find the API docs (or at least the API source code) for Loki?
2. Hmm, please open an issue regarding this with the use-case that prompted this? This is not a use-case we have but if this is important for more people, we'd be happy to support it!
3. You can only select based on the metadata, the metadata is just a set of kv pairs. Like {app=cassandra, namespace=prod, instance=cassandra-minion-0}
4. Yes, we used gzip. Please see the sheet referenced at the end of the design doc [0] to see what we compared with.
5. It's mainly protobufs right now [1] over HTTP, but we'll be adding more docs soon. Mind opening an issue for this so that we don't forget?
[0] https://docs.google.com/document/d/11tjK_lvp1-SVsFZjgOTr1vV3... [1] https://github.com/grafana/loki/tree/master/pkg/logproto
While a lot of our logging needs seem like they would be fulfilled by this system — because we attach trace IDs to log messages, and because (at least in Payments) you can usually find the appropriate trace ID by searching for a Payment ID, which could be annotated too — there are definitely many times I've copy/pasted the text in quotes from a log-generating line of code in a Java or Go file, to find out if it's being executed, or as a handle into a subsection of code/logging.
In the linked design doc, you include a motivating tweet near the top, saying, “just give me log files and grep, I am dying”. But unless I'm misreading things, there's no `grep` here. Right?
I'm guessing you could narrow down (using metadata) and then grep, but if the narrowest metadata you have is app name and time range, you're still going to be grepping over a lot of data…
Will make it more obvious. Davkals has an iteration of the UI that makes it a separate field, which will also help.
While this will definitely be slower than something that indexes the contents, you'll be able to store much more in Loki at much lower costs.
For your hosted service, will you put in place any restrictions on, for example, the size of the time range that can be queried?
Also for your hosted service, will the degree of parallelisation vary by pricing tier?
Looks like the Go regex lib, which isn't super performant, so it could potentially be improved if it ends up being an issue.
http://docs.grafana.org/features/explore/#logs-integration-l...
There is a lightweight CLI for Loki too, you don't have to use the Grafana UI: https://github.com/grafana/loki/blob/master/docs/logcli.md
Three questions:
1. It sounded like it was easy to set key-value metadata in the promtail tool. Things like hostname, availability zone, etc.
But can you also append metadata via the log-files themselves?
Basically, our log files are JSONL (http://jsonlines.org/) and look like this:
{ "user-id" : "abc", "client-type" : "mobile", "etc" : "…", "messages" : ["Parsing incoming http request", "Saving user data, valid model", "Finishing up http request, sending response to client"] }
{ "user-id: " "def", "etc" : "you get the idea…" }
{ "third log line here" : "and so on" }
Can we ingest these into Loki and have the "user-id" metadata appended as key-value labels to each message?2. Sometimes it makes sense to group related log lines. In the example above, we'll have ~20-30 log lines from a single http request. It would be convenient if we somehow could group them together, for example based on a unique X-Request-Id value. And then use that in the UI to see all related log lines together easily.
3. We currently store metadata about log-lines that is numeric. For example, things like request time. Will it be possible to query on that type of numerical value, to e.g. find all the log lines where requests took more than 500 ms?
The current design inherently limits cardinality to the number of pods you have running and the various labels applied to them.
For 3 I'd say that doesn't make sense with this design. I'd suggest taking a look at something like scalyr.com. You can configure a log parser and then your logs become both queryable, and you can create time-series on numeric fields such as request time and look at the 99th percentile, min, max, etc.
> Loki is meant to be complementary to existing solutions like Elasticsearch and Splunk that do full text indexing
Can you elaborate a bit on this please? If I'm already using Elasticsearch, Splunk or the like, why would I want to add on another, less powerful logging service? (not trying to be a dick, genuinely want to understand why I'd want this!)
When debugging you'd want as much info as possible and you'd want to be able to simply tail + grep it. I've been told to log less because the amount I was logging would burn a hole in the pocket when deployed to production. Sometimes, at scale people only send WARN (maybe even only ERROR) and above in production which is sometimes not enough when trying to debug a system on fire.
While ELK does a great job of indexing the contents, and if you depend on it for BI, you should definitely still keep it for use-cases where just select+grep won't suffice. But for just storing logs, and being able to select, stream and grep laaarge quantities of logs, Loki will come in handy.
Or maybe you'd send everything to Splunk, but with a small retention period, whereas you'd send everything to Loki but with a 90 day retention period (or whatever)?
1. Batching logs will allow better compression ratios and bigger blobs (which means lower per-operation costs), but must be balanced with the risk of data loss - what is your strategy here?
2. Will this handle multi-line logs? Say a regex matches part of a multi-line log, will it then return all the lines for that log?
3. Can you add your own labels, or are you limited only to those assigned by Loki?
4. If you are limited to labels assignd by Loki, how will you handle labelling as you expand out of k8s and accept logs for other sources (e.g. syslog)?
Will Loki be able to parse the log generation time out of logs (where it's included, and it usually is), or will it only use the log ingestion time for time range searches?
Grafana is offering a logging UI for Loki in the upcoming v6 release called Explore; you can enable it on the master builds right now, see https://grafana.com/blog/2018/09/21/grafanas-explore-ui-taki...). It makes is super easy to start exploring and sifting through your logs.
Also, Loki allows you to push regexp matches server side, so you can distribute the "grep" among multiple machines for extra points :-)
I've been kind of building something similar myself, but this looks much better.
The free cloud demo does look it is getting hammered right now ..
Exactly! Elastic is a really powerful system but I think there is growing sentiment that its overkill for container logs. This is exactly where we see Loki really helping - almost complimenting Elastic even.
> I've been kind of building something similar myself, but this looks much better.
:blush: thanks! We'd love you input on Loki too...
> The free cloud demo does look it is getting hammered right now ..
Yeah, looks like I'm going to be scaling that all day... Our motivation for over the free service for the next few months is to really battle harden the system - anyone sending us data is really helping us iron out the kinks and improve the open source. Its early days, but expect it to get much better over the next weeks and months..
I don't know much about kubernetes, but Loki looks super interesting for our application logs too.
Could someone maybe ELI5 what the difference is between "container logs" and, well, any other kind of logs? Don't most dockerized application send their stdout to Docker, and aren't those, then, the container logs? What kind of logs _aren't_ "container logs" and therefore a better fit for ElasticSearch than Loki?
Thanks :-)
Having said that, this will work for any logs as long as you can tag the logs meaningfully. We'll soon be releasing packages for all major distros and journald.
[0] https://www.timescale.com/
[1] https://blog.dbi-services.com/optimized-row-columnar-orc-for...
Real columnar databases like MemSQL or Clickhouse are a different beast -- for example they give very good column-wise compression in the dataset, which can save dramatic amounts of space. They're also good very for use cases like this, since they're heavily optimized for OLAP style workloads.
There is also cstore_fdw which does offer columnar, compressed storage for PostgreSQL as a foreign table, but it won't hold a candle to something like MemSQL or Clickhouse in terms of raw performance. Maybe one day.
Ultimately it's not about columnar storage or partitioning support, though, it's about the data and the queries you want to run on it, in what amount of time. Timescale can do pretty good for a lot of cases like this I bet, and I'm investigating it myself for a project.
I have been evaluating TimescaleDB and my company currently uses Splunk, ES, and Prometheus. I'm going to be giving Loki a go this weekend for our k8s cluster.
Your explanation here really helped clarify some things I've only explored a bit.
Postgres full text search, and Timescale extension would be ideal for logging kind of application. I could just set data retention to something like 2 months to keep the data volume simple and manageable.
But, all "data explorer" like tools for SQL databases obviously focus on having their end users write SQL. I'm hard pressed to find a good search interface that talks to Postgres at the backend..
It's still focused on SQL but the "Question" flow is quite simple, and doesn't require you to write SQL. You can, of course, write raw SQL if you want to and have those results still visualised.
Metabase seems excellent for building a dash board composed of several SQL based questions, and to generally have a searchable list of questions(which are internally SQL queries). But, searching for logs seems to not be one of its use cases...
Last I used(2 years ago), Postgres weren't able to keep up with it. Though I just use default RDS and didn't do anything to tune it.
I asked myself that question over the past few months and ended up building a proto-type logging pipeline using Postgres/TimescaleDB.
I wrote about it in a blogpost; https://www.komu.engineer/blogs/timescaledb/timescaledb-for-...
Logging in plain text - syslog
Logging to ELK(+indexes)
Logging to a DB like Postgres
Systemd's binary logs
This Grafana Loki
10+ GB/day is big to deal with as you say. But, I'm wondering it would be around the same regardless of your choice above..
With all of the choices though, we can always set staging rules and stage old data out to archives.
Personally, I'm struggling to find an equivalent search UI for Postgres.
We're heavy users of Grafana, Graphite, Prometheus, and Elasticsearch, and while the latter started in a BI role it's expanded to take on pretty much anything we can throw at it. However, there's still tons of system and service logs we're not gathering yet, because the effort to get them to Elasticsearch and store/maintain them is not worth it, especially since the value is not always clear until well after the pipeline is set up.
I'm definitely excited to see what loki can do for us, and just upgraded one of our test Grafanas to nightly to start playing with Explore, looks great!
What we tried to get right is the seamless switch from a Prometheus to Loki where it's retaining the labels of the query to essentially find the logs that come from the same e.g., "job". The assumption is that you need to be consistent with your relabelling rules of Prometheus and Loki.
Thanks for the feedback, and feel free to reach out with any questions.
Before clicking to look at the comments, I totally thought this was about some interesting creation mythos with a character that taught humanity how to harness wood from nature to build the first wood dwellings.
I mean, that's totally something that could feasibly get posted to HN and do well.
That's how I learned the phrase.
There are only two hard problems in computer science: cache invalidation, naming things and off-by-one errors.
Keep track of all items (like "main page", "search results for $blah") that component items (like "body of post 234", "avatar of user 3540") are used in, and when a component item changes, invalidate its cache and that of all items it was used in.
I'm pretty sure the quote was originally talking about function, variable and data structure naming, in which Google may or may not be help. Really, I've always found naming problems in that are more to do with correctly predicting the future, since the most egregious problems I encounter are when something that used to be named at least passably has changed over time to be very confusingly, if not outright misleadingly, named.
We used Grafana + KairosDB and turned the logs into tags essentially.
Our entire productions system had easy 'taps' that you could annotate to monitor everything about our crawler.
For example, number of HTTP requests, their status codes, the language of the content.
We also record intersections of the tags like lang+domain.
The downside of a system like this is that you have to know all your metrics apriori... If not and you need them at runtime you're out of luck.
The UPSIDE is that you use like 1/100th of the total size you would originally need for raw logs.
What we found is that you quickly converge on the tags you need and then you don't end up adding many more.
Everyone says 'disk' is cheap but in our situation our logs outpace the amount of data we would collect. We'd have 1000s of petabytes of logs by now.
> We will be able to pass on these savings and offer the system at a price a few orders of magnitude lower than competitors. For example, GCS cost $0.026/GB/month, whereas Loggly costs ~$100/GB/month[0]
Maybe I'm misunderstanding, but I don't see how Loggly costs anything like $100/GB/month? It seems clear to me that the $99 plan costs $100/GB/day (or $100/30GB/month). Not exactly cheap, but nowhere near as expensive as the design doc reckons.
The rationale still holds and one of the primary reasons we built Loki is to have a an easy to manage scalable "open-source" solution.
I personally have been told to log less at a previous company because of the associated costs of logs and I don't think I utilised anything beyond grep (with a few exceptions). I personally feel the trade-offs are right and a simple greppable/tailable solution that is cheap is missing from the eco-system.
What you should use depends on what tools you have available. We currently use Datadog, so we use their API for publishing events. We can then search for them in the event stream or overlay them on dashboards.
You could look at the datastores Grafana supports for annotations and choose whichever of them you're comfortable operating.
One example - https://github.com/joelparkerhenderson/architecture_decision...
And not shutdown at all! Everything we did at Kausal is now part of Grafana and Grafana Cloud - David's PromQL-completion UI is in Grafana v6 as the explore view, and Cortex is the backend for Grafana Cloud's hosted Prometheus...
And Cortex is now a CNCF project: https://github.com/cortexproject/cortex https://www.cncf.io/sandbox-projects/
Do you have integration with Docker ? Especially Logspout ?
There is a free preview right now: http://grafana.com/loki We'll be introducing a paid tier in the next few months, once things are more stable.
> Do you have integration with Docker ? Especially Logspout ?
Not yet! That sounds like a great idea though - mind filling an issue in the repo for it so I don't forget?
We're soon coming out with a blog post explaining the motivations and architecture in detail. But yes, the pricing is one of the motivations to build this, and as logs will be stored in S3 or an object store, the cost will several orders of magnitudes less.
[1] https://speakerdeck.com/davkal/on-the-path-to-full-observabi...
From Loki's readme, It still stores the full log (but compressed), and indexes values in each logline.
Inside kubernetes the metatadata would be podname, namespace, deployement name, container name, etc...
We batch together sets of lines into compressed "chunks", and only add index entries per chunk. In practices, chunks have 1000s of lines in them - so the index tends to be many orders of magnitude smaller than in other systems.
To be clear: Loki indexes metadata about the streams, and the streams themselves are indexed by time. We don't full-text index the streams though. We think, when combined with metrics, this represents a nice trade off in terms of complexity vs features. This simplification not only makes Loki easier to understand, but also scale and operate.
[0] https://docs.google.com/document/d/11tjK_lvp1-SVsFZjgOTr1vV3... [1] https://github.com/oklog/oklog
What would be the reason that promtail is not tailing log files for all of the running pods?
We initially started using fluentd as the agent, but we found its metadata "enrichment" facilities weren't reliable enough - we'd get log lines without the pod tags, for instance. For something like Loki, which depends really heavily on the metadata for index, this was super important. So we wrote promtail.
[0] https://github.com/grafana/loki/blob/master/docs/api.md [1] https://github.com/grafana/loki/blob/master/pkg/logproto/log...
With graphite protocol the metric stream is already plain-text with just 3 fields per line (path value time), its pumped over a socket to a collector.. but that's where the hard stuff starts.
Store it where ever you want, this isn't a magical datastore that makes things faster, use clickhouse-client, whatever it doesn't matter.
There is a widening disconnect between the unix way and how new projects are created.
Logs contain a lot of detail that are not appropriate for metrics. Logs output tends to be expensive, it costs a lot of time and IO to emit, compress, send, store logs. Logs also tend to contain PII things. Usernames, IP addresses, trace IDs, etc. We don't want that kind of cardinality in metrics.
Metrics are another thing. They tend to be internal counters inside a single process. With multi-threaded languages like Go, Java, C++, we can track metric data in memory very efficiently.
For metrics, see https://openmetrics.io/
What you are proposing for deriving metrics from logs exists in plenty of forms. Check out https://github.com/google/mtail , for example. Of course, you still have to store them somewhere, visualize them somehow, etc. That won't simply be part of your shell pipeline.
I think I'll try out this new logging solution. I only want basic indexing.
Unless you have a small load, six months from now you'll say "I've spent the last 6 months setting up and tuning an ELK cluster. It looks like it might be ready for production workloads now, but let's make it a beta for the time being – I don't want to see a crash like we had last time..."
There is another school of thought that is "Always Be Sneaking Dicks Into Your Art". I have some respect for people in the second camp.
But on a serious note the developers should indeed look into this. I recall this comment from HN [0] which highlights why neutral branding and naming are important.
That being said I'm also not really sure what it's supposed to look like anyway (is it an L?) so it's probably worth changing for that as well.
Immature or not, the are at least a reasonable number of people that see just that when they look at the logo.
At the end of the day, is all immaterial, they change it, or they don't, I'm not sure it really matters.
If I'm not mistaken, some designers even go to lengths to make their mundane logos more memorable.
It's not likely I'll be forgetting the name of Loki any time soon so their branding probably (?) succeeded?
They are either going to have to take it off the readme, shrink it, or likely, redesign. Maybe just rotate it 90-120 degrees clockwise.
But Loki is not a replacement for Elastic - we don’t have any complex query support, and we don’t do full text search. Elastic is great for analytics or BI, but we think Loki is the way forward for your container/pod logs.