While a lot of our logging needs seem like they would be fulfilled by this system — because we attach trace IDs to log messages, and because (at least in Payments) you can usually find the appropriate trace ID by searching for a Payment ID, which could be annotated too — there are definitely many times I've copy/pasted the text in quotes from a log-generating line of code in a Java or Go file, to find out if it's being executed, or as a handle into a subsection of code/logging.
In the linked design doc, you include a motivating tweet near the top, saying, “just give me log files and grep, I am dying”. But unless I'm misreading things, there's no `grep` here. Right?
I'm guessing you could narrow down (using metadata) and then grep, but if the narrowest metadata you have is app name and time range, you're still going to be grepping over a lot of data…
Will make it more obvious. Davkals has an iteration of the UI that makes it a separate field, which will also help.
http://docs.grafana.org/features/explore/#logs-integration-l...
While this will definitely be slower than something that indexes the contents, you'll be able to store much more in Loki at much lower costs.
For your hosted service, will you put in place any restrictions on, for example, the size of the time range that can be queried?
Also for your hosted service, will the degree of parallelisation vary by pricing tier?
Looks like the Go regex lib, which isn't super performant, so it could potentially be improved if it ends up being an issue.
There is a lightweight CLI for Loki too, you don't have to use the Grafana UI: https://github.com/grafana/loki/blob/master/docs/logcli.md
Three questions:
1. It sounded like it was easy to set key-value metadata in the promtail tool. Things like hostname, availability zone, etc.
But can you also append metadata via the log-files themselves?
Basically, our log files are JSONL (http://jsonlines.org/) and look like this:
{ "user-id" : "abc", "client-type" : "mobile", "etc" : "…", "messages" : ["Parsing incoming http request", "Saving user data, valid model", "Finishing up http request, sending response to client"] }
{ "user-id: " "def", "etc" : "you get the idea…" }
{ "third log line here" : "and so on" }
Can we ingest these into Loki and have the "user-id" metadata appended as key-value labels to each message?2. Sometimes it makes sense to group related log lines. In the example above, we'll have ~20-30 log lines from a single http request. It would be convenient if we somehow could group them together, for example based on a unique X-Request-Id value. And then use that in the UI to see all related log lines together easily.
3. We currently store metadata about log-lines that is numeric. For example, things like request time. Will it be possible to query on that type of numerical value, to e.g. find all the log lines where requests took more than 500 ms?
The current design inherently limits cardinality to the number of pods you have running and the various labels applied to them.
For 3 I'd say that doesn't make sense with this design. I'd suggest taking a look at something like scalyr.com. You can configure a log parser and then your logs become both queryable, and you can create time-series on numeric fields such as request time and look at the 99th percentile, min, max, etc.
> Loki is meant to be complementary to existing solutions like Elasticsearch and Splunk that do full text indexing
Can you elaborate a bit on this please? If I'm already using Elasticsearch, Splunk or the like, why would I want to add on another, less powerful logging service? (not trying to be a dick, genuinely want to understand why I'd want this!)
When debugging you'd want as much info as possible and you'd want to be able to simply tail + grep it. I've been told to log less because the amount I was logging would burn a hole in the pocket when deployed to production. Sometimes, at scale people only send WARN (maybe even only ERROR) and above in production which is sometimes not enough when trying to debug a system on fire.
While ELK does a great job of indexing the contents, and if you depend on it for BI, you should definitely still keep it for use-cases where just select+grep won't suffice. But for just storing logs, and being able to select, stream and grep laaarge quantities of logs, Loki will come in handy.
Or maybe you'd send everything to Splunk, but with a small retention period, whereas you'd send everything to Loki but with a 90 day retention period (or whatever)?
1. Batching logs will allow better compression ratios and bigger blobs (which means lower per-operation costs), but must be balanced with the risk of data loss - what is your strategy here?
2. Will this handle multi-line logs? Say a regex matches part of a multi-line log, will it then return all the lines for that log?
3. Can you add your own labels, or are you limited only to those assigned by Loki?
4. If you are limited to labels assignd by Loki, how will you handle labelling as you expand out of k8s and accept logs for other sources (e.g. syslog)?
Will Loki be able to parse the log generation time out of logs (where it's included, and it usually is), or will it only use the log ingestion time for time range searches?
Grafana is offering a logging UI for Loki in the upcoming v6 release called Explore; you can enable it on the master builds right now, see https://grafana.com/blog/2018/09/21/grafanas-explore-ui-taki...). It makes is super easy to start exploring and sifting through your logs.
Also, Loki allows you to push regexp matches server side, so you can distribute the "grep" among multiple machines for extra points :-)