Some notes on Grafana Loki's new "structured metadata"
utcc.utoronto.ca
utcc.utoronto.ca
Loki has its idiosyncrasies but they are there for a good reason. Anyone who has sat there waiting hours for a Kibana or Splunk query to run to get some information out will know what I'd referring to. You don't dragnet your entire log stream unless your logs are terrible, which needs to be fixed, or you don't know when something happened, which needs fixing. I watch many people run queries that scan terabytes of data with gay abandon on a regular basis on older platforms and still never get what they need out.
The structured metadata distinction is important because when you do a query against that you are not using an index, just parsed out data. That means explicitly you're not filtering, you're scanning and that is expensive.
If you have a problem with finding things, then it's not the logging engine, it's the logs!
In Kibana, if something is there I will find it with ease and it doesn't take a lot of time to investigate issues in a microservice based application. It is also quite fast.
For grafana cloud and Loki you can close to a good usability with LBAC (label based access control) but you still need have many data sources to map onto each “team view” to make it user friendly.
What is missing for me is like in elastic a single datasource for all logs which every team member across all teams can see and you scope out the visibility level with LBAC
So good but not perfect! When we have the time we'll look for alternatives
Elastic was kind of a resource hog and much more expensive for the same amount of data.
That might be dependent on your use case though.
2. It has a convenient and simple query language.
3. It works very well with traces and metrics.
the pain part:
1. It struggles to query logs over a wide time range.
2. Its indexing (or labeling) capabilities are very limited, similar to Prometheus.
3. Due to 1 and 2, it is difficult to configure and use correctly to avoid errors related to usage limits (e.g., maximum series limits).
IMHO, Loki query language is the most inconvenient language for logs I've seen:
- It doesn't support calculating multiple stats in a single query. For example, it cannot calculate the number of logs and the average request duration in a single query.
- Its' syntax for aggregate functions is very unintuitive and is hard to use, especially if you aren't familiar with PromQL.
- It requires putting an annoying "|=" separator between words and phrases you are searching in logs.
- You need to use a hack with JSON parsing when filtering or stats calculations on log fields is needed.
Also out of box configuration sinks 1TB/hr quite happily in microservices mode.
Could you share Loki config, which can deal with 1TB/hr volume of logs?
Modern columnar SQL such as ClickHouse are 10+ times more efficient in real-world use cases.
I'm a CEO and founder of Quesma, which, let's use Kibana with ClickHouse: https://quesma.com/
Forever free, source-available license.
why do you need storing JSON-encoded string inside log field? It is much better parsing the JSON into separate fields at log shipper and storing the parsed log fields into Elasticsearch. This gives better query performance and may also reduce disk space usage, since values for every parsed field are stored separately (this usually improves compression ratio and reduces disk read IO during queries if column-oriented storage is used for per-field data).
I tried explaining this at https://itnext.io/why-victorialogs-is-a-better-alternative-t...
I was not able to do that with the log shipper. If I configured parsing, then the not-JSON messages would get dropped.
At the end of the day ELK was throwing us a bunch of roadblocks in order to solve problems we didn't need solved. Maybe if we were trying to build some big analysis layer on top of our logs that would've been nice. VL has worked great for our use case of needing to centralize and view logs.
This is explained in more details at https://itnext.io/why-victorialogs-is-a-better-alternative-t...
important because the title includes _new_
We are still waiting for a compelling implementation that will show the way.
I think that expressiveness, performance and level of fluency by base language models (i.e. the amount of examples in training set) are the key differentiators for query languages in the future. SQL ticks all those boxes.
There are a lot of fundamentals in observability, but there are very verbose in SQL:
- rate operator, which translates absolute value to rate, possible with SQL and window functions, but takes many lines of code
- pivot, where you like to see the top 5 counts of errors of most hit-by-error microservices plus others over time
- sampling is frequent in observability and will be useful for LLMs, it is a one-liner in SQL with pipe syntax, even customizing specific
I actually believe LLM gen AI plays extremely well with pipe syntax. It allows us to feed partial results to LLM, sampling as well as show how LLM is evolving queries over time. SQL troubleshooting is not a single query but a series of them.
Still, SQL with pipe syntax is just syntactical sugar on SQL. It let's you use all SQL features as well as compiles to SQL.
Maybe that would work as well with traces and logs but IMO the problem space is quite different and not sure how much value we’d get from a unified language where some subsets only apply to parts, ie traces and logs and metric, as opposed to spiritually similar but distinct languages.
I'm very glad to know about VictoriaLogs issues, so they could be addressed quicker. If you hit such issues, then please file them at https://github.com/VictoriaMetrics/VictoriaMetrics/issues or publish an article highlighting these issues.