I can't recommend serious use of an all-in-one local Grafana Loki setup
utcc.utoronto.ca
utcc.utoronto.ca
The criticisms brought by Chris are valid, and it's good feedback. I just wish it was done so in a more constructive manner.
I personally know what Loki is capable of, I've watched it grow from something tiny to something I'm incredibly proud of and in awe of.
I run Loki on several Raspberry Pi's ingesting 20-100GB a day, I also run Loki clusters on thousands of cores ingesting hundreds of TB's a day. It's an amazingly flexible and capable project built by an amazing team by a company I'm incredibly proud to work for.
That being said....
> Loki doesn't feel like it's been built to be operable by people who don't know its code and its internal details
How true this is... while I know for a fact there are hundreds to thousands of folks out there who are successfully running Loki (and thank you to those who share your success stories, it means the world to us), it can be very rough around the edges...
But we are working hard to improve this, we are doing the best we can.
All I really ask is for folks who want to give feedback about Loki, PLEASE DO! but please be patient.
The author of this article says they tried to work with us to improve Loki and that it didn't work out. I'm truly sorry for that, these posts are full of constructive feedback and I would really love to see this put into issues and pull requests to make the project better.
But honestly I sort of assumed that if you are going to manage hundreds of machines you should probably look at the more scalable configurations anyway. But if you are just doing a handful of machines it's more than capable in the local store only version.
- Single binary - Written in Rust (lightweight, fast and stable) - Use S3 bucket or Mount point - Visualize with Grafana
https://github.com/parseablehq/parseable
(founder here)
If you don't follow the one-true-architecture you will get bitten in a million ways.
* Log ingestion on the host pulls logs from the application/system/whatever, timestamps the logs itself (bc when you're interested in failure states do you really trust the log emitted by a broken app? Also because devs are famously bad a timezones), adds it's own metadata, and stores them in a local outbox queue.
* Local log ingestion determines where to send logs based on service discovery and periodically updates.
* Log ingestor ships the logs to a durable queue and flushes only after getting an ACK from the queue.
* Log processor reads from the queue and ships the logs off to persistent storage or a dead letter queue where you get an alert if it ever has something in it. Log processor only ACKs back to the queue only once it gets an ACK from the db. Logstash used to sin in this regard.
* Persistent storage treats logs as opaque blobs from the perspective of how they're physically stored. Indexes are time-window based depending on your volume, usually daily, and shipped off to different tiers / deleted on that basis.
This stack can horizontally scale indefinitely up to (and past since the queue backups allow you to temporarily fake more throughput than you really have) the throughput of your backing database.
I loathe how complicated and brittle the ELK stack is but they get this exactly right and if you implement it it becomes nigh-impossible to lose data. The market for "ELK style architecture but not the size of a 400 lb gorilla" has got to be huge but is seemingly untapped last I checked.
Why do we need the durable queue in between? Why not let the Log Ingestor ship the logs off to persistent storage?
Queue is useful if you want to write those logs into multiple places at once
We have DNS, we don't need log sender to have a service discovery mechanism on top of that. Set it to log server address and be done, scale at that point if you need to, we know how to do it.
Log processor doesn't need a fucking queue. Log sender does, for network reliability one. And that gives you ability to restart log processor quickly (only need to process current message in transit and close) and with zero impact (as long as you're down shorter than the logger's queue)
Only reason to add queue is if you have multiple readers for logs. That also conveniently gives you a form of QoS on log processor, if you read with equal rate from all sources the most spammy ones will hit their own internal queue limit first and wont cause other servers to miss the logs. Even then you might just op for the loggers sending things into 2 places at once.
The "shit logs" (whether by volume or needing messaging decoding) is a problem that's complex but IMO most of that should be within log processor, as it should be. That's also a good place to resolve any geoip or DNS if needed.
Having service discovery solves some issues.
* DNS TTL and applications holding on to DNS names indefinitely (prometheus, haproxy, nginx, and I bet your app somewhere all do this).
* Applications that don't support DNS record priorities.
* Serving different results to different clients based on their identity that isn't isn't random.
> Log processor doesn't need a fucking queue. Log sender does, for network reliability one
Yes. That's what the queue is for. The log sender also has a queue but as it lives on the host itself minimizing its use is how you don't lose logs on server crashes. If your architecture is the log processor accepts logs, and stores them in a queue for buffering then you've implemented the same architecture. But if that queue lives on the log processor itself then you risk data loss if that server dies. Having a shared queue in front of the pool of log processors, is simpler, has better throughput, easier to shard, and more reliable. Logs can't get stuck on a particular processor anymore because its lease will end and another worker will pick it up.
a) simplicity of rsyslog
b) monstruosity of ELK | Grafana | etc.
Somehow I like Prometheus (I think it's "simple"), but it's not enough to display and search for logs. Somehow, none of the companies I have worked for, have used "simple tools" like rsyslog to handle logs. They all used cloud (Datadog, New relic) or self hosted (ELK, Prometheus + Grafana). I wonder why (I guess it's because "money buys you simplicity")
I just want the following:
- on each machine I want to get logs from: install the agent (a simple binary) + simple /etc/myagent.conf. The agent forwards logs to my "main log server"
- on my "main log server": install the "log processor" (again, just a binary please!) + simple /etc/mylogprocessor.conf. The "log processor" shows me a nice localhost:9090/ web interface in which I can search for logs (indexed by any field I want).
Easy, no? My use case is not thousands of machines nor Terabytes of data logs per second. I just have a few machines and I don't want to deal with multi-clustered solutions or anything like that. Just 2 binaries! Does that exist?
The problem I have encountered that even "simple" (just my home NAS + few devices) setups require some log mungling to get useful info into whatever system uses it. Many apps don't have "log in JSON" option in the first place, and near-always there is no real standard in fields of that message either.
And also near-always I want to filter out or rate-limit some particularly spammy message or service just because I don't even want to look at it when browsing logs as it is just noise
> Easy, no? My use case is not thousands of machines nor Terabytes of data logs per second. I just have a few machines and I don't want to deal with multi-clustered solutions or anything like that. Just 2 binaries! Does that exist?
...graylog I guess ? I looked at it and it is apparently pretty integrated, but price on higher volumes made us do ELK on "actual big stuff"
With Splunk, ELK, Greylog, it feels insanely pokey. I know they have the parsers and such. At times I've kind of boned up on their search syntax but I've never gone "all in" with any of them, maybe because they all don't seem like a really solid long term solution. They seems to have a different kind of model than what I want, the time range is kind of nice but often times I won't have a time range until later. My model involves winnowing down the the data I want and then extracting pieces and viewing the data different ways. Am I just using all these tools the wrong way? Is my mental model off? Maybe it's a log consistency thing, it's always sort of a great day when you get "Error: abc failed because xyz and def." and that's the answer to everything. Many times I'll be spending time looking at logs and I'll notice an increase in a certain behavior happened before the outage happened and that's the give away.. Then a new grafana dashboard is created with a new metric to try and identify that before it happens again.
Loki kind of looks like it supports my method but again, I'm back to that "I haven't gone all in" with it problem. As I'm rambling, I've seen these sexy dashboards with like red/yellow/green lights and some latency graphs and cool looking stuff and then a little table of the last 20 "log messages" and maybe I'm used to looking at logs that you don't show in your dashboard or something like that.
They all feel like a square hole to my round peg. Maybe it's just me.
mine come in to an anycast IP on the network, one file per host, the syslog stamps the receive time at the start in "y-m-d-h-m-s+0000" format, in a y/m/d directory struture, bzip2 after a few days
I have a few scripts which I use to parse the logs and pull reports out (BGP drops/recover times for example), but most of the time tail/cat/sort/grep/cut/etc does the job. Where I differ is I use perl rather than awk.
Sure it doesn't scale to millions of terrabytes a second of minable personal information or whatever the average modern LAMP stack generates, but it currently records about 15G a day from 400 different devices just fine.
I imagine if you implement it "correctly" you don't lose data but my experience with Elastic Search has been horrible.
I've lost data many times, for things like logs reaching an artificial maximum number of indices, and ES shutting down, or just not being able to support the simple case of a log coming both as a json and as a plain-text; there's no setting to say "just cast to text if there's a conflict", it drops the log and the workaround is to find among the many outdated ES posts out there, a piece of Ruby code to fix that one case. Many other issues (I compiled a list of like 20 stupid things about ES and the many ways I've lost data and gave up adding stuff).
Like said in one comment here, it works well on billions of logs on one modest instance. And Grafana integration is on the way :)
https://github.com/quickwit-oss/quickwit
(disclaimer: I'm one of the cofounders)
Unlike index-free solutions like Loki or Parseable, Quickwit is built on top of a modern full-text search index (tantivy). At query time, Quickwit produces much faster results (all other things being equal: CPU, memory, etc.), especially when the volume of data to analyze is large or queries are complex (high cardinality values, aggregations). Quickwit also stores data in a columnar format, so it's also good at OLAP-style queries (no joins though).
This comes with a cost during ingestion; Quickwit is more resource hungry than Loki but can still ingest at 20MB/s to 40 MB/s on a commodity instance with 4CPU. Similarly, regarding storage footprint, Loki compresses logs better because it does not maintain those extra data structures. Still, a Quickwit index tends to be much smaller than an Elasticsearch index.
The next release of Quickwit (may) will be shortly followed by the publication of a benchmark against Elasticsearch/OpenSearch, and by another one later, against Loki. You'll be able to see for yourself.
Due to a buffering issue, Loki would exit in case of configuration error without printing any error message or anything at all.
There is definitely something weird about how the project is run.
N.b., I've run this in both AWS/GCP in a k8s scenario, against S3/GCS respectively as a long-term store.
I've also run Mimir, and they both work _fantastic_ in this deployment scenario (as described).
These components are typically build to be "cloud-native" which often means run on kubernetes. If you already run on kunernetes, grafana products are typically straightforward to run.
In general I think that is the target customer and use case for out of grafana cloud deployments. The all in one binaries are more toys to say hello world or maybe test things in local development.
Red Hat logging product manager says: "We made the decision to move to Loki and Vector" https://www.youtube.com/watch?v=QZ4Hv85lEJ0&t=938s
For example, another comment here talks about Loki locking up Docker if it's the logging backend and the container crashes. I suspect that wouldn't be possible, or would be less likely and more manageable with Vector in the middle because it will buffer. I've also dealt with normalising logs from different sources and it can be a pain, but Vector will do some or all of that already, reducing the requirements put on Loki.
I found it easier to setup and configure than Loki.
The UI is very basic for now but I’m excited to see what the future holds for this project!
To complete your description of Quickwit, It is a distributed search engine for logs and traces. It's written in Rust, ingest at speed, horizontally scalable, and separates compute from storage.
Last but not least, Grafana integration is planned for next month :)
It's notoriously difficult to know why ingestion of certain logs failed, to the point where I run a staging monitoring environment to debug issues like these.
*until recently... the new "time series" panel is a disaster
What's a good all-in-one cloud monitoring solution that might optionally deal with logs as well?
Log monitoring is not actually a priority at this stage, I just want something to track metrics, chart them and to alert me when the servers are on fire.
Something open-source adjacent and not crazily expensive would be best. As I said I know Prometheus, but for some reason it's so flexible and free form I really do not enjoy using it.
For your use case, Coralogix is awesome. It comes with a managed Grafana instance, if you wish to use that interface, or a custom dashboarding solution, metric driven alarms, release tagging, and much more.
You can find out more at https://coralogix.com/platform/metrics/
All in all I find a pretty new player on the market, but there is not much to compare it with. Given other product from Grafana I guess it will mature as well. There's always more mature projects like Graylog, but compared to that Loki is pretty small. But yeah, it got teeths, but dang it's fast!
I tried using it in a small k8s cluster on digital ocean. Initial installation using the recommended helm package was easy enough. However, it only saved a very short period of log data. I spent a fair amount of time searching docs and the web about how to increase the storage, with no luck. Such an obvious and common need should not be so difficult to configure. You should not have to deep dive, reverse engineer, and read source code in order to solve such simple problems.
A broader solution like Coralogix (https://coralogix.com/) is more appropriate if you're venturing into the unknown and you need more of a data discovery capability. The problem with most of THESE "all in one" platforms is they tried flexibility and feature breadth for cost. Coralogix is a wee bit different, and it gives users an awesome set of cost optimization tools, as well as a simple pricing model, to offer data discovery without a sudden increase in costs (or worse, surprise overages!).
That said: 1) I feel like Loki is languishing and not reaching its full potential. User experience needs a lot of work.
2) Grafana is a for profit company
3) Grafana sees its future in the margin rich Saas offering
4) open source is still supported, but only to a certain point and the rest is commercial. I wouldn’t expect material support if you’re not paying for it.
Among other licenses I found this in one of the sub folders (/EE).
> This software and associated documentation files (the "Software") may only be used in production, if you (and any entity that you represent) have agreed to, and are in compliance with, the SigNoz Subscription Terms of Service, available via email (hello@signoz.io) (the "Enterprise Terms"), or other agreement governing the use of the Software, as agreed by you and SigNoz, and otherwise have a valid SigNoz Enterprise license for the correct number of user seats. [...]
I guess it is for enterprise edition or something but it was not immediately obvious to me what parts are under EE and which parts are under the MIT Expat license.
Top level license:
Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software.
This was one of the reasons why we at VictoriaMetrics decided to start working on better solution for logs - VictoriaLogs [4].
[1] https://utcc.utoronto.ca/~cks/space/blog/sysadmin/GrafanaLok...
[2] https://grafana.com/docs/loki/latest/fundamentals/labels/
[3] https://utcc.utoronto.ca/~cks/space/blog/sysadmin/GrafanaLok...
[4] https://www.youtube.com/watch?v=Gu96Fj2l7ls&t=1950s
[5] https://www.slideshare.net/VictoriaMetrics/victorialogs-prev...
ingester:
chunk_retain_period: 30s
chunk_idle_period: 5m0s
chunk_block_size: 262144
chunk_target_size: 1572864
chunk_encoding: gzip
max_chunk_age: 2h0m0s
[1] https://grafana.com/docs/loki/latest/operations/storage/file...Moreover, the post speaks volumes about the potential pitfalls of "all-in-one" solutions. These can seem appealing due to their simplicity, but as this author has experienced, they often come with their own challenges, especially when things go wrong.
The friction with the Loki setup underscores the importance of having a good contingency plan in place for log management systems, as loss of log data can be a severe issue.
As for the author's point about the devs wanting you to use their cloud service, it's a common business model. Free or cheap software that's difficult to set up on your own, but with a paid service that makes it easy. The question is whether this trade-off is worth it for your specific use case.
This sucks, but it’s also why you take filesystem snapshots or perform a backup before upgrades.
> PS: Loki also has some container-ized multi-component run-it-yourself example setups. I don't have any experience with them so I have no idea if they're better supported and more reliable in practice than the all-in-one version (which isn't particularly, as we've seen). A container based setup ingesting custom application logs with low label cardinality and storing the actual logs in the cloud instead of the filesystem may be a much better place to be for using Loki in practice than 'all in one systemd journal ingestion to the filesystem'.
Author may be holding the tool wrong, using it for a scenario it was not optimized for
However, I would have used vector anyway locally as I prefer to have a centralised collector.
We store maybe 250GB worth of logs in each instance, and ingest an estimated 1-2k lines a second.
I'm considering right now to implement this.
Like this article might just be for show but its the first thing that came up ¯\_(ツ)_/¯ https://grafana.com/docs/grafana/latest/developers/contribut...
All of this stands in the way of contributing. And this is their decision to make, of course, but it is hostile to would-be contributors.
I could try explain to you what the purpose of a CLA is, but you could also easily put in the effort by going on Google.
Make sure to also look into the implications of contributing code with an OSS license, any license. That's as much contract as a CLA is.
Yes. Everything. That's the point of near-every CLA.
So the corporation behind it have option to close the code if they want to, taking your contributions with it.
Some corps might not ever do it, but any company is one MBA away from "what we can cut from OSS version and move to enterprise to get more customers"?
> Except for the license granted herein to Grafana Labs and recipients of software distributed by Grafana Labs, You reserve all right, title, and interest in and to Your Contributions. [emphasis mine]
You can always use the version of the code your contributions went into. You still own the code you wrote, and still retain the rest of theirs under its original license. You're just not entitled to future versions of the source like a copyleft license without a CLA would grant.
I understand not being excited to serve corporate interests (that's what a CLA does), but posting intellectually uncurious flamebait as a result makes for boring reading.
Do you hire a lawyer any time you encounter an unfamiliar OSS license? I assume you don't, or you'd have a hard time using any modern packaging ecosystem.
I can see someone giving their work to OSS project not wanting the corporation that took it to have option to just take it and close.
Linux kernel uses "Developer Certificate of Origin" which is basically just "I certify that I contribute stuff I have rights for". That is enough.
CLA is entirely to detriment of actual OSS
Am I missing something? Do you have any specific problems with the CLA? Is there an alternative option to a CLA that ensures the original copyright owner can continue offering the code base under multiple licenses after accepting an external contribution?
We did OSS for decades without CLAs and most projects still do not require CLAs.
Your good faith doesn’t hold up in court and I understand why they’d want to clarify ownership of the contributions. Just because we’ve always done it this way doesn’t mean people aren’t open to liability. Just because another project accepts the risk of a random contributor winning a lawsuit against them doesn’t mean Grafana should. I’m surprised CLAs aren’t more common.
I was personally surprised at how generous their CLA was with ownership rights for you and your contribution to the project. You retain a lot when contributing.
CLAs are becoming common because litigation is becoming common. It's a product of our times and mostly a safeguard for companies in case a person is able to slip in some malevolent code or write hate speech in docs, or if someone tries to claim copyright on docs/code.
The last contributor to the docs was an hour ago (at time of writing this comment) and came from a maintainer not employed by Grafana Labs.
Looking down the recent commits I see lots of activities from non-Grafana employees that have been accepted.
If there are specific issues with contributing docs or code please do point me towards them.
Every day I wake up and spend at least an hour reading and responding to issues and PRs in the Tempo repo.