The Syslog Hell
techblog.bozho.net
techblog.bozho.net
It's in general a trend for old unix tools to work better in reality then in theory something thats rare for more modern tools.
Sure it's nice been able to use more modern query tools and have graphing libraries available but syslog grep and awk does get the job done and dont require a lot of resources to set up and maintain.
just write a script, email the people regularly in charge with the CSV, and let the PHB make it look pretty.
thank you for the good laugh :)
That requires resources to design and keep running in practice. After dealing with a few systems for logs, I'd rather choose them for non-trivial setup now, than bare syslog and redo all of this from scratch.
Also, simplifiation does not never happen. E.g., I have a near trivial impl (under 300 lines of Nim) at https://github.com/c-blake/kslog { yeah, it may not be of very wide appeal..See the first point. :-) }
I feel a better factoring/separation of concerns is to disentangle file/data distribution from getting in-socket data somewhere persistent.
* Some stuff logs with syslog and some doesn't.
* Formats vary, oh the fun of implementing the different variations.
* You have to parse text dates into unix, when whatever wrote the file converted unix to text. It probably lacks milliseconds, which is all kinds of fun in a modern setting where there can easily be a hundred things happening during any second.
* Various edge cases. Where exactly does every field in the log file end? Can there be a newline? (yes, guaranteed). Can there be random binary junk (yup, sometimes).
* How do you keep track of where you stopped parsing? How do you deal with that the old log might have been removed and a new one with the same name now appeared?
* Dealing with compression, log rotation, race conditions.
It's simple on the surface. Actually writing a program that deals with all that stuff is bloody annoying, because none of it actually gets what you want done. You typically want to detect important events happening, or graphing something. Instead, 95% of the time goes on mind-numbing minutia dealing with parsing because the system was built for an admin using `tail` and `grep` in the 80s. It wasn't planned for a modern admin maintaining a few dozen computers each of which log multiple megabytes of stuff every hour.
Never understood people who complain about journald, because this stuff was a pain in my butt a good decade before journald existed. I certainly don't have any fond memories from dealing with it.
And personally I see the "hydra" as a benefit -- everything integrates well with everything else, because it's all designed to go well together.
I was trying to reach a compromise in a simple syslog standard so it would be easier to authenticate and analyze. And trying to make it good enough for non-*nix systems. Nobody else cared about this.
It was one of the worst time wasters in my life. It was all politics. The syslog-ng guys were adamant with their proposal which was a very, very over-complicated idea based on another standard (BEEP). And I strongly suspect the Cisco/Microsoft guys were intentionally trying to make the group not work in subtle ways. After months, I just left.
They eventually published RFC 3195. And it's barely used, of course.
It seems Cisco's implementation still uses DIGEST-MD5 for authentication.
> But what makes things hell is the fact that too many vendors decided not to care about what is in the RFCs, they decided that “hey, putting a year there is just fine” even though the RFC says “no”, that they don’t really need to set a host in the header, and that they didn’t really need to implement anything new after their initial legacy stuff was created.
It sounds like the author and I are doing similar work, so he knows my pain: if you make a product which can parse syslog, and somebody selects your product for parsing syslog, and they they feed it non-syslog logs from Company Y's product... it's now your problem, instead of Company Y's, even though you're perfectly capable of parsing syslog! Luckily, regular expressions and beer eventually get most things sorted out. :)
* Want to parse stuff? journalctl -o json
* Lots of stuff going on, need more precise timestamps? -o short-precise
* Want metadata, like the pid? It's in there.
* Want to know where to continue parsing? It supports cursors.
* Want to save disk space? It uncompresses logs transparently and can trim logs to whatever size you want.
Programs should be logging because there's some value in the information being sent to the log. If it annoys everyone and serves no purpose, the program needs fixing.
> You can't get back stuff you exclude by mistake.
Ok. That's fine - it's my system, my data, and my ass on the line if we lose data by mistake. I don't know see why a developer that knows nothing about my problem domain and the constraints under which I'm operating gets to dictate to me that my choice is wrong.
If you went ahead and implemented thoughtfully and sent a patch there’s a good chance it would get merged. Everyone has a idea about what “I just need basic filtering” and so accommodating anyone but not everyone is a recipe for making even more people mad. Telling people to just disable journald on-disk storage and forward to rsyslog and friends if you need complex filtering and don’t want to pay for double storage is a solution that feels bad but is actually generic enough to be useful.
Poettering has made it abundantly clear that he will not accept patches that implement arbitrary filtering. There has been very little engagement on his part with people who have legitimate issues that this would resolve, and appears unwilling to accept that there are use-cases where arbitrarily filtering log data at ingestion time has benefits.
Why anyone, myself included, would put the time and effort in to developing a patch that it's clear is ideologically opposed by the gatekeepers of that project is beyond me.
A colleague of mine run on evaluation of remote log collection using journald's remote support and syslog.
He found many problems with jornald remote logging. We have mobile connections and they can be poor occasionally. Syslog scored better.
Whether the evaluation was fair I have no it's idea. It would have required me the same time to do my own one. But I would have preferred to see systemd win. At least in the local case the benefits mentioned by the parent comment have convinced me years ago.
[1]: https://www.digitalocean.com/community/tutorials/how-to-cent...
* Want log input from processes that exited? Nope, information is lost (bug 2913)
* Want to validate the certificate of a remote log recipient? Nope! (bug 4092)
* Want to submit changes to the specification to help fix any of the above problems? What specification? The code is the spec!
I am a huge journald fan, but some of these bugs have been around for years and years. It's frustrating as hell, so I don't blame the syslog diehards.
If you've ever used reiserfs you'll know fsck isn't guaranteed to make things better.
Config was also way less painful than traipsing through whatever hellscape the FluentBit/FluentD configs are.
When applications on Windows fail they never think to generate an event, not even for something as simple as a "permission denied attempting to open file 'c:\blah'". Instead it's chock full of useless noise from daemons that activate every 2 seconds to poll something and then log that everything is still ok.
https://docs.microsoft.com/en-us/windows/security/threat-pro...
I have seen and hope to never see again worst cases such as the COM control for all of .NET Framework event log messages (everything from .NET system messages to just the mostly plain text storage from applications written in .NET) accidentally badly unregistered leaving all of the event log messages unreadable.
Don't even get me started on this... Microsoft is actually bad for this with even their own .NET-based enterprise applications. As well - it also makes gathering logs from production servers, then performing analysis on a different machine difficult, as that machine will likely have none of the dependencies required.
Text... text and more text, that is universal.
Whilst that's possibly true, at least you/we have the possibility to do so.
SystemD's built in log thing at least has a text dump feature, but god help you if the database gets corrupted.
What is this FUD nonsense you're attempting to spread here?
systemd-journald assumes journals marked as ONLINE when opened for writing are potentially corrupted, and renames them rather than attempting to write to them, which produces messages like "journal corrupt, renaming".
Those renamed journals still participate in journalctl queries, the data isn't lost. It's just a bit of wasted space in the interest of not risking actually corrupting an uncleanly shutdown journal by writing to it.
I guess I have been looking at Loki and is the issue in configuration or operation? If it’s harder to operate than ELK I’m not going to touch it though.
The hard part is the message field's content and format. It basically boils down to actions of thousands of individual developers. They will never agree on a format and logging style.
When a "standard" sticks around this long and needs to support so many legacy devices things can get a bit messy. At least syslog is human readable, while things may not be as machine parsable as you'd like, the info you need is usually only a few greps away.
The big problem with SNMP is that the MIBs have to be handled out of band. If there was some part of the standard where you could query the device to get its MIB in some standard format it would be so so much better. The daemon could be small because it wouldn't have to ship with hundreds of megabytes of data for devices built over the past 40 years. You wouldn't have to go on a hunt to track down where the vendor hid the MIBs for oddball and obsolete equipment, often times only available with a support contract on a website that was decommissioned years ago.
Alternatively it could have a query type that gives you a description of every field, so when you walk the tree you get all of the data that you would otherwise need the MIB for.
There are a lot of things you need in order to prevent the lazy from fucking up:
- A version. A lazy programmer might either ignore or hard-code a version number, but at least you have a hint as to which standard you're trying to conform to, and can retain backwards compatibility. (The new RFC has a version, but the old one doesn't, preventing interoperability)
- A format that's easy enough for programmers to understand, but difficult (or "feature-filled") enough that they won't attempt to implement it all themselves and will reach for real libraries.
- Extensions. Vendors will always want to do something different than everyone else, so if you don't add the option of extensions, they will either fork the protocol and make breaking changes, or try to sneak changes into other parts of the protocol/format.
- A standard reference implementation + tests. Make it easy for vendors to test their versions against another one, so the developers don't have to do busy-work like "read a standard" or "write tests".
- Think about the future. Does your standard include a specific width integer? Does it preclude a specific network payload size? Will addressing change in the future? Can your data payload support arbitrary binary data? Can your standard change later and still be backwards compatible?
A very well established and expensive product apparently thought the way to write json logs is something akin to:
printf("{date: \"%s\", msg: \"%s\"}\n", date, msg);
That turned out real good when msg contained quotes.Json isn't a magic bullet here. Most log pipelines end up with custom logic for all sorts of reasons that mostly shouldn't be an issue.
And delivery mechanism is so powerful! Online or batched, pull or push, with clear logic and great documentation. After looking at things like RELP, I just want rsyslog to go away and everyone switch to journald. It is time we stop losing syslog entries just because the net was down for a bit!
The cherry is it's much harder to fill the disk with logs.
No matter the service I'm using the same journalctl command to find what I want.
This is why syslog is completely disabled on all our servers
It’s like having a black box that forgets the last 30 seconds of the flight otherwise.
This is incredibly difficult to get right.
Most of the issues I’ve dealt with shaft the network before the filesystem.
Opening an EC2 console in AWS is still a simple and reliable way to find out what's going wrong with your instance. Wouldn't be possible if we didn't already have the convention to have the kernel, syslog, etc print to tty1.
Normally, Unix-like tools are not as painful as journalctl. But Poettering's interfaces are a gauntlet of unusual concepts and hidden inter-relationships, with no common examples or intuitiveness. And things like dbus make it worse, now that there are many parts of a modern Linux system that have no console interface, because nobody's written one for the particular application you need to view or change the right dbus settings. And /sys/ is literally a wilderness of random undocumented settings that are often the only interface to critical system functions.
Linux distributions are now a tiresome no-mans-land of overcomplicated mysterious crap. I'm willing to bet the major Linux distributions will be abandoned over the next decade for simpler systems that are cloud-native, mobile-friendly, and have less Kafkaesque interfaces.
You started with a strawman about caring about logs, and ended complaining about sysfs.
> I'm seriously struggling to understand the complaint here
The complaint was trying to explain a parent commenter's point about journald "They’re really nice until everything on the node is completely broken. Then they are a massive obstruction to access and understanding what went wrong due to the opaqueness."
Point: Logs are really annoying to manage on systems with journald.
Counter-point: Ship your logs somewhere else / you probably don't have these problems in real life
Counter-Counter: If we're not supposed to use these tools on our hosts, exactly why are they installed?
I would argue that logs are less annoying to manage on systems with journald, once you take the time to learn how to leverage the tools.
I would also argue that shipping mission-critical logs off-server is a worthy endeavor, regardless of logging system used.
I like journald because it lets me isolate logs for a particular unit without grepping and accidentally including output from unrelated services. It's faster to find the data that I need between time ranges rather than manually comparing time stamps.
I can count on one finger the number of times the journal has been corrupted on the servers that I manage, and it was because of hardware failure.
So far, I tried looking at logs from the dead system using “journalctl -D” - it seemed to work. And the way how the log files from each boot are always separate is pretty handy. Other than that, the only problems I have seen were having more to type and having to learn more commands.
Am I in for a nasty surprise?
In particular, I have just verified with "strace" that running journalctl with either "-D /var/log/journal" or "--file" option does not even open machine-id file nor a d-bus socket -- so whatever problems you had with machine-id would be gone with.
Also, when you said, that you "cp'ed the files off it", did you mean you copied off the contents of /var/log/journal? Were you able to open the files on the other machine?
If that is your concern, you can disable compression in the journald configuration so that the contents can be read with "strings" or similar tools.
I imagine the journald logs are just files at the end of the day and you can just read them with some tool?