See also the design rationale: https://docs.google.com/document/u/0/d/1IC9yOXj7j6cdLLxWEBAG...
It's basically comparing an append-only fixed-format text file with a queryable database. Of course the former is going to be more performant on writes.
A huge feature list doesn't matter much if the software is bad. Given that journald still irrecoverably corrupts its logs even after all these years and -apparently- suffers from substantial write amplification, I'm gonna stick with my ordinary syslog implementations, thanks.
Also, in regards to your original comment:
> This issue report feels like it ought to be accompanied by a fix.
This smells a lot like the "Don't come to me with problems, come to me with fixes." order that a lot of mid-level and director-level management really loved to make five, ten years back. [0] While this sounds like a hard-charging order and gives the impression that it's bringing much-needed discipline to lazy-ass subordinates, the truth of the matter is that its actual effect [1] is to get people to shut the fuck up about the company's problems. The job of most mid-level and nearly all director-level management is to do inter-organization coordination. Most low-level folks don't come to mid- or director-level management with problems they can solve. After all, if they could solve them, they would... talking to folks in that layer of management is usually a huge drag. Most low-level folks only come to these sorts of folks with issues that require inter-organization coordination!
So, yeah... the only obligation of someone who's reporting a bug is to provide a reasonably well-written bug report accompanied with reproduction instructions and diagnostics that are as clearly written as is reasonably possible. Reporters of performance bugs are under no obligation to suggest how to eliminate the bug... especially not if the project they're reporting the bug against has both paid maintainers and claims it's the infrastructure on top of which all Linux systems should be built. Corporate-backed projects that make such grand claims put themselves in a radically different class than the one that covers hobby or small-time projects.
[0] AIUI, it came out of Google, but my understanding might be incorrect.
[1] ...regardless of whether or not that effect is intentional...
On LKML you might get cursed out, but if your fix is solid fix, it has high chances of getting through. Regardless of how true it would be in reality, the atmosphere created by upstream is that I do not expect the same with journald unless you convince redhat management
> e.g. hook in a senior engineer that you know is intimate with the system.
unless, of course, you don't know said engineer because you don't even work in the same company, you're a just a user seeing a problem in an app you use
Or they work in a different part of the fairly-large company that you both work for.
I guess emmelaich either missed the part of my commentary where I talked about handling inter-organization communication, and/or has never worked at a company where it's simply impossible to know everyone who could reasonably be relevant to the stuff that the company works on.
At least triage it with your best effort.
(Also, I'm not entirely sure this is a bug so much as an inefficiency report. Consumption of storage space isn't a documented or promised behavior, nor is the behavior technically incorrect. It's just wasteful.)
Nobody's talking about an obligation here.
Also you: [1] If you think you can do better than journald's existing format, propose a new one with tests to prove it.
The fact that you're personally powerless to enforce an obligation doesn't change the fact that you're talking about creating an obligation.> I'm not entirely sure this is a bug so much as an inefficiency report.
Performance bugs absolutely are bugs... especially when they're in a long-running corporation-backed project that presents itself as the project atop which all Linux systems should be built.
Per our Guidelines:
> Please respond to the strongest plausible interpretation of what someone says, not a weaker one that's easier to criticize. Assume good faith.
That's what I did. So, right back at you.
Please explain, because it's not coming across that way. It's coming across as needlessly picky and combative, especially after I told you what I meant (or, at least, didn't mean) and you continued to argue with me.
I'm neither required nor strongly obligated to do so, nor do I see significant personal benefit to doing so. So, I will not.
However, these days it's quick and easy to command an LLM-based system to generate most any text. Before one demands an explanation from a human, perhaps one should machine-generate a plausible-sounding explanation and present that along with one's demand for a human-synthesized one?
Way to double down on the “needlessly picky and combative” angle, dude.
Sit and consider the points of similarity between my refusal-shaped reply and the entire conversation we had prior to it and you might find enlightenment, in the style of those classic Zen tales. Perhaps an LLM-based tool might be able to assist you in this, or maybe it will be distracting and misleading.
GLHF and all that.
Doctor, heal thyself.
Modern drives will read data at 500MB/s, sometimes even more. Your log files are approaching tens if not hundreds of gigabytes before a sequential read stops being a viable option. Tinies modicum of partitioning by date and source basically makes it a complete nothingburger.
Like ultimately it isn't even fast, journalctl is so bad at rendering text that it's approximately still as slow as seeking in a 400 MB .log-file using less.
Anyone with any sort of scale where you actually need indexing immediately drops journald and uses loki or elasticsearch instead. Journald is not even remotely a contender in that space.
That I agree with. I don't personally use journalctl much these days, particularly now that practically everything's a container and all their logs are getting shipped off-host for indexing. But I get why, 14 years ago, it was considered a good idea.
We have some services at work that log to text files and some to journald.
The log volume to file is >> the log volume to journald. Yet `rg query myservice.2026-08-01.log` seems to always wind up being faster and better than something like `journalctl -u myservice.service --since '2026-08-01' --until '2026-08-02' -g 'query'`. (The tab completion and discoverability is also better, I guess)
On a file of ~1M lines (350 MB), ripgrep returns a query matching 22k lines in 1.9s.
(Separately: We have gzip log rotation for old logs. A log with 1.9M lines (106 MB gzipped) on disk can be searched from cold in 1.3s by ripgrep, whilst still being ergonomic. Maybe a different compressor is better?)
On a service unit filter over ~218k lines, `journalctl -u service --since <...> -g query` returns 4.8k matching lines in 4.127s (and this is after having warmed the cache by returning the unfiltered query cold, in 11.9s. I don't have root to clear the disk cache
I should have ran the filtered query first but ah well, since the difference is already so large it doesn't matter that the filtered journalctl query gets an unfair advantage)
Comparing input rates:
ripgrep (plain): 576k line/sec
ripgrep (gz): 1.4M line/sec
journalctl (-u, --since, -g): 53k line/sec
I don't have a great deal of understanding of journald's internals, so perhaps there is some variable here that is unreasonably unfavourable to journald.Transcript
user@machine:~/log$ time rg query aservice.log | wc -l
22628
real 0m1.859s
user 0m0.109s
sys 0m0.909s
user@machine:~/log$ wc -l aservice.log
1071606 aservice4.log
user@machine:~$ time journalctl --since '2026-08-16' -u myservice@1.service | wc -l
218661
real 0m11.914s
user 0m10.729s
sys 0m0.798s
user@machine:~$ time journalctl --since '2026-08-16' -u myservice@1.service -g query | wc -l
4804
real 0m4.127s
user 0m3.727s
sys 0m0.186s
Probably more painful on the day to day is how `journalctl -fu myservice` seems to stall for on the order of 5 to 10s, sometimes. I can't reproduce at the moment and maybe it only happens on some machines, but if you were interested in 'real-world anecdata' it's something to maybe note.$ time journalctl > /tmp/all.log
real 1m11.364s user 0m52.299s sys 0m6.540s
$ time wc -l /tmp/all.log 3659597 /tmp/all.log
real 0m0.152s user 0m0.056s sys 0m0.096s
$ time journalctl | grep sshd | wc -l
12944
real 0m53.973s user 0m49.535s sys 0m5.210s
$ time grep sshd /tmp/all.log | wc -l 12944
real 0m0.429s user 0m0.332s sys 0m0.100s
https://github.com/systemd/systemd/issues/2460#issuecomment-...
The reads from /tmp/all.log are almost certainly cached since you just wrote the file, and will basically boil down to a memcpy call, rather than actual disk I/O. Speed difference isn't as big as you would think on a modern SSD, but it isn't nothing either.
Running this between calls should flush the changes to disk and then drop the page cache, making for a fairer test.
$ sudo sync
$ echo 3 | sudo tee /proc/sys/vm/drop_caches
journalctl >/tmp/all.log 0m59.763s
journalctl | grep -c sshd 0m57.767s (otherwise >all is the only example encumbered by write speeds)
wc -l /tmp/all.log 0m0.336s
grep -c sshd /tmp/all.log 0m1.776sThis is what you should be running for a proper comparison:
1. `echo 3 | sudo tee /proc/sys/vm/drop_caches`
2. `time journalctl -u ssh.service >/dev/null`
3. `journalctl >/tmp/all.log`
4. `echo 3 | sudo tee /proc/sys/vm/drop_caches`
5. `time grep -q sshd /tmp/all.log`