1.2.3.4 - - [08/Aug/2023:12:48:11 +0200] "GET /wp-config.php.bak HTTP/1.1" 404 196 "-" "Mozilla/5.0 (Windows NT 6.1; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/47.0.2503.0 Safari/537.36"
1. Apache has a perfectly coherent and structured view of the situation: Host 1.2.3.4 made a GET request of this particular URL, with this particular UA, at this particular date.2. Apache proceeds to shove all this into a string, destroying the structure, and probably some of the data (it likely knows the timestamp with >1s precision)
3. Then to do some sort of analysis we need to manually reverse this process: write some convoluted regex to match this kind of line, turn an IP address back into a number, turn a timestamp into a number, find the delimiters. We're now undoing the damage Apache did. Why!?
4. For extra fun this adds extra problems to deal with log rotation, log compression, newlines and special characters in fields, and so on. All kinds of problems that are ultimately unproductive to solve and get in the way of actual useful work.
Logging should be structured by nature. I shouldn't be writing weird regexps. I should have every message structured from the start, where I can search by "ip_address is in 1.2.3/24", or "request == 'GET'", or such. Every bit of data should be logged in its pristine form with zero ambiguity, with its original type, and analyzable accordingly.
I think systemd has the beginning of a good idea in here, in that you can actually do this if you care to, and can send arbitrary chunks of data (even binary) to the log if needed.
People who complain about journald not logging in plaintext honestly baffle me, because come on, whose idea of fun is it to do log parsing? And why are we spending CPU cycles on converting timestamps to text and parsing them back?