Audit logs are a distinct feature.
Excess load can't take down a well-engineered log collection infrastructure. There can be overload, but the backpressure should propagate downstream and senders and intermediaries should buffer locally if needed. Once the collectors are able to catch up again, the spooled messages will be dispatched, and the backlog should recover.
A well-engineered logging system for sites that care about integrity and durability should look a lot like a distributed message queue.
> Audit logs are a distinct feature.
In my experience, this is not always as distinct as one might hope. On multiple occasions in my career, a customer demanded we perform research using our logs to answer, and the information they sought were not in the class of logs that were considered "audit logs" in advance. Everyone chooses differently what qualifies as "audit logs"; it doesn't have an objective definition.
You're only delaying the inevitable. Even with local buffering, you can arrive at a point where you can buffer no more and have either to choke the production workload or start dropping messages.
TCP alone won't get you there. It's certainly not durable in and of itself. All TCP can do is ensure that streamed data is received in the correct order, and confirm that a segment's data was successfully placed into the right buffer on the receiver side. You also need immutable storage, stronger integrity checks than what TCP itself provides, and many other requirements I haven't researched in a while.
If you truly believe OTel is a “baby fart,” I’d recommend you go to SRECon, KubeCon, and other similar conferences and make your opinion loudly known. Nobody talks about syslog there, and these are big and well-respected companies represented there.
Syslog is mature and battle tested for what it is: schlepping unstructured plaintext logs from one place to another. It’s not really purpose-built to meet the needs of a full-fledged observability solution, though, of which logs are but a component.
In fact, the queue management and at least the possibility of some rudimentary end-to-end cryptographic integrity checks are some of the stronger points of rsyslog. Splunk Cloud and Elastic, as far as I know, lacks the latter completely which rules them out as a single log sink for environments with that type of requirements.
A modicum of research reveals that even the rsyslog documentation starts out with UDP for remote delivery: https://docs.rsyslog.com/doc/getting_started/beginner_tutori...
Well, maybe go observe how a broad array of sites implement it in practice, then you might take it more seriously. Maybe you don't implement it that way, but a lot of people will just follow the tutorials or shortcut their way to something that works (but is brittle).
At any rate, I was responding directly to the claim that "No one has suggested running syslog over unreliable transport" which is obviously untrue.
That documentation link probably isn't as telling as is suggested, because the next example is for tcp. In the old days before tcp support was widespread (looking at you, Java) it was common to listen for udp on localhost so it was probably a common configuration.
There's not much to debate here. Syslog is used everywhere and the main reasoosn are that it is very reliable, trivial to load balance, and popular implementations have integrity checking that is permissible in regulatory environments. You can criticize it for many things, for example that most parsers are much too liberal or that the facility and severity fields are clearly dated, but not for being unreliable.
This was a solved problem a long time ago in rsyslog. One can define a local spool and enable TCP (and optionally encryption) to multiple syslog servers. If something interrupts the flow the syslog messages will queue locally and then de-spool when communications are restored.
Edit: being pedantic -- it's syslog-ng actually.