So in theory we could uniquely identify all logs as a much more compact binary representation + a lookup table we ship with the executable.
So in theory we could uniquely identify all logs as a much more compact binary representation + a lookup table we ship with the executable.
You can see a fairly general explanation of the concept here: https://messagetemplates.org/
It was hard to figure out how to use it without documentation so it wasn’t very popular. No idea if they still ship it in the DDK.
It was a preprocessor that converted logging macros into string table references so there was no runtime formatting. You decoded the binary logs with another tool after the fact.
Vaguely remember some open source cross platform tool that did something similar but the name escapes me now.
Is be curious how that stacks up against something like zstd with a predefined dictionary.
Originally there was only one system-wide application event log and you needed to be admin to install your message definitions but it all changed in Vista (IIRC). I'd lost interest by then so I don't know how it works now. I do know that the event log viewer is orders of magnitude slower than it was before the refit.
ETW is for high speed general purpose logging of low-level events with multiple collection mechanisms including realtime capture.
Sure, zipping your logs gives you a LOT, and you should do it, but if the result disappoints it is not the end of the road, at all.
But someone who has tried to wrangle gazillion row dumps from a variety of old msgpack protobuf etc and make sense of it all will hate it.
Zipped text formats are infinitely easier to own long term and import into future fancy tools and databases for analysis.
This is a logging library that does lazy formatting, - https://defmt.ferrous-systems.com/
In most modern distributed tracing, "observability", or similar systems the write amplification is typically 100:1 because of these overheads.
For example, in Azure, every log entry includes a bunch of highly repetitive fields in full, such as the resource ID, "Azure" as the source system, the log entry Type, the source system, tenant, etc...
A single "line" is typically over a kilobyte, but often the interesting part is maybe 4 to 20 bytes of actual payload data. Sending this involves HTTP overheads as well such as the headers, authentication, etc...
Most vendors in this space charge by the gigabyte, so as you can imagine they have zero incentive to improve on this.
Even for efficient binary logs such as the Windows performance counters, I noticed that second-to-second they're very highly redundant.
I once experimented with a metric monitor that could collect 10,000-15,000 metrics per server per second and use only about 100MB of storage per host... per year.
The trick was to simply binary-diff the collected metrics with some light "alignment" so that groups of related metrics would be at the same offsets. Almost all numbers become zero, and compress very well.
HTTP/2 would also improve efficiency because of its built-in header compression feature, but again, I've not seen this used much.
The ideal would be to have some sort of "session" cookie associated with a bag of constants, slowly changing values, and the schema for the source tables. Send this once a day or so, and then send only the cookie followed by columnar data compressed with RLE and then zstd. Ideally in a format where the server doesn't have to apply any processing to store the data apart from some light verification and appending onto existing blobs. I.e.: make the whole thing compatible with Parquet, Avro, or something other than just sending uncompressed JSON like a savage.
Weird perspective, yours.
It would cut costs a lot if the source agents did this (pre)processing locally before sending it down the wire.
The answer to "this thing is horrendously inefficient because of misaligned incentives" isn't to be frugal with the thing, but to make it efficient, ideally by aligning incentives.
Open source monitoring software will eventually blow the proprietary products out of the water because when you're running something yourself, the cost per gigabyte is now just your own cost and not a profit centre line item for someone else.
IP addresses are obviously part of the formatting data, not the fixed strings.
IPv6 is used internally a lot more than externally, so I would expect to see a LOT of commonality in the network ID fraction of the addresses- essentially all the bits of representing your IPv6 network ID get reduced to the number of bits in a compression token, in the worst case. In the moderate case, you get a few chatty machines (DNS servers and the like) where the whole address is converted to a single compression token. In the best case, you get that AND a lot of repetition in the rest of the message, and you reduce most of each message to a single compression token.
It's hard to explain if you haven't actually experimented with it, but modern variants of LZ compression are miraculous. It's like compilers- your intuition tells you hand tuned assembly is better, but compilers know crazy tricks and that intuition is almost always wrong. Same with compressors- they don't look at data the same way you do, and they work way better than your intuition thinks they would.