The reason is, it's thousands of times cheaper to +1 a counter/record a gauge and flush it every once in a while, than to serialize per-request metrics into HTTP headers, log them, and re-parse them for analytics later.
The reason is, it's thousands of times cheaper to +1 a counter/record a gauge and flush it every once in a while, than to serialize per-request metrics into HTTP headers, log them, and re-parse them for analytics later.
If all you care about is overall latency, awesome! Use a TSDB. Once you care about latency per endpoint/user agent/customer ID/client platform (or combination thereof), you need the flexibility associated with structured log data, stored in something meant for fast analytical querying.
0: https://www.honeycomb.io/blog/the-problem-with-pre-aggregate...
Use a basic sampling approach (recommended by Honeycomb and others) to sample e.g. one percent or less of your “boring” transactions, 10% of your suspect transactions, and all of your errors. By suspect I mean too slow by some cutoff, generated more work than expected, etc. Those numbers are arbitrary and could change depending on the environment. For CDN logs of static assets, 0.1% or less might be sufficient because there is hopefully not much going on there, but you’d like to know if there were.
You then (importantly) store the sampling rate along with the event itself to reconstruct a usable comparison between the different types of events. For example, one event for each of the previous sampling rates I gave would be three total events stored but represent 111 transactions.