A Practitioner's Guide to Wide Events
jeremymorrell.dev
jeremymorrell.dev
This is also complicated by the possibility to apply various filters for the events before and after ststs' calculations.
Wide events can be stored in traditional databases. But this approach has a few drawbacks:
- Every wide event can have different sets of fields. Such fields cannot be mapped to the classical relational table columns, since the full set of potential fields, which can be seen in wide events, isn't known beforehand.
- The number of fields in wide events is usually quite big - from tens to a few hundreds. If we are going to store them in a traditional relational table, this table will end up with hundreds of columns. Such tables aren't processed efficiently by traditional databases.
- Typical queries over wide events usually refer only a few fields out of hundreds of available fields. Traditional databases usually store every row in a table as a contiguous chunk of data with all the values for all the fields of the row (aka row-based storage). Such a scheme is very inefficient when the query needs to process only a few fields out of hundreds of available fields, since the database needs to read all the hundreds fields per each row and then extract the needed few fields.
It is much better to use analytical databases such as ClickHouse for storing and processing of big volumes of wide events. Such databases usually store values per every field in contiguous data chunks (aka column-oriented storage). This allows reading and processing only the needed few fields mentioned in the query, while skipping the rest of hundreds fields. This also allows efficiently compressing field values, which reduces storage space usage and improves performance for queries limited by disk read speed.
Analytical databases don't resolve the first issue mentioned above, since they usually need creating a table with the pre-defined columns before storing wide events into it. This means that you cannot store wide events with arbitrary sets of fields, which can be unknown before creating the table.
I'm working on a specialized open-source database for wide events, which resolves all the issues mentioned above. It doesn't need creating any table schemas before starting ingesting wide events with arbitrary sets of fields (e.g. it is schemaless). It automatically creates the needed columns for all the fields it sees during data ingestion. It uses column-oriented storage, so it provides query performance comparable to analytical databases. The name of this database is VictoriaLogs. Strange name for the database specialized for efficient processing of wide events :) This is because initially it was designed for storing logs - both plaintext and structured. Later it has been appeared that it's architecture ideally fits wide events. Check it out - https://docs.victoriametrics.com/victorialogs/
[1] https://clickhouse.com/blog/a-new-powerful-json-data-type-fo...
[1] https://clickhouse.com/docs/en/sql-reference/data-types/newj...
For this clickhouse wide event lib I'm working on (not worth anyones time atm) I am still using this schema https://www.val.town/v/maxm/wideLib#L34-39 (which is from a Boris Tane talk https://youtu.be/00gW8txIP5g?t=801) for good multi-tenant performance.
I hope clickhouse performance here can still be vastly improved, but I think it is a little awkward to get optimal performance with wide events today.
Just my feeling would be that I’d add the tenant ID before the timestamp as it should filter the parts more effectively
The novelty of wide events is that it is recommended to:
- emit an event (leg entry) once per every processed request. Previously it was OK to emit many logs per request. This could complicate degugging and analyzing such logs.
- don't afraid to add fields to the event if these fields can help debugging and/or analyzing the logs. Previously it was recommended artificially limiting the number of log fields to some small value. This could prevent from debugging and analyzing such logs in the future.
However for understanding how your system or your users are behaving, querying the wide or “main” events will be far better as entry points for exploration.
However the practice of collecting a lot of context per-transaction / unit-of-work and emitting that as one piece of data, storing it in a place that can quickly query across these and visualize is not very common across most orgs and teams. Feedback on this article has been a mix of “I’ve never heard of this before” and “we’ve been doing this for a decade, didn’t know anyone had a name for it” with not a lot of in-between.
It’s not a new idea, which I call out in the intro. It’s not even a very fancy idea. It is really, really helpful if implemented though. Modern OLAP column stores help a lot here too since they make this type of exploration cheap and quick.
"Observability" seems like a weird term for that to me, but okay.
But I don't understand why not just give the appropriate context in the submission, rather than keeping a title that only makes sense to a very specific niche audience and then not saying up front what the niche is.
The concept of an "event" is coherent in many other programming contexts, so the possibility that one could be coherently "wide" is at least plausibly interesting. But then I get there and find myself completely disoriented, and eventually figure out that it's not actually relevant to anything I do. And anyway it looks like a lot of this jargon is really just not necessary to convey the core ideas... ?
Sure, but it would be nice if title submissions made it feasible to predict the topic category of the article for people who are not already in the relevant niche.
> Adopting Wide Event-style instrumentation has been one of the highest-leverage changes I’ve made in my engineering career. The feedback loop on all my changes tightened and debugging systems became so much easier.
What I get is: here's a thing that made a big improvement to how I debug systems.
Except, it turns out that the systems in question are very specific ones.
> The tl;dr is that for each unit-of-work in your system (usually, but not always an HTTP request / response) you emit one “event” with all of the information you can collect about that work.
Okay, but... as opposed to what? And why is it better this way?
>“Event” is an over-loaded term in telemetry so replace that with “log line” or “span” if you like. They are all effectively the same thing.
In the programming I do, "event" doesn't mean anything to do with logging or telemetry.
I had to lookup wide events in the middle of the article, and I can’t say I can viscerally see and feel the benefits the OP was espousing. Just felt like an adderall-fueled dump of information being thrown at me.
(I wrote it mostly so I could stop re-explaining this concept from first-principles and how to go about implementing it over-and-over again )
That's how I would have titled it.
I get that HN isn't appealing to the general population, but the world of programmers etc. is still quite broad.
You are obviously the one who is not understanding or is perhaps misunderstanding something.
Observability is a pretty standard term in software development.
Events have nothing per se to do with logging or tracing, but you can visualize/trace events with logs/spans.
From my perspective, you seem to misunderstand a lot in the article, I am not judging you for that, just observing this.
I suggest you try to understand the gist of the article instead of scolding the language used.
You’re being unreasonable about this IMO.
Again: not judging, just observing.
Consider that you are perhaps the minority ¯\_(ツ)_/¯