Why observability requires a distributed column store
honeycomb.io
honeycomb.io
Also while PostgreSQL as well as MySQL were initially row-based DBs, they both nowadays support columnt store. PostgreSQL supports it via Citus and MySql (MariaDB) supports it natively.
Retained spans don't include the whole trace with them and you can't make them do that. You can't filter traces by fields in multiple spans. You have to set facets beforehand on fields you want to filter on, otherwise you can't.
So we created a small fork of the dd tracer repo which sends all spans to firehose on the side and stores them as compressed JSON files on S3. There we query them using Athena.
It's cheap, fast, and has 100% retention for an extended period of time. If you get the partitioning right and create a temporary table with just the time period you're interested in before starting a debugging session, then Athena will also be incredibly cheap.
Datadog APM is still nice for exploration and initial analysis, but Athena is great for any required deep dives.
PS: We'll probably write a blog post about the specifics of that setup at some point.
[0]: https://spacelift.io
It really like using Honeycomb vs Grafana + Prometheus for figuring out what is broken. I cannot wait until their dash boarding and alert/slo management mature enough that I don’t have to use separate products for it.
Pre-public-Git repository attempts at Prometheus' data store had used Cassandra in prototyping, but the Thrift protocol, which Cassandra then used, was unfortunately riddled with bugs in the wireformat (how were enums indexed over the wire? turns out it varied by what kind of client you were using even in the same distribution release, which was to me incredulously bad), which resulted in a lot of byzantine bugs. I eventually gave up on Cassandra and settled on a different data model using LevelDB (Levigo bindings). It worked well and was stable and allowed me to not need to worry about:
https://danluu.com/file-consistency/
and
https://danluu.com/deconstruct-files/
I had a feeling in the back of my mind once Prometheus matured a bit it could have fallen back onto Cassandra for distributed series storage (again, ca. 2012 thinking), since it largely could have used an append-only storage model for series archival. Even the metrics metadata (metric families, fields, and indexing) could have been stored in such a way.
It would be fun to revisit these design problems with today's technologies and knowledge.
This is worth a read.