TrailDB – An Efficient Library for Storing and Processing Event Data
tech.adroll.com
tech.adroll.com
What use-case would TrailDB be the obvious, hands-down way to go (vs maybe Postgres + Citus)?
Btw, I am happy to answer any questions about the project here.
We use Kinesis amongst other things to stream raw data to S3.
How do you handle continuous data with TrailDB? It seems that you store raw events Kinesis, buffer events and periodically write TrailDB files to S3. When you want to process events for a specific user, normally the events might be in a random TrailDB file so the timeline would be mixed up. Do you merge TrailDB files in a single instance when processing the data or have a sharding mechanism?
And continuous data is handled by sharding TrailDBs across some fields in our log lines, this way all the related events for a cookie in a given day belong in the same shard, each day the shard mapping is the same and we can just download the same shard ids from S3 and process the files sequentially using our DSL language. With a bit of code you can make this whole process of downloading from S3 and processing completely automated, this is in fact what we do with our data pipeline[0][1].
[0]: http://tech.adroll.com/blog/data/2015/09/22/data-pipelines-d... [1]: http://tech.adroll.com/blog/data/2015/10/15/luigi.html
Thank you so much Ville and AdRoll for sharing this with the world! Your talk on the Trillion Row Data Warehouse in Python was an inspiration for much work I have done since. This library will carry that even further.
Its about efficiency. TrailDB compresses the hell out of event streams grouped by some key (typically user). I think it's generally something like 100x - it does so by making certain assumptions about your data. Similarly, when operating on your event sequences, many of those operations can be performed as integer instructions on the compressed data.
TrailDB gives you order(s) of magnitude of leverage. That's the difference between doing analytics on a single machine and a cluster in many cases. There aren't too many other databases that were designed to gracefully handle tens of trillions of events, so its really not surprising that you need a different tool for that job.
[1] http://akorotkov.github.io/blog/2016/04/06/extensible-access...
PipelineDB is all about querying streaming, real-time data in SQL whereas TrailDB is great for computationally intensive analysis of historical data using any programming language.
TrailDB is more about a different way of grouping and events, and granularly querying and analysing each trail of data. TrailDBs are materialized and stored in S3 typically.
This is a library to read and write a data format that is optimized to give access to granular events and actors within an event stream. For example this could be used to trail all of the events generated by one entity (credit card, cookie, email, account and so on) over a dataset. At that point you can choose what you want to do with it: extract features for ML, train ML directly on raw data, run arbitrary queries for outliers and anomaly detection and what have you.
Also, another event DB project: https://github.com/benbjohnson/skydb.io/blob/master/source/b...
> So I understand, by "pure go" you mean you've thrown out the lua, julia and ruby bits?
wow, that was quite the polyglot codebase...
TrailDB is a C library, Parquet and Avro are not. Depending on your use case this might be a pro or a con.
Most time-series databases manage to store a large number of data points by aggregating data over time. Discrete events can't be aggregated. Instead, TrailDB leverages predictability of events to compress data over time.
> InfluxDB is better geared towards the following use cases: Storing all individual events, not just time series of values. E.g. storing every HTTP request with full metadata vs. storing the cumulative count of HTTP requests for certain dimensions.
[0] https://prometheus.io/docs/introduction/comparison/#promethe...
However, last time I looked that didn't seem like a use case InfluxDB invested a ton of effort optimizing for, especially if you tried to store a lot of separate "event histories" either as separate measurements or by using tags. By "a lot" I mean 100K-100M histories. It may have improved since.
Besides, there is somewhat fundamental tradeoff between allowing efficient realtime granular writes, which I believe a priority for InfluxDB, and building efficient indexed store for millions of histories/event trails, which is more of a TrailDB use case. Kind of OLAP vs OLTP.
Contributions are welcome :)
Let me know if you have any use cases that are in the API but not covered yet by this library :)
It would be interesting to benchmark Judy against another similar well-optimized data structure.