Beringei: A high-performance time series storage engine
code.facebook.com
code.facebook.com
They speak about compressing the data before "storing" it.
I don't have a lot of experience with inmemory anything, but are we talking about retaining the compressed format in server memory here? Ie, RAM is your datastore.
Then, at some point, to serve requests/queries for the data don't you have to get it "out of" RAM and uncompress it, also an inmemory operation?
Did I get this right?
Somebody implemented the algorithm in go based on the paper here: https://github.com/dgryski/go-tsz
A short answer that may work for your question: the bits that are set in RAM are xor values relative to previous values. To provide an answer as to what the value is, a series of read|xor operations are performed.
For series without periodic sampling, delta-of-deltas performs similarly to normal deltas.
I thought they were talking about healthcare data too when I first read it but seeing as this is Facebook I realized they are talking about server health not human health.
Context: initial commit with 277,989 additions [0].
[0] - https://github.com/facebookincubator/beringei/commit/17a6c2d...
Additionally, Dynamodb can be decent for this as long as you only need to sort on time.
Same way most web servers log traffic.
With some basic assumptions on my part including that you can have a delay in writing data to the database (since it's archival and analysis), but you don't want the application to be delayed, putting a fast queue like Kafka or RabbitMQ in place could help. It'll buffer when the database is under load, isolating your application from that.
Offline analysis might also skip the database entirely: columnar data formats (parquet) or log structured merge / sorted string tables (rocksdb) could work very well for archival or offline analysis in MR type environments.
1Gb per day is not a lot - that's 12k/sec which is trivial. Your Postgres must be badly misconfigured!
The "hacked" way is to go CSV files, store them on s3. It is very limited but it is also very simple. (Not sure what software can run queries on CSV files directly).
The "doing things right" way is to go for AWS RedShift or Google BigQuery. There is a learning curve at the start but it's really REALLY good and it will pay off by many folds later.
Have a separate connection to a DB on another machine that is handling the load. Do that within the transaction/request from the client, or pop it into a message queue.
The queue will "persist" the data, but will also schedule it for when your other box/DB can ingest it. At the end of the day, if you're pushing too-much data, you need to have a mechanism in place to let you scale. To me, a simple solution would be a message queue server with plenty of redundant storage space. From there on, you keep adding more consumers.
SQL docs: https://github.com/axibase/atsd-docs/tree/master/api/sql
Analytics examples: https://github.com/axibase/atsd-use-cases
* https://prometheus.io/docs/operating/storage/#chunk-encoding
* https://github.com/prometheus/prometheus/blob/127332c56f85b7...
https://github.com/facebookincubator/beringei#license says "We also provide an additional patent grant."
None of this sounds to me like you have to grant Facebook any patents to use Beringei.
That's exactly what we're using TrailDB for. Works great.
Individual events are typically hundreds to thousands of nanoseconds long, with about 20 nanoseconds of overhead to grab a timestamp using the rdtscp instruction.
Hope that helps!
We also do some hdr_histogram processing for our dashboard in parallel to storing traces in TrailDB, and then we retain the individual traces to do longer term processing or when we're tracking down issues in production.
• we track function entry/exit/throw, offset (in nanoseconds) from the initial timestamp, and in the case of a throw, we also capture a stack trace (which is not stored in TrailDB)
• our server is written in C++
We would love to hear how you are using TrailDB and what could be improved for your use case. Feel free to open issues in GitHub or drop by at our Gitter channel, https://gitter.im/traildb/traildb