7 karma · joined December 26, 2021
And it gets even harder when the clock never stops and live data keeps flowing.
Client visualization layer: Timeplus Vistral (https://github.com/timeplus-io/vistral)
Server data processing layer: Timeplus Proton (https://github.com/timeplus-io/proton)
Streaming-native from processing to insights in motion
Streaming joins require maintaining state for both sides of the join
High-cardinality data (millions of unique keys) means huge state sizes
Traditional approach: Keep everything in memory will make memory exhausted
The high-cardinality join memory problem isn't unique to Timeplus. Apache Flink also uses hybrid hash joins that spill to disk (RocksDB) when memory fills, Materialize shares indexed state across multiple queries (but still requires keeping full datasets in memory), and RisingWave stores state in cloud object storage (S3/GCS) with LRU caching for hot data. What makes Timeplus different is its purpose-built optimization for the Pareto Principle, where a tiny fraction of data generates the vast majority of activity - keeping hot data in memory and cold data on disk for dramatic memory savings.
1. Table A : fact events, high-throughput (10k~1M eps), high-cardinality
2. Table B, C, D : couple of dimension tables (fast or slow changing).
The use case is straightforward : join/enrich/lookup everything into one big flattened, analytics-friendly table into ClickHouse.
What’s the best pipeline approach to achieve this in real-time and efficiently?
[1] Dynamic Tables: One of Snowflake’s Fastest-Adopted Features: https://www.snowflake.com/en/blog/reimagine-batch-streaming-...
Domain.spawn (fun _ -> print_endline "I ran in parallel")
Anyway, love the simplicity of this expression!
-> Streaming Queries - Process large datasets with constant memory usage
-> Async Inserts - High-throughput data ingestion with automatic batching
-> Compression - LZ4 and ZSTD support for reduced network overhead
-> TLS Security - Secure connections with certificate validation
-> Connection Pooling - Efficient resource management for high-concurrency applications
-> Rich Data Types - Full support for #ClickHouse types including Arrays, Maps, Enums, DateTime64
-> Idiomatic OCaml - Functional API leveraging OCaml's strengths
That's why we love ClickHouse, the fastest and most lightweight approach for real-time analytics. Furthermore, data stream processing should also uphold these same principles, without unnecessary complexity.
With Timeplus Proton - a single-binary, fast and efficient streaming processing engine, ClickHouse users can now natively and effortlessly utilize SQL queries and Materialized Views to enable fast and scalable incremental streaming processing from Kafka or any other data stream source, turbocharging broader real-time streaming analytics use cases such as data stream pipelines, unified online/offline ML features, infrastructure monitoring, or any other latency-sensitive applications.
Please check this PR from the proton team: https://github.com/ClickHouse/ClickHouse/pull/54870
If you are looking for a high-performance and lightweight streaming processing engine capable of running everywhere, or gaining streaming insights with historical context, you should try Proton: https://github.com/timeplus-io/proton!