HouseWatch: Open-source tool for monitoring and managing ClickHouse clusters
github.com
github.com
OTOH, monitoring long running background jobs on CH cluster is very valuable to have. Its real pain to verify, if parent and child queries have executed correctly. I would suggest doubling down on features that users cannot readily get via grafana or Otel.
that kind of thinking (that it's too hard to learn a second tool) is how datadog gets away with charging $$$$ for mediocre versions of 10 different products that cost an order of magnitude more than they would individually. the benefits you get from combining everything into one tool are vastly overstated compared to the benefits you get from having the in-house expertise to use the right tool for the job.
What type of data benefits from that type of Database?
https://blog.cloudflare.com/http-analytics-for-6m-requests-p...
Ps, great and still highly relevant resource covering all the major database system designs, their advantages and drawbacks: https://www.oreilly.com/library/view/designing-data-intensiv...
Large aggregations, massive datasets, large joins, and workloads that are ready heavy and eschew row-level mutations.
They get used for data analysis frequently, time series data and associated analysis meshes quite nicely too. ClickHouse itself was originally built to support arbitrary analytical queries on clickstream data at pretty massive scale. Cloudflare uses it for live analytics, Uber uses it for logs.
It's originally designed to handle hundreds of columns, and billions of rows, but I think it can still apply to much smaller use cases that value performance. I'm implementing it currently in a similar scenario, and I'm using AirByte OSS version to ELT from postgres. Then I'm using tableau or some other BI tool to analyze that data much more effectively (I will be trying to perform complex aggregations/group by reports on 100mm rows)
If you are _actually_ interested I suggest using google search to find some good sites that go over what a column oriented database does/is used for.
This isn’t hard; I’ll get you started:
https://www.kdnuggets.com/2021/02/understanding-nosql-databa...
Columnar stores are optimized for reads. Row stores are optimized for writes.
Technically the overall idea is that if you have lots of queries that only read certain columns and your database stores rows contiguously it's a waste to read a whole row and then discard columns.
Also compression (such as run length or delta or even ztsd) often works better if you give it a block of data that's from one column (such as a timestamp or tag value).
Everyone doesn’t need to cater to the lowest common denominator of knowledge.