ClickHouse Cloud is now in Public Beta
clickhouse.com
clickhouse.com
Good to see transparent comparisons available now for Cloud performance vs. self-hosted or bare metal results as well as results from our peers. The ClickHouse team will continue to optimize further - as scale and performance is a relentless pursuit here at ClickHouse, and something we expect to be performed transparently and in a reproducible manner. Public benchmarking benefits all of us in the tech industry as we learn from each other in sharing the best techniques for attaining high performance within a cloud architecture
Full disclosure: I do work for ClickHouse, although have also been a past member of SPEC in developing and advocating for public, standardized benchmarks
Check out the response below that has a reference to some of our billing FAQs.
Is one block on one non-partitioned non-distributed table one write unit? What about one insert that's two blocks on such a table? What about one block on a null engine with two MVs listening to insert into two non-partitioned non-distributed tables? What if the table is a replacing mergetree, do I incur WUs for compactions? etc.
My worry is that it is essentially 1 WU = 1 new part file, which I understand makes sense to bill on but is tremendously intransparent for users - at least I have no clue how often we roll new part files, instead I'm focused on total network and disk i/o performance on one side and client query latency on the other.
For example, I just checked that uploading 1.1GB example table(cell_towers with 14 columns) cost me 0.38 write units.
Nonetheless you can't insert a-whole-file-and-just-that-file in less than one write.
This doesn't imply to me that each individual INSERT costs 1 WU, but that it could be fractional. I guess it depends on how you read it?
(See https://news.ycombinator.com/item?id=33081099 for the original wording.)
There's no way to think about what an actual write unit means. You could measure the costs on a sample workload, but that's far from ideal. Some transparency here would be nice.
I understand the answer is complicated, based on hairy implementation details, and subject to change. Give me the complexity and let me interpret it according to my needs.
Working on updating the FAQ and tooltips now and sharing your feedback. <3
If you are doing INSERT in batches with one million rows, it will give
SELECT formatReadableQuantity(1000000 * 100 / 0.0125)
8.00 billion
inserted rows per dollar. Pretty good, IMO.If you are doing millions of INSERT queries with one record, without "async_insert" setting, it will cost much more.
That's why we have "write units" instead of just counting inserts.
As I've said several times in this thread, I understand why you don't count inserts or rows. What I don't understand is what unit a WU does actually correspond to. In particular I don't understand its relation to e.g. parts or blocks, which are the units one would focus on optimizing self-hosted offerings.
For those complex pipelines you may find more useful to run tests during trial. Data distribution, partitioning and so on can change actual cost significantly so estimates can be too pessimistic or optimistic
Right, that's exactly what I don't want to deal with. Unless I have even just a ballpark estimate of complex pipelines both before I commit to any sales crap and afterwards when we're designing new pipelines, it's just not an option for us at all. I have no clue if it's going to cost us $10, $100, or $10000.
There are other responses from ClickHouse in the comments on the pricing, so I'll defer to their expertise on that topic there. Thank you for your feedback and ideas, as normalizing a price-based benchmark is an interesting concept (and where ClickHouse would expect to lead also given the architecture and efficiency)
I was curious to compare MonetDB, DuckDB, ClickHouse-Local, Elasticsearch, DataFusion, QuestDB, Timescale, and Athena. Amazingly, MonetDB shows up better than DuckDB in all metrics (except storage size), and Athena holds its own and fares admirably well, esp given that it is stateless. While, Timescale and Quest did not come up as good as I hoped they would.
https://benchmark.clickhouse.com/#eyJzeXN0ZW0iOnsiQXRoZW5hIC...
It'd be interesting to see how rockset, starburst (presto/trino), and tiledb fare, if and when they get added to the benchmark.
As it is right now the benchmark is not particularly representative of DuckDB's performance. Check back in a few months :)
[1] https://github.com/duckdb/duckdb/issues/3969#issuecomment-11...
Goes without saying, if there are cost advantages to be had due to DuckDB's unique strengths, then serverless DuckDB Cloud couldn't come here soon enough.
> despite DuckDB freeing more buffers than it is allocating
Can you please clarify how is that even possible?
[1] https://cpp4arduino.com/2018/11/06/what-is-heap-fragmentatio...
Technically, I hope you understand that this isn't possible but maybe I am misinterpreting what you're trying to say.
auto buff = malloc(N);
free(buff);
free(buff);
is one way to free "more" buffers than allocated but this will lead to an UB and depending on the underlying system allocator implementation it may or may not crash.However, given how silly this would be I believe this is not what you're trying to convey?
So, I assume, the context is, DuckDB allocates x buffers, frees x - m buffers at some point later, then allocates n buffers where n <<<< m, and yet malloc fails.
In the GitHub thread mytherin linked to above, Alexey Milovidov, ClickHouse CTO, points out that ClickHouse uses jemalloc and makes for a better choice than glibc malloc given the issue with fragmentation. It is likely that DuckDB switches to jemalloc, too.
The scenario I am describing is roughly the following:
Suppose we allocate 100K buffers that all have an equal size, and our memory usage is now 10GB. After that point we free 20K buffers, but allocate 10K more. In other words, from that point on we are freeing more buffers than we are allocating.
Now, since we are freeing more than we are allocating, you would expect our memory usage to go down. However, when using the standard glibc malloc on Linux, our memory usage unexpectedly goes up. After this happens several times in a row the system runs out of memory and new calls to malloc fails.
"The dataset is represented by one flat table. This is not representative of classical data warehouses, which use a normalized star or snowflake data model. The systems for classical data warehouses may get an unfair disadvantage on this benchmark."
Taking a look at the queries [0], it looks like it mostly consists of full table scans with filters, aggregations, and sorts. Since it's a single table, there are no joins.
[0]: https://github.com/ClickHouse/ClickBench/blob/main/snowflake...
For example, the c6a.4xlarge instance type in AWS has 16 vCPUs, 8 cores and "max_threads" in ClickHouse will be 8.
For full transparency, I think you should do the same in ClickHouse. Or is there a strong reason not to run benchmarks on standard analytical workloads like TPC-H, TPC-DS or SSB?
There are numerous benchmarks that use similar to TPC queries, but those are not standardized and can be misleading. For example a lot of work was done by Fivetran to get this report [0], but they show only overall geomean for those systems and you can't understand how they actually differ. Anyway their queries are not original TPC - variables are fixed in queries, they run first query when official query is a multiquery.
Contributors from Altinity run SSB with flattened and original schemas [2]. SSB is not well standardized and we see a lot of pairwise comparisons with controversial results - generally you can't just reproduce them and get all the results in single place for the same hardware.
[0] https://www.fivetran.com/blog/warehouse-benchmark [1] https://www.tpc.org/tpcds/results/tpcds_results5.asp?orderby... [2] https://altinity.com/blog/clickhouse-nails-cost-efficiency-c...
>c. Public Disclosure: You may not publicly disclose any performance results produced while using the Software except in the following circumstances: (1) as part of a TPC Benchmark Result. For purposes of this Agreement, a "TPC Benchmark Result" is a performance test submitted to the TPC, documented by a Full Disclosure Report and Executive Summary, claiming to meet the requirements of an official TPC Benchmark Standard. You agree that TPC Benchmark Results may only be published in accordance with the TPC Policies. viewable at http: //www.tpc.org (2) as part of an academic or research effort that does not imply or state a marketing position (3) any other use of the Software, provided that any performance results must be clearly identified as not being comparable to TPC Benchmark Results unless specifically authorized by TPC.
But given that each database system has its own flavor of SQL, vanilla TPC benchmarks may not work out of the box so one needs to tweak them a bit and this might be what actually disqualifies the published results from all of the clauses from above being applicable.
I can also anticipate that combination of clause (2) and (3) is what some that publish the results are also taking advantage of.
[1] https://www.oracle.com/mysql/heatwave/performance/ [2] https://www.singlestore.com/blog/tpc-benchmarking-results/ [3] https://docs.pingcap.com/tidb/v6.2/v5.4-performance-benchmar... [4] https://www.monetdb.org/blogs/learning-from-benchmarking/
Had a pricing question. Say we connected a log forwarder like Vector to send data to Clickhouse Cloud, once per second. If each write unit is $0.0125, and we execute 86,400 writes over the course of the day, would we end up spending $1080? Do you only really support large batch, low frequency writes?
The cloud offering seems like an amazing product for companies that they could afford I am not sure if the billing is right, but for 5M inserts per month the total bill would be $62K.
Based on the comments here, and my own confusion, I think figuring out a different way of billing read/write operations is in order.
Batch upload is of course more cost effective, but I would expect that to be more typical in a traditional data warehouse where we are uploading hourly or daily extracts. Clickhouse would likely show up in more real time and streaming architectures where events arrive sporadically.
I am a huge fan and advocate of Clickhouse, but the concept of a write unit is strange and the likely charges if you have to insert small batches would make this unviable to use. A crypto analytics platform I built on top of Clickhouse would cost in the $millions per month vs $X00 or low $X000 with competitors or self hosting.
1) I had no idea what Clickhouse was for the first 30 seconds looking at the homepage. I now understand it to be a database of some sort. I shouldn't have seen the words "performance" and "cloud" and "serverless" before seeing the word database, right? I'm starting off confused. There shouldn't be an assumption that I know what you all do.
2) I have no idea what a column oriented database is. I've been a developer for 29 years (mostly frontend but I do a lot of full stack too). If I need an explainer, a lot of devs will.
Aside from that, it looks like a nice offering and I wish you all the best!
The pitch on the landing page is that ClickHouse Cloud is great if you love ClickHouse. If you don't know what ClickHouse is, you have to do some work to find out.
Whereas traditional db’s, data like [first name, last name], the columns may have way less meaning on their own and you need both columns to have the data make sense.
A traditional db with B(-)tree storage is much slower when dealing with that type of usage. Storing (stock close) in a single column format makes that type of query much faster.
And thanks for the honest feedback.
It's always an interesting balance of promoting a new thing (Cloud) and explaining an existing thing. This might be helpful.
https://clickhouse.com/docs/en/home
(note: I work at ClickHouse)
We all have our specialties, and that is fine. It is a common pattern that a developer gets comfortable with a particular tech stack, and then uses it for many years without seeing the need for much else. Some developers use Ruby on Rails plus Postgres for everything, others use C# and .NET and SQL Server. It's fine, if that's all you need.
Still, this is the year 2022. Cassandra, to take one example, was released in 2008. For everyone who has needed these fast-read databases, they've been much discussed for 14 years, including here on Hacker News, and on every other tech forum. At this point I think a company can simply assume that most developers will have some idea what a column database is.
Does anyone know how much of the Clickhouse team (or ownership) is still located in Russia?
They can show support by donating to UA defence and showing proof (important - not some neutral org). Otherwise it is not support but a bunch of bs
So, yes, words, but the potential consequences of these words have more significance than empty air.
Show proof
From their "We Stand With Ukraine" page. [1]
> The formation of our company was made possible by an agreement under which Yandex NV, a Dutch company with operations in Russia
While Yandex NV is registered in the Netherlands, it's pretty clear that Yandex NV is directly related to Yandex (specifically, according to Wikipedia, it's a holding company of Yandex). For those who don't know, Yandex is basically the Google of Russia and holds 42% of market share among search engines in Russia.
The blog post does not seem to make any claims that Russia is not benefitting financially from the commercial success of Clickhouse. And given the above, such claim would unlikely be true. As such, I still think it's pretty much safe to assume that a portion of any $$$ paid to Clickhouse ultimately goes to fund the war and kill people.
That said, I sincerely hope they could find a way to stop that flow of money from happening somehow, as otherwise it's a nice technology and a great technical team behind it...
Zero.
No employees in Russia, no ownership from Russia, no influence from Russia.
Probably the best comparison is CockroachDB Cloud. They have a "serverless" offering based on unit pricing and a dedicated offering based on provisioned servers + support/maintenance overhead. I think that would be the ideal place to go long-term, but I'm super excited for this current one. I love ClickHouse and want to support them.
ClickHouse is also an interesting case because there's lots of options to migrate clusters, use S3 as long-term storage, etc to where I don't particularly feel locked into this offering if I ever wanted to shift into my own.
Disclaimer: I work for ClickHouse.
Disclaimer: I’m a customer of altinity.
Thanks for this comment. We'll be publishing a blog at Altinity to compare both models. My view is that they both have their place. The BigQuery pricing model is great for testing or systems that don't have constant load. The Altinity model is good for SaaS systems that run constantly and need the ability to bound costs to ensure margins.
Having a selection of vendors that offer different economics seems better for users than everyone competing for margin on the same operational model.
Disclaimer: I'm CEO of Altinity.
Take some notable examples from the list: https://clickhouse.com/docs/en/about-us/adopters/, something around web analytics, APM, ad networks, telecom data... ClickHouse is perfectly suited for these use cases. But if you try to align these scenarios with, say, BigQuery, they will become almost impossible or prohibitively expensive or just slow.
There are specialized systems for real-time analytics like Druid and Pinot, but ClickHouse does it better: https://benchmark.clickhouse.com/
There are specialized systems for time-series workloads like InfluxDB and TimescaleDB, but ClickHouse does it better: https://gitlab.com/gitlab-org/incubation-engineering/apm/apm... https://arxiv.org/pdf/2204.09795.pdf http://cds.cern.ch/record/2667383/
There are specialized systems for logs and APM, but ClickHouse does it better: https://blog.cloudflare.com/log-analytics-using-clickhouse/
There are specialized systems for ad-hoc analytics, but ClickHouse does it better as well: https://github.com/crottyan/mgbench
Well, even if you want to process a text file, ClickHouse will do it better than any other tool: https://github.com/dcmoura/spyql/blob/master/notebooks/json_...
And ClickHouse looks like a normal relational database - there is no need for multiple components for different tiers (like in Druid), no need for manual partitioning into "daily", "hourly" tables (like you do in Spark and Bigquery), no need for lambda architecture... It's refreshing how something can be both simple and fast.
Even join-heavy queries work great.
I've found the documentation to be pretty good and there's a really active telegram group where devs offer help (if you get noticed, but that's not too hard).
As far as whether it works on a single core, you probably need to load data and test. That's the only way to answer that question for sure.
IMO the real question for any database is whether you can invest the time to learn how to fix problems and deal with boring but necessary operations like upgrade. If not, it's probably better to use a managed service. We operate a SaaS for ClickHouse but use managed MySQL and Kubernetes. That choice freed up resources to focus on making ClickHouse run well.
Disclaimer: I work for Altinity.
Altinity has been in the space for a while offering Clickhouse services as well: https://altinity.com/
EDIT: brazenly was the wrong word here :)
alexey-milovidov 17,522 commits 5,648,537 ++ 5,580,633 --
alesapin 4,618 commits 1,024,262 ++ 932,501 --
KochetovNicolai 3,867 commits 377,420 ++ 314,035 --
That said I would like to defend ClickHouse and their use of the logo. They acquired the IP and it's their right to use it as they please. It's hard to imagine them not using the ClickHouse name. There's a crowd of companies using/supporting ClickHouse so it's the obvious way to market themselves.
The interesting question is whether it's a good thing for users to have a large company controlling ClickHouse rather than having it in a foundation.
My advice would be to pivot away from the Clickhouse brand and towards a great OLAP service that happens to run on Clickhouse.
I wish you guys luck but you’ll probably be used by customers mainly to negotiate with Clickhouse on better rates. And they’ll undercut competition when needed since they have so much cash and investor mindshare.
At the same time we contribute to open source and work to make ClickHouse better, just as many other businesses built on ClickHouse do. We've done quite a bit of work including many ClickHouse PRs, maintenance of the Grafana ClickHouse plugin since 2019, Altinity Kubernetes operator for ClickHouse, etc. We plan to step up our contributions as the market grows.
FWIW, here is Alexey's perspective on the topic directly for those of you following along at home.
It is included in many adblocker lists.
https://github.com/StevenBlack/hosts/issues/1781
Unfortunately mvps.org no longer seems to reply to emails.
It will be interesting to watch this unfold from a code licensing perspective. Will Clickhouse Inc. move to a more restrictive license to block all these other ClickHouse services?
I don't see it as a reasonable move.
If they're really doing elastic reads as suggested in another comment I don't think this can be the standard CH server (at least for reads) - or I've missed something very exciting in recent versions.
* https://clickhouse.com/docs/en/guides/sre/configuring-s3-for...
* https://clickhouse.com/docs/en/integrations/s3/s3-merge-tree...
I can vaguely imagine how you'd build such a thing, you'd spin up new instances that knows your write-primary's sharding/partitioning scheme - heck you could even approximate it with slightly-augmented clickhouse-clients doing the S3 pulls directly and no real additional "server" - but as far as I know there's no way out of the box.
https://clickhouse.com/docs/en/integrations/migration/
Depending on what you are trying to achieve, you could use clickhouse-local with the MySQL engine to move data, or could use an ETL/ETL tool to migrate/sync
https://github.com/ClickHouse/ClickHouse/issues/3311
I also had some pretty bad join performance (CH table joined to MySQL table), the quick solution to both of these is that we instead use the table function (https://clickhouse.com/docs/en/sql-reference/table-functions...) to copy the data periodically.
https://altinity.com/blog/connecting-clickhouse-to-external-...
Disclaime: I work for Altinity.
1. Data rebalancing is not automatic.
2. It doesn't really have a concept of cold-tier pushed to S3 directly, so cluster management is not simple for a small team.
Other than that, ClickHouse looks super amazing.
Another approach is to store all in S3 with local caching, it is much more easy and somewhat more efficient.
ClickHouse Cloud covers these concerns.
https://clickhouse.com/docs/en/integrations/s3/s3-merge-tree...
Have anyone migrated from InfluxDB to ClickHouse?
At a previous company, I wrote a simple TCP server to receive LineProtocol, parse it and write to ClickHouse. I was absolutely blown away by how fast I could chart data in Grafana [1]. The compression was stellar, as well...I was able to store and chart years of history data. We basically just stopped sending data to Influx and migrated everything over to the ClickHouse backend.
[1] https://grafana.com/grafana/plugins/grafana-clickhouse-datas...
A single node instance with a fast disk is more than sufficient for most needs: https://hub.docker.com/r/clickhouse/clickhouse-server
If you need a cluster, https://github.com/Altinity/clickhouse-operator makes things easy
IMHO, the unique value-prop of this offering is that it elastically scales reads (compute), like Redshift/Snowflake/BigQuery/other cloud data warehouses do, while also "being Clickhouse", and so giving you very traditional SQL querying capabilities (where those others all ask you to bend your SQL to fit the DB, sometimes in pretty ridiculous ways.)
I would suggest not thinking of this offering as "ClickHouse, but in the cloud"; but rather, thinking of this as "a cloud data warehouse, but rather than using its own proprietary query engine, it's just ClickHouse."
If you haven't evaluated other cloud data warehouses as alternatives to solving your problem (and found them wanting for one reason or another), then you're likely not in the niche that would see a positive profit margin on using ClickHouse Cloud.
If you have a lot of log data and want something open source and serverless you can self host, check out Matano (https://github.com/matanolabs/matano).
The only acknowledged downside is that it does not have automatic rebalancing of data if new nodes are added.
OLAP/Column dbs generally work well for large bulk inserts with many analytical queries…
How do clickhouse (or other column dbs) work with larger inserts (e.g. log data type inserts)?
Clickhouse inc is subsidiary of Clickhouse B.V in Netherlands and is controlled by a bunch of russians.
Can we leave the russiaphobia in the political threads at least?
Not super classy if you ask me.
For those unaware, Yandex is a Russian internet megacorp, think Google but in cahoots with the authoritarian government: cherrypicked news coverage friendly to the government, effectively a monopoly across multiple verticals (eg they bought out Uber in Russia), etc. In 2020, the year that deck seems to be from, they were already censoring their news feed [2] and tweaking search ranking to promote pro-government results [3] for years.
[1]: https://presentations.clickhouse.com/meetup40/introduction/
[2]: https://meduza.io/feature/2022/05/05/my-zamuchilis-borotsya in Russian, but google translate does a reasonable job
[3]: https://www.svoboda.org/a/30580605.html same
Sure, but the split appears have been very successful in this regard.
A cynical view would be that it's only now that the association became bad for business.
[0]: By the way, Yandex was showing Crimea as unqualifiedly Russian territory since 2014, until the war, when they just removed borders between countries completely
Edit: qualified the remark
Google and Apple used to show it as part of Russia to Russian users as well.
https://www.google.com/amp/s/techcrunch.com/2019/11/27/apple...