* Citus has a great clustering story, and a small data warehousing story, afaik no timeseries story;
* TimescaleDB has a great timeseries story, and an average data warehousing story;
* Clickhouse has a great data warehousing story, an average timeseries story, and a bit meh clustering story (YMMV).
(Disclaimer: I work for a competitor)
This is a really great comparison. I might borrow it in the future :-)
But yes, if you have classic OLAP-style queries (e.g., queries that need to touch every database row), Clickhouse is likely the better option.
For anything time-series related, and/or if you like/love Postgres, that is where TimescaleDB shines. (But please make sure you turn on compression!)
TimescaleDB also has a good clustering story, which is also improving over time. [0][1]
What competitor btw? I tried to open a link from your profile but it does not work.
The TimeScale licensing approach, the way it is written, perhaps accidentally, has lots of hidden landmines. The TimeScale license slants toward cloud giant defense to the extent that normal use is perilous.
For example, timescale can be used for normal data (postgres) as well, so any rules seem to apply to all your data in the database. The free license only usable available if:
the customer is prohibited, either contractually or technically, from defining, redefining, or modifying the database schema or other structural aspects of database objects, such as through use of the Timescale Data Definition Interfaces, in a Timescale Database utilized by such Value Added Products or Services.
My read is that if you let a customer do anything that adds a custom field, or table, or database, or trigger, or anything that is "structural" (even in the regular relational stuff) anywhere in your database (metrics or not), you are in violation. There doesn't seem to be a distinction about whether this is "direct" control or not, or whether a setting indirectly adds a trigger. I don't want to be in a courtroom debating whether a new metric is a "structural change!"
Now, none of that might be the intent of the license, but you have to go by what it says, not intentions.
The sad part of that is, I, and I'm sure many folks, have no interest in starting a database company, but we can't rally timescale because of legal risk. Looks awesome otherwise, though.
Hi Eric, thanks for taking a close look at our license.
I'd like to dispel some misconceptions:
The core of TimescaleDB is Apache2. Advanced features are under the Timescale License.
Regarding this:
the customer is prohibited, either contractually or technically, from defining, redefining, or modifying the database schema or other structural aspects of database objects, such as through use of the Timescale Data Definition Interfaces, in a Timescale Database utilized by such Value Added Products or Services.
My read is that if you let a customer do anything that adds a custom field, or table, or database, or trigger, or anything that is "structural" (even in the regular relational stuff) anywhere in your database (metrics or not), you are in violation. There doesn't seem to be a distinction about whether this is "direct" control or not, or whether a setting indirectly adds a trigger. I don't want to be in a courtroom debating whether a new metric is a "structural change!"
That's not correct, and we took pains to clarify that in the license: 3.5 "Timescale Data Definition Interfaces" means SQL commands and other interfaces of the Timescale Software that can be used to define or modify the database schema and other structural aspects of database objects in a Timescale Database, including Data Definition Language (DDL) commands such as CREATE, DROP, ALTER, TRUNCATE, COMMENT, and RENAME. [0]
Strictly speaking, if you provide Data Definition Interfaces (DDL) to customers via a SaaS service (ie you are running a TimescaleDBaaS - which applies to < 0.000001% of all possible users) you are in violation of the license. But otherwise you are fine.If you are looking for more votes of confidence, today there are literally millions of active TimescaleDB instances, including by large companies like Walmart, Comcast, IBM, Cisco, Electronic Arts, Bosch, Samsung, and many many smaller ones. [2]
If you have any other questions, I'm happy to answer them here, or offline (ajay at timescale dot com).
[1] https://www.timescale.com/legal/licenses#section-3-5-timesca...
The phrasing "such as" in "such as through use of the Timescale Data Definition Interfaces" looks to me like it can be interpreted as saying "Including but not limited to"
It was a little cumbersome to list every DDL SQL command, which why it uses that language. But that is the intent.
If you have a specific question, happy to answer it here (or offline)
Also, we used this language deliberately to provide more clarity. DDL vs DML is a pretty clear line to most developers who use TimescaleDB (vs some other companies who use language like, "you can't compete with us" etc).
My interpretation doesn't hinge on the DDL/DML question. If what you said is your intent, the legal language used is wrong and you should fix it. Consider this a bug report. Here's a source!
https://www.law.cornell.edu/definitions/uscode.php?width=840...
In order for the license to be usable, you need to be limitative here.
The issue is the 'such as', which reads to me as indicating that providing an API endpoint that can add fields would also be a way of letting the customer modify the database schema and therefore covered.
Which means that building a SaaS app backed by timescale that has any level of customisation exposed to the user appears to be prohibited.
This seems a rather stronger level of prohibition than stopping people directly competing with you, and would suggest that if it's intended it would help to make it more explicit, and if it isn't then an explicit statement of that would be worth adding.
But I appreciate the feedback on how we could make our language clearer. Will share with the team!
I understand from your perspective this whole discussion probably looks silly, because you know what’s in your mind.
There are just too many stories of this blowing up in someone’s face (mostly due to legal, not due to actual enforcement) :/
a) Oracle buys TimescaleDB
b) Oracle sues any of 1000+ SaaS apps that have decent revenue and they can identify
c) The other 900 suddenly have massive due diligence issues even if not sued
If I have a table that records timeseries data and then another table that has a customer-provided extensible set of metadata where a customer can define columns and other related tabular data, would that violate the license? The customer doesn't have a direct, like, psql level of access but the API intentionally provides a very similar level of interaction.
Does this qualify as providing Data Definition Interfaces? If none of those additional columns and such appear on tables set up as timescale tables does that make any difference?
The Timescale License only covers the TimescaleDB code. Postgres code continues to be covered by the OSS PostgreSQL License [0].
So putting aside the question about API vs. psql level (again, the Timescale License was drafted to enabled this for "Value Added Services", e.g., where "such value-added products or services are not primarily database storage or operations products or services"), this license wouldn't apply for non-Timescale code.
[Timescale co-founder here]
[0] https://github.com/timescale/timescaledb/blob/master/NOTICE
One other bug report on the license language front, this language could be construed to prohibit uses like Prometheus--unless there's a definition of operations products I missed (possible).
> are not primarily database storage or operations products
Suggested edit:
> are not primarily database storage or *database operations products*
Yes, database is a modifier to both storage and operations: database (storage or operations) products.
It's a good edit; will keep in mind.
(For good reason, we generally just don't like to "update" the text of the license too much, even for minor nits.)
Thanks!
Understood. If you see the other thread regarding the "such as" language, there is a serious edit that you can batch-up and repair to reflect your intentions with this one.
The "such as" language, retains your right to sue anyone for license violations if their API allows any customer action that causes structure changes indirectly, via the DDL, even under the hood (materialized views, too, presumably). That's way way more use cases than just repackaging TSDB as a service. That's a landmine, which when people compare and choose databases, they'd just assume avoid, even if otherwise comfortable with a cloud-protective license. Making this clearer and less onerous probably will probably pay for itself with a wider top-of-funnel for the product with more people more confident in the license.
The "We Clarified It In a Thread on Hacker News Public License" is probably not as ideal as updating the places that need clarification. :-P
I appreciate from your POV this probably is silly but for us it’s very helpful to have explicit clarity since investors are questioning and demand certainty. I once had to rewrite https handshakes because a lawyer thought export compliance laws would be violated, so hopefully you understand my trepidation :-)
How so? An end user should prefer a database under a license that protect the developer and users from cloudification/proprietization/SaaS
This worries me and makes me wonder if they are going for the open-source-only-by-name model.
[Please reply instead of giving silent dowvotes.]
> As an alternative, you can provide DCO instead of CLA. You can find the text of DCO here: https://developercertificate.org/ It is enough to read and copy it verbatim to your pull request.
> If you don't agree with the CLA and don't want to provide DCO, you still can open a pull request to provide your contributions.
https://github.com/ClickHouse/ClickHouse/blob/master/CONTRIB...
Anyway, Yandex CLA will be removed in the upcoming days (it should be already removed).
I don't know about ClickHouse but the other 2 uses bitmap indexes to make storing petabytes of data affordable.
Row oriented databases would struggle to compete against ClickHouse. They are easily an order of magnitude slower.
For example, there are a couple varieties of Bloom filters, which allow you to test for presence of string sequences in blocks. This allows ClickHouse to skip reading and uncompressing blocks (actually called granules) unnecessarily.
By Splitbee: https://github.com/ClickHouse/ClickHouse/issues/22398#issuec... By GitLab: https://github.com/ClickHouse/ClickHouse/issues/22398#issuec... And others: https://github.com/ClickHouse/ClickHouse/issues/22398#issuec... https://github.com/ClickHouse/ClickHouse/issues/22398#issuec...
If you'll find more, please post it there.
TimescaleDB can work pretty fine in time series scenario but does not shine on analytical queries. For most of time series queries, it is below ClickHouse in terms of performance but for small (point) queries it can be better.
The main advantage of TimescaleDB is that it better integrates with Postgres (for obvious reasons).
There are also many comparisons of ClickHouse vs Citus. The most notable is here: https://blog.cloudflare.com/http-analytics-for-6m-requests-p...
ClickHouse can do batch DELETE operations for data cleanup. https://clickhouse.com/docs/en/sql-reference/statements/alte... It is not for frequent single-record deletions, but it can fulfill the needs for data cleanup, retention, GDPR requirements.
Also you can tune TTL rules in ClickHouse, per table or per columns (say, replace all IP addresses to zero after three months).
@zX41ZdbW@ - Thanks for pointing out the various benchmarks that have been run by other companies between Clickhouse and TimescaleDB using TSBS[1]. As we mentioned, we'll dig deeper into a similar benchmark with much more detail than any of those examples in an upcoming blog post.
One notable omission on all of the benchmarks that we've seen is that none of them enable TimescaleDB compression (which also transforms row-oriented data into a columnar-type format). In our detailed benchmarking, queries on compressed columnar data in Timescale outperformed Clickhouse in most queries, particularly as cardinality increases, often by 5x or more. And with compression of 90% or more, storage is often comparable. (Again, blog post coming soon - we are just making sure our results are accurate before rushing to publish.)
The beauty of TimescaleDB columnar compression model is that it allows the user to decide when their workload can benefit from deep/narrow queries of data that doesn't change often (although it can still be modified just like regular row data), verses shallow/wide queries for things like inserting data and near-time queries.
It's a hybrid model that provides a lot of flexibility for users AND significantly improves the performance of historical queries. So yes, we do agree that columnar storage is a huge performance win for many types of queries.
And of course, with TimescaleDB, one also gets all of the benefits of PostgreSQL and its vibrant ecosystem.
Can't wait to share the details in the coming weeks!
But it can't be updated or deleted, so what do you mean by this?
The overall concept that I was intending to highlight is that you can benefit from both row & columnar store in TimescaleDB. Chunks that are not yet compressed (row store data) can be modified (INSERT/UPDATE/DELETE) as usual and it's transactional - so you're assured it's been completed.
As of TimescaleDB 2.3, compressed chunks (columnar) do allow INSERTS but UPDATES/DELETES on compressed chunks are not yet supported natively. You _can_ decompress any chunk and modify the data (again, transactionally) as needed and recompress.
Which database would be a good fit for this? There isn't too much data, maybe tens of thousands of rows eventually. Would Timescale be a good fit? I'd prefer that, due to existing familiarity with Postgres, but if ClickHouse is better, that's good too.
Initial discussion: https://github.com/ClickHouse/ClickHouse/issues/19627
Being implemented: https://github.com/ClickHouse/ClickHouse/pull/24755
Clickhouse has ALTER ... DELETE and ALTER ... UPDATE functionality now! (and TTLs)
We've recently been working through a detailed benchmark of TimescaleDB and Clickhouse. The DELETE/UPDATE question has been an intriguing story to follow - and I honestly hadn't considered the GDPR angle.
ATM, Clickhouse is still OLAP focused and their MergeTree implementation does not allow direct DELETE (or UPDATE) of any data. All DELETE/UPDATE requests are applied asynchronously by (essentially) re-writing/merging the table data (it's referred to as a "mutation") without whatever data was referenced in the DELETE/UPDATE. [1]
[1]: https://clickhouse.com/docs/en/sql-reference/statements/alte...
Data for in-active users gets deleted because our clickhouse retention policy is lower than the in-active-user timeout
Most products do the asynchronous rewrite, especially if they're based on immutable storage. That's fine, but it should be tested to verify that it's not triggering on every delete, for example, and that it's resource-efficient.
I use them every now and then, but I prefer working with partition strategies when I have to these programmatically.
Altinity is fixing this. The project is called Lightweight Delete and it's for exactly the GDPR reason cited. The idea is that there will be a SQL DELETE command that causes rows to disappear instantly. What actually will happen is that they will be marked as deleted, then garbage collected on the next merge.
Disclaimer: I work for Altinity.