ClickHouse, Inc.
github.com
github.com
As we did not want to go into the HA/backup/restore details at that time we created a solution that can be quickly recreated from data in other databases.
Interesting presentation from Alexey about Features and Roadmap from May 2021:
* Citus has a great clustering story, and a small data warehousing story, afaik no timeseries story;
* TimescaleDB has a great timeseries story, and an average data warehousing story;
* Clickhouse has a great data warehousing story, an average timeseries story, and a bit meh clustering story (YMMV).
(Disclaimer: I work for a competitor)
This is a really great comparison. I might borrow it in the future :-)
But yes, if you have classic OLAP-style queries (e.g., queries that need to touch every database row), Clickhouse is likely the better option.
For anything time-series related, and/or if you like/love Postgres, that is where TimescaleDB shines. (But please make sure you turn on compression!)
TimescaleDB also has a good clustering story, which is also improving over time. [0][1]
What competitor btw? I tried to open a link from your profile but it does not work.
The TimeScale licensing approach, the way it is written, perhaps accidentally, has lots of hidden landmines. The TimeScale license slants toward cloud giant defense to the extent that normal use is perilous.
For example, timescale can be used for normal data (postgres) as well, so any rules seem to apply to all your data in the database. The free license only usable available if:
the customer is prohibited, either contractually or technically, from defining, redefining, or modifying the database schema or other structural aspects of database objects, such as through use of the Timescale Data Definition Interfaces, in a Timescale Database utilized by such Value Added Products or Services.
My read is that if you let a customer do anything that adds a custom field, or table, or database, or trigger, or anything that is "structural" (even in the regular relational stuff) anywhere in your database (metrics or not), you are in violation. There doesn't seem to be a distinction about whether this is "direct" control or not, or whether a setting indirectly adds a trigger. I don't want to be in a courtroom debating whether a new metric is a "structural change!"
Now, none of that might be the intent of the license, but you have to go by what it says, not intentions.
The sad part of that is, I, and I'm sure many folks, have no interest in starting a database company, but we can't rally timescale because of legal risk. Looks awesome otherwise, though.
Hi Eric, thanks for taking a close look at our license.
I'd like to dispel some misconceptions:
The core of TimescaleDB is Apache2. Advanced features are under the Timescale License.
Regarding this:
the customer is prohibited, either contractually or technically, from defining, redefining, or modifying the database schema or other structural aspects of database objects, such as through use of the Timescale Data Definition Interfaces, in a Timescale Database utilized by such Value Added Products or Services.
My read is that if you let a customer do anything that adds a custom field, or table, or database, or trigger, or anything that is "structural" (even in the regular relational stuff) anywhere in your database (metrics or not), you are in violation. There doesn't seem to be a distinction about whether this is "direct" control or not, or whether a setting indirectly adds a trigger. I don't want to be in a courtroom debating whether a new metric is a "structural change!"
That's not correct, and we took pains to clarify that in the license: 3.5 "Timescale Data Definition Interfaces" means SQL commands and other interfaces of the Timescale Software that can be used to define or modify the database schema and other structural aspects of database objects in a Timescale Database, including Data Definition Language (DDL) commands such as CREATE, DROP, ALTER, TRUNCATE, COMMENT, and RENAME. [0]
Strictly speaking, if you provide Data Definition Interfaces (DDL) to customers via a SaaS service (ie you are running a TimescaleDBaaS - which applies to < 0.000001% of all possible users) you are in violation of the license. But otherwise you are fine.If you are looking for more votes of confidence, today there are literally millions of active TimescaleDB instances, including by large companies like Walmart, Comcast, IBM, Cisco, Electronic Arts, Bosch, Samsung, and many many smaller ones. [2]
If you have any other questions, I'm happy to answer them here, or offline (ajay at timescale dot com).
[1] https://www.timescale.com/legal/licenses#section-3-5-timesca...
The phrasing "such as" in "such as through use of the Timescale Data Definition Interfaces" looks to me like it can be interpreted as saying "Including but not limited to"
It was a little cumbersome to list every DDL SQL command, which why it uses that language. But that is the intent.
If you have a specific question, happy to answer it here (or offline)
Also, we used this language deliberately to provide more clarity. DDL vs DML is a pretty clear line to most developers who use TimescaleDB (vs some other companies who use language like, "you can't compete with us" etc).
The issue is the 'such as', which reads to me as indicating that providing an API endpoint that can add fields would also be a way of letting the customer modify the database schema and therefore covered.
Which means that building a SaaS app backed by timescale that has any level of customisation exposed to the user appears to be prohibited.
This seems a rather stronger level of prohibition than stopping people directly competing with you, and would suggest that if it's intended it would help to make it more explicit, and if it isn't then an explicit statement of that would be worth adding.
But I appreciate the feedback on how we could make our language clearer. Will share with the team!
a) Oracle buys TimescaleDB
b) Oracle sues any of 1000+ SaaS apps that have decent revenue and they can identify
c) The other 900 suddenly have massive due diligence issues even if not sued
I understand from your perspective this whole discussion probably looks silly, because you know what’s in your mind.
There are just too many stories of this blowing up in someone’s face (mostly due to legal, not due to actual enforcement) :/
My interpretation doesn't hinge on the DDL/DML question. If what you said is your intent, the legal language used is wrong and you should fix it. Consider this a bug report. Here's a source!
https://www.law.cornell.edu/definitions/uscode.php?width=840...
In order for the license to be usable, you need to be limitative here.
If I have a table that records timeseries data and then another table that has a customer-provided extensible set of metadata where a customer can define columns and other related tabular data, would that violate the license? The customer doesn't have a direct, like, psql level of access but the API intentionally provides a very similar level of interaction.
Does this qualify as providing Data Definition Interfaces? If none of those additional columns and such appear on tables set up as timescale tables does that make any difference?
The Timescale License only covers the TimescaleDB code. Postgres code continues to be covered by the OSS PostgreSQL License [0].
So putting aside the question about API vs. psql level (again, the Timescale License was drafted to enabled this for "Value Added Services", e.g., where "such value-added products or services are not primarily database storage or operations products or services"), this license wouldn't apply for non-Timescale code.
[Timescale co-founder here]
[0] https://github.com/timescale/timescaledb/blob/master/NOTICE
I appreciate from your POV this probably is silly but for us it’s very helpful to have explicit clarity since investors are questioning and demand certainty. I once had to rewrite https handshakes because a lawyer thought export compliance laws would be violated, so hopefully you understand my trepidation :-)
One other bug report on the license language front, this language could be construed to prohibit uses like Prometheus--unless there's a definition of operations products I missed (possible).
> are not primarily database storage or operations products
Suggested edit:
> are not primarily database storage or *database operations products*
Yes, database is a modifier to both storage and operations: database (storage or operations) products.
It's a good edit; will keep in mind.
(For good reason, we generally just don't like to "update" the text of the license too much, even for minor nits.)
Thanks!
Understood. If you see the other thread regarding the "such as" language, there is a serious edit that you can batch-up and repair to reflect your intentions with this one.
The "such as" language, retains your right to sue anyone for license violations if their API allows any customer action that causes structure changes indirectly, via the DDL, even under the hood (materialized views, too, presumably). That's way way more use cases than just repackaging TSDB as a service. That's a landmine, which when people compare and choose databases, they'd just assume avoid, even if otherwise comfortable with a cloud-protective license. Making this clearer and less onerous probably will probably pay for itself with a wider top-of-funnel for the product with more people more confident in the license.
The "We Clarified It In a Thread on Hacker News Public License" is probably not as ideal as updating the places that need clarification. :-P
How so? An end user should prefer a database under a license that protect the developer and users from cloudification/proprietization/SaaS
This worries me and makes me wonder if they are going for the open-source-only-by-name model.
[Please reply instead of giving silent dowvotes.]
> As an alternative, you can provide DCO instead of CLA. You can find the text of DCO here: https://developercertificate.org/ It is enough to read and copy it verbatim to your pull request.
> If you don't agree with the CLA and don't want to provide DCO, you still can open a pull request to provide your contributions.
https://github.com/ClickHouse/ClickHouse/blob/master/CONTRIB...
Anyway, Yandex CLA will be removed in the upcoming days (it should be already removed).
I don't know about ClickHouse but the other 2 uses bitmap indexes to make storing petabytes of data affordable.
Row oriented databases would struggle to compete against ClickHouse. They are easily an order of magnitude slower.
For example, there are a couple varieties of Bloom filters, which allow you to test for presence of string sequences in blocks. This allows ClickHouse to skip reading and uncompressing blocks (actually called granules) unnecessarily.
By Splitbee: https://github.com/ClickHouse/ClickHouse/issues/22398#issuec... By GitLab: https://github.com/ClickHouse/ClickHouse/issues/22398#issuec... And others: https://github.com/ClickHouse/ClickHouse/issues/22398#issuec... https://github.com/ClickHouse/ClickHouse/issues/22398#issuec...
If you'll find more, please post it there.
TimescaleDB can work pretty fine in time series scenario but does not shine on analytical queries. For most of time series queries, it is below ClickHouse in terms of performance but for small (point) queries it can be better.
The main advantage of TimescaleDB is that it better integrates with Postgres (for obvious reasons).
There are also many comparisons of ClickHouse vs Citus. The most notable is here: https://blog.cloudflare.com/http-analytics-for-6m-requests-p...
ClickHouse can do batch DELETE operations for data cleanup. https://clickhouse.com/docs/en/sql-reference/statements/alte... It is not for frequent single-record deletions, but it can fulfill the needs for data cleanup, retention, GDPR requirements.
Also you can tune TTL rules in ClickHouse, per table or per columns (say, replace all IP addresses to zero after three months).
@zX41ZdbW@ - Thanks for pointing out the various benchmarks that have been run by other companies between Clickhouse and TimescaleDB using TSBS[1]. As we mentioned, we'll dig deeper into a similar benchmark with much more detail than any of those examples in an upcoming blog post.
One notable omission on all of the benchmarks that we've seen is that none of them enable TimescaleDB compression (which also transforms row-oriented data into a columnar-type format). In our detailed benchmarking, queries on compressed columnar data in Timescale outperformed Clickhouse in most queries, particularly as cardinality increases, often by 5x or more. And with compression of 90% or more, storage is often comparable. (Again, blog post coming soon - we are just making sure our results are accurate before rushing to publish.)
The beauty of TimescaleDB columnar compression model is that it allows the user to decide when their workload can benefit from deep/narrow queries of data that doesn't change often (although it can still be modified just like regular row data), verses shallow/wide queries for things like inserting data and near-time queries.
It's a hybrid model that provides a lot of flexibility for users AND significantly improves the performance of historical queries. So yes, we do agree that columnar storage is a huge performance win for many types of queries.
And of course, with TimescaleDB, one also gets all of the benefits of PostgreSQL and its vibrant ecosystem.
Can't wait to share the details in the coming weeks!
Which database would be a good fit for this? There isn't too much data, maybe tens of thousands of rows eventually. Would Timescale be a good fit? I'd prefer that, due to existing familiarity with Postgres, but if ClickHouse is better, that's good too.
But it can't be updated or deleted, so what do you mean by this?
The overall concept that I was intending to highlight is that you can benefit from both row & columnar store in TimescaleDB. Chunks that are not yet compressed (row store data) can be modified (INSERT/UPDATE/DELETE) as usual and it's transactional - so you're assured it's been completed.
As of TimescaleDB 2.3, compressed chunks (columnar) do allow INSERTS but UPDATES/DELETES on compressed chunks are not yet supported natively. You _can_ decompress any chunk and modify the data (again, transactionally) as needed and recompress.
Initial discussion: https://github.com/ClickHouse/ClickHouse/issues/19627
Being implemented: https://github.com/ClickHouse/ClickHouse/pull/24755
Clickhouse has ALTER ... DELETE and ALTER ... UPDATE functionality now! (and TTLs)
We've recently been working through a detailed benchmark of TimescaleDB and Clickhouse. The DELETE/UPDATE question has been an intriguing story to follow - and I honestly hadn't considered the GDPR angle.
ATM, Clickhouse is still OLAP focused and their MergeTree implementation does not allow direct DELETE (or UPDATE) of any data. All DELETE/UPDATE requests are applied asynchronously by (essentially) re-writing/merging the table data (it's referred to as a "mutation") without whatever data was referenced in the DELETE/UPDATE. [1]
[1]: https://clickhouse.com/docs/en/sql-reference/statements/alte...
Data for in-active users gets deleted because our clickhouse retention policy is lower than the in-active-user timeout
I use them every now and then, but I prefer working with partition strategies when I have to these programmatically.
Most products do the asynchronous rewrite, especially if they're based on immutable storage. That's fine, but it should be tested to verify that it's not triggering on every delete, for example, and that it's resource-efficient.
Altinity is fixing this. The project is called Lightweight Delete and it's for exactly the GDPR reason cited. The idea is that there will be a SQL DELETE command that causes rows to disappear instantly. What actually will happen is that they will be marked as deleted, then garbage collected on the next merge.
Disclaimer: I work for Altinity.
I love the confidence here.
Clickhouse optimizes on the 2 most important things for OLAP - minimal disk space due to compression benefits of columnar storage and minimal compute for the same reason - and therefore fast.
However it isn’t flexible when you want to expand the use case. You can’t do any sort of text search, no complex joins (there are no foreign keys), and you need to order you tables there way you want to sort them.
For certain things it’s perfect. It was built to solve a problem Yandax had and that’s notable. But it doesn’t have anywhere near the flexibility of Elasticsearch for example.
But yes it’s purpose built to be extremely fast and minimize storage for the types of use cases it is built for.
That is false. I have built a large scale system that does tons of text searches, complex joins and even queries on top of JSON objects with performance that rivals BigQuery and surpasses it in terms of cost.
Edit: actually much better performance when you account for BigQuery’s cold start scenario.
ClickHouse is used for everything from log management to managing CDN delivery to real-time marketing and many other applications. It's gone far beyond the web analytics use case for which it was originally developed at Yandex.
Edit: clarification
We still call this OLAP but it's quite different from traditional uses. In particular the core data come event streams.
https://altinity.com/presentations/2020/06/16/big-data-in-re...
https://altinity.com/webinarspage/2020/6/23/big-data-and-bea...
In any case though I expect a lot of growth in Clickhouse community now and investment both engineering and most importantly Marketing - I think Clickhouse technology has a lot more adoption potential than it currently has
But I presume the GitHub link (https://github.com/ClickHouse/ClickHouse/blob/master/website...) has been submitted because clickhouse.com is going to be blocked for a large fraction of HN users (Peter Lowe’s Ad and tracking server list, which I think uBlock Origin has enabled by default, includes ||clickhouse.com^). I’m actually a bit curious why clickhouse.com (or more likely a subdomain?) would be being used this way; I’d have thought that they’d separate any such uses to a different domain so as not to hinder their main domain which is about the software and nothing to do with ads or tracking at all (even if that’s probably the main end use of such an OLAP DBMS).
This was a very old entry - it was added on Fri, 06 Jun 2003 19:53:00. Back then it was a marketing company that served ads.
I pride myself on knowing the entries in my list very well, but I have to admit I forgot about this one, which is ironic because I use Clickhouse at my job these days.
Thank you for your hard work. Every day it makes my experience of the internet 100x better.
Thank you for maintaining such lists. You and a few others like you save me much time and aggravation.
I don't look forward to when Chrome enforces Manifest v3 when I'll probably have to wait for a whole extension to be updated instead of just a list file.
When you don't have VC money and you HAVE to profit, you really learn optimize the workflow.
What do you mean by "do a lot". Can you deliver as quickly as a team? If so, do you work more hours or are you just better? If you're just better, why do you decide to stay with your largish tech company when we're acknowledging SV pays more? A remote role would increase your salary, no?
Breaking this down:
Russia has a great mathematics and engineering education system. Many graduates, unable to leave, take jobs with Russian tech companies. Russian tech companies pay less than US tech companies.
That's why the situation may be as is with ClickHouse. You're not in Russia.
Texas isn't particularly known for running lean. Every Big Tech has a presence in Austin. Dallas is filled with legacy financial companies burning money on IT.
I would validate this by getting an offer. Even adjusting for location, my guess is it's still significantly higher than what you're making locally. The adjustment isn't, say, you live in SF so you make 400k, you live in Houston you make 150k. It's ~15-20% for most places, at most.
> Yeah, I use RN to do all the platforms.
I want to point out those bloated teams that aren't lean from SV? They made ReactNative. They made Flutter too if you were thinking of swapping.
I know Google and FB made Flutter and RN respectively, and I thank them for that. That doesn't mean other SV companies aren't bloated, FAANG has a lot of money from their spigots (Google & FB have their ad money streams) and live on another level, not VC funding.
Cross platform has come a long way and enables small companies to do a lot more with less. Flutter is nice but has flaws, RN is the sweet spot. I could talk for days about the pros and cons of both.
A previous comment of mine regarding of Flutter vs RN: https://news.ycombinator.com/item?id=28394396
I was just talking with a friend who recently left Google. He's now trying to figure out what to work on next and has spent the months since reading widely. As we were talking, he gestured at the wall of bookcases behind him and said, "I'm not really concerned with maximizing income. I can already buy all the books I can read."
And personally, I'm at a not-for-profit because I want impact. I could make a lot more money elsewhere, and I certainly have in the past. But when I look back on that stuff, a lot of it just looks like a waste of time to me. The financial traders I worked for took in money that would otherwise have been hoovered up by other traders. The excellent code base that never got any users because the business side was kinda fucked up. The enterprise system that limped along a while longer thanks to our stress and overtime. Life's too short.
And I really get hunterb123's perspective here. I'd rather be part of a small team getting shit done rather than a highly paid developer on a vast effort to shift some ad-revenue metric by 0.2% over the next quarter. Some people like that and it's fine. But in interviews I've asked enough former FAANG developers, "So why did you leave?" that I know it's not for me.
tldr; Please consider contextual economic realities before applying reductive backhanded comments.
[0] I'm shooting from the hip, here. Apologies if my example numbers are way off for part or most of the regions I'm mentioning. Even if so, there was a day -- not very long ago -- when they were quite close.
And also Altinity is a trustable partner with a great know-how about Clickhouse internals. They have started to offer managed instances in AWS: https://altinity.com/altinity-cloud-test-drive/
At last, Alibaba Cloud has an option to use: https://www.alibabacloud.com/product/clickhouse
Are there any other ones?
Despite Yandex (who originally built Clickhouse) offering a managed solution, a substantial investment outlay from the VCs does come off as a huge vote of confidence in the founders.
BTW congrats to Alexey on the new company.
There's been so many queries where I've thought 'that's going to need a join and aggregation across tens of billions of rows, no way!' - and then Clickhouse spits back a query result in 10 seconds...
Disclaimer: I work for Altinity, which operates the Altinity.Cloud service.
1.) Use Clickhouse as infrastructure to build a product similar to MixPanel / Amplitude 2.) If I wanted a basic MVP of above can anyone point to me in steps (like 1., 2., 3., etc.) on what I would need to do to have a basic MVP ready. (Note: I am already very familiar with Docker, Kubernetes and writing rest APIs) Would greatly appreciate this since it would clear up a lot of questions I have
(Or you can use PostHog, which has essentially done all this for you and has all the functionality that Mixpanel/Amplitude has, but you're able to self host it!)
Can you explain this in a little more detail:
“Which would be read by Clickhouse” are talking about something like a Kafka connector? Or some Ksql type query?
Edit: I tested it but for some reason either the docs are strange or wrong. but TOAST tables are actually replicated?! or at least I see the data?
ClickHouse is a columnar, analytic, close-to-a-DBMS, but not a full-fledged one. The "100x-1000x faster" is compared to row stores. Last time I checked it was mostly single-table-oriented.
We compared Timescale, Clickhouse, Snowflake and Firebolt. Ended up really liking Firebolt, some amazing tech with a few roughedges (its pretty new), basically Clickhouse speed meets Snowflake simplicity definitely one to watch.
I guess this is used heavily in advertising?
Yandex N.V. is the largest
internet company in Europe
and employs over 14,000 people
I’m quite sure there are larger “internet companies” in the EU such as Booking, Zalando, etc.