TiDB – cloud-native, distributed SQL database written in Go
github.com
github.com
1. SQL front end nodes 2. Distributed shared nothing storage (TiKV) 3. Meta data server (PS) 4. TiFlash column store
1 and 3 are written in Go 2 is written in Rust and uses RocksDB 4 is written in C++
2 & 3 are graduated CNCF projects maintained by PingCAP.
Disclaimer: I work for PingCAP
A side aspect of this is that they also destroy any motivation for developing free software other than as a marketing tool, a way to drive growth. While that is perfectly fine for companies to do this, it has long term consequences that lead to bad behavior that mean you should be extra wary of relying on them.
TiDB is Apache licensed, that should be enough no? Because if we ban all projects from VC-backed startups, you're going to clear your tech stack pretty quickly.
I'm afraid that they can always do that with or without a CLA. Apache is quite a permissive license.
So if I make a contribution under some very restrictive license, but sign a CLA allowing relicensing, all bets for the future are off.
Therefore, it doesn't make a difference to have a CLA or not. What you trying to avoid is already allowed as per the license.
However, some CLAs require you to declare the originality and ownership of your contributions. Maybe that's why they require a CLA to be signed for contributions.
I think the best way would be something similar to the Linux Foundation. Companies in need of a certain type of database would pool resources to develop and maintain it.
PostgreSQL was a community developed database that was funded in something approaching this way as most of the developers were either students or worked for someone who paid them while they worked on it.
> TiDB is Apache licensed, that should be enough no?
It's good enough to use it with the knowledge that it isn't truly free (libre) because the CLA gives the owning company the right to change that at any time. So you can use it but don't build a business around it or make it a critical part of your infrastructure as they might pull the rug out from under you.
Do you have any links to the Core Storage CNCF project? I couldn't find it on the CNCF website under the Database and Cloud Native Storage sections?
It isn't copyleft, so even without a CLA, they could make a proprietary or open core version of the product, as can anyone else.
Disclaimer: I currently work at PingCAP and previously worked at the Linux Foundation.
Foundation (whether it's the LF or others) isn't a panacea either. It's fine when projects get started and there are plenty of willing member companies, but for many projects, companies often lose interest, need to cut down on open source related investments, etc. and the project funding dwindles. The reality is that the member companies need to both fund and provide software developers for projects, and it's difficult to expect them to keep the same commitment for more than a few years these days....
A popular more lightweight solution is the DCO[0].
a distributed database is potentially more complicated to operate, and optimize, and because its new and potentially has more sharp edges maybe less reliable (?) but the extraction of moderate sized datasets doesn't really seem to be an obvious failing.
People who bang on about Postgres replication have rarely setup replication in Postgres themselves and that too in the 100a of Pb scale.
MySQL replication works well and can be scaled more easily (relative to Postgres) but has its own problems. eg., DDL is still a nightmare, lag is a real problem, usually masked by async replication. But then eventual consistency makes the application developers life more complicated.
My fear is companies without in-house RDBMS expertise see these products as a way to continue to avoid getting that expertise.
I have strong doubts about distributed DBs in general, but I also can’t see them blatantly lying in a talk.
Most tech companies have poor knowledge of proper data modeling and SQL, leading to poor schema design, and suboptimal queries. Combine that with the fact that networked storage (e.g. EBS) is the norm, and it’s no wonder that people think they need another solution.
The amount of QPS you can get out of a single DB is staggering when it’s correctly designed, and on fast hardware with local NVMe disks (or has a faster distributed storage solution). Consider that a modern NVMe drive can quite easily deliver 1,000,000+ IOPS.
> Consider that a modern NVMe drive can quite easily deliver 1,000,000+ IOPS.
There could be other bottlenecks, for example I am consistently experiencing that linux kernel doesn't handle memory pages allocations fast enough once disk traffic hits few GB/s, because it does it in single thread.
Now go to OLAP, and a single query might be doing multiple table joins. It might be scouring billions of records. It might need to do aggregations. Suddenly "millions of ops" might be reduced to 100 QPS. If you're lucky.
And yes, that's even using fast local NVMe. It's just a different kind of query, with a different kind of result set. YMMV.
But yes, OLAP is of course its own beast, and most DBs are suited for one or the other.
The organization might say "Okay. Maybe you should do your ad hoc exploration on an OLAP system. Preferably our data warehouse where you can let your report run for hours and we won't see a production brownout while it's running."
So complexity of ad hoc joins in the warehouse generally can get more complex.
Which means interest in Postgres (specifically) is only 9.13% of overall interest in databases; MySQL another 14.04%. Combined 23.27%.
Is that a significant percentage of interest? Yes. Many others are a fraction of 1% of mindshare in the market.
Yet the reason there are 423 systems ranked in DB-Engines is because no one size fits all data, or data query patterns, or workloads, or SLAs, or use cases.
PostgreSQL and MySQL are, at the end of the day, oriented towards OLTP workloads. While you can stretch them to be used for OLAP, these are "unnatural acts." They were both designed in days long ago for far smaller datasets than typical for modern-day petabyte-scale, real-time (streaming) ingestion, cloud-native deployments. While many engineering teams have cobbled together PostgreSQL and MySQL frankenservers designed for petabyte-scale workloads, YMMV for your data ingest, and for p99s and QPS.
The dynamic at play here is that there are some projects that lend themselves to "general services" databases, where MySQL or PostgreSQL or anything else to hand is useful for them. And then there are specialized databases designed for purpose for certain types of workloads, data models, query patterns, use cases, and so on.
So long as "chaos" fights against "law" in the universe, you will see this desire to have "one" database standard rule them all, versus a Cambrian explosion of options for users and use cases.
I’m not saying it fixes everything, but knowing how a DB works, and applying proper data modeling and normalization could severely reduce the size of many datasets.
I think people’s idea of scale and operating at scale is limited to their experience.
You can get MySQL to run at any scale, look at Meta and Shopify. Operational complexity at that scale is a different story.
Distributed databases reduce a lot of the operational complexity. To take one example:
Try a DDL on a 5 TB table in any replicated MySQL topology of choice and compare it with TiDB’s Distributed execution framework.
That’s a new one to me; never have I ever heard someone claim that making a system distributed reduces complexity.
I’ve operated EDB’s Distributed Postgres at a decent scale (< 100 TB of unique data); in no way did it enable reduced complexity. For the devs, maybe – “chuck whatever data you want wherever you want, it’ll show up.” But for Ops? Absolutely not, it was a nightmare.
https://jepsen.io/analyses/tidb-2.1.7
The devil is in the details, and anyone who is looking to implement TiDB for data correctness should read through not just this but other currently-open correctness-related Github issues:
e.g., https://github.com/pingcap/tidb/issues?q=is%3Aissue%20state%...
TiDB currently has 74 open issues related to "correctness," and has closed 163.
In the previous event (2023) there were speakers from Airbnb, Databricks, Flipkart, PayPay, and others sharing their experiences as well - https://www.pingcap.com/htap-summit/sept-2023/
Disclosure: Employee of PingCAP the company behind TiDB
Wasn't last year 2024?
Reference: https://docs.pingcap.com/tidb-in-kubernetes/stable/get-start...
(Also, bold move to write a db without fully being able to manage memory).
> Distributed Transactions: TiDB uses a two-phase commit protocol to ensure ACID compliance, providing strong consistency. Transactions span multiple nodes, and TiDB's distributed nature ensures data correctness even in the presence of network partitions or node failures.
> High Availability: Built-in Raft consensus protocol ensures reliability and automated failover. Data is stored in multiple replicas, and transactions are committed only after writing to the majority of replicas, guaranteeing strong consistency and availability, even if some replicas fail. Geographic placement of replicas can be configured for different disaster tolerance levels.
See https://github.com/pingcap/tidb?tab=readme-ov-file#key-featu...
Correctness has been a focus for a long time for TiDB, including working on passing Jepsen Tests back in 2019, see https://www.pingcap.com/blog/tidb-passes-jepsen-test-for-sna... and https://jepsen.io/analyses/tidb-2.1.7
Disclosure: Employee of PingCAP the company behind TiDB
The first paragraph I find under Key Features is:
Distributed Transactions: TiDB uses a two-phase commit protocol to ensure ACID compliance, providing strong consistency. Transactions span multiple nodes, and TiDB's distributed nature ensures data correctness even in the presence of network partitions or node failures.
It claims to ensure correctness in the presence of partitions, not availability, so I don't think that's a claim that CAP is wrong? I would expect it to be unavailable if it can't provide a consistent response.
I don't know about memory management. It seems go may not be used for the storage layer, but in any case, it wouldn't be the first attempt at a DB on a garbage collected runtime. I can think of at least Cassandra on the jvm and CockroachDB also on Go. I do prefer that DBs do their own memory management (as well as storage) but this is not as unique as it used to be though.
Additionally, it can push down the DAG to the TiKV storage nodes, written in Rust, to reduce movement of data and work closer to the physical data.