TiDB – A distributed NewSQL database compatible with MySQL protocol
github.com
github.com
Most distributed newSQL databases are an SQL layer on top of a distributed KV store. What they try to do is hide the distributed reality of their database so it acts like a regular database from the client side. Of course there are always caveats that might not be completely obvious but can cause terrible performance.
We take the opposite approach. We make the user aware of the distributed nature and force the user to use the distributed database like a distributed database should be used. You must split your data into chunks (actors) and you have a full raft replicated SQL engine (SQLite) within that chunk.
I think TiDB is comparable more to CockroachDB.
Basially, ActorDB doesn't hide the fact that it's partitioned. Rather, it forces the client to deal with partitions (or "actors") at the application level.
In particular, this means that while you can have transactions spanning multiple actors, queries can't. So you can't do joins across actors, and if you want to select from multiple actors in one network roundtrip, you have to do a special kind of looping statement that first finds the actors to operate on, then executes a statement on each.
You can work with multiple actors at once, but the database doesn't pretend that it's a single database; rather, each actor is sort of like a separate database, and it's up to you to design the data model and the SQL statements to distribute the data in an optimal way.
In those cases where you do need joins or aggregations that span multiple actors, you'll have to jump through some hoops. Joins, in particular, are probably not going to be very efficient. You might precompute some data that other databases would figure out on the fly. Again, since I've not used it, I don't know all the ways you would work around such limitations.
Unlike ActorDB, TiDB and CockroachDB present the illusion of a single database, on top of a distributed key/value store. There's a magical execution engine that takes selects, even with joins, and automatically splits the query plan into multiple parallel requests to the shards that hold the data, and then merges the results back together. You can do "select * from sometable" and it will return your table in one piece, no matter how it's distributed physically.
There are certainly benefits and drawbacks to both approaches.
MongoDB's docs: "To capture a point-in-time backup from a sharded cluster you must stop all writes to the cluster"
Cassandra's docs: "To take a global snapshot, run the nodetool snapshot command using a parallel ssh utility ... This provides an eventually consistent backup. Although no one node is guaranteed to be consistent with its replica nodes at the time a snapshot is taken"
Riak's docs: "backups can become slightly inconsistent from node to node"
CockroachDB's docs: "The table data is dumped as it appears at the time that the command is started ... there is no guarantee that NOW() is monotonic in transaction order"
Does TiDB support consistent backups? Are there any docs covering how it is implemented?
So we can use any MySQL tools such as mysqldump or mydumper to backup the database consistently.
See the MVCC implementation here: https://pingcap.github.io/blog/2016/10/17/how-we-build-tidb/
Also, I'd assume that Rust would be better at expressing and pattern-matching against the sort of complex execution plan trees and query expressions you need for this sort of thing.
I was looking at Apache Spark the other day, which is written in Scala, and which has a planner that goes SQL -> AST -> logical plan -> physical plan. The planner/optimizer relies extensively on pattern matching to apply rules to the logical plan (pushdown and so on), and it manages a lot of this with "match" statements and not a lot of recursion.
But that stuff is murder in Go, which doesn't have pattern matching, or generics for that matter. I know it's bad because I'm in the middle of something similar right now, in Go. The Spark code is exactly how I wanted to organize it (minus the awful class inheritance they do), but that's not possible in Go.
Other than going through the kernel to talk to a library..
Actually we are using RPC to call Rust from go, and it has nothing to do with Go or Rust, because either way we have to do RPC call(by encoding and decoding data), currently we depend on protocol buffer and customized RPC , and we are trying to migrate to gPRC. We know Go doesn't have pattern matching and generics, and it's kind of pain sometimes. But still we like Go very much because of simplicity and concurrency.
Go and Java both share good concurrency support, reasonable developer productivity, moderate type safety and good code generation.
Also the whole thing reminds me of this: https://mobile.twitter.com/edd/status/400190499585544192/pho...
It is missing NewSQL, which is happening right now, but it shows how we are going in circles. There are benefits of NoSQL, particularly for type of data that it is ok once in a while to have individual values wrong or missing (tracking users, shopping carts etc). CRDTs are also useful. I'm wondering what NewSQL would bring, but I'm thinking that we will go back to traditional relational databases once again. We probably end up with some kind of hybrid approach, and decide for given piece of data whether it should be distributed (at the cost of consistency) or vice versa.