What to Expect in CockroachDB 1.0
cockroachlabs.com
cockroachlabs.com
Do range queries on a simple table also translate to underlying key-value range queries with minimal overhead?
Also, how's the performance nowadays?
> FDB-SQL stores metadata in the KV store itself. This is great from a distributed correctness point-of-view. It’s also great from a simplicity point-of-view. If I trust the underlying system to be safe and consistent, then my metadata is also safe and consistent. But now I need to do a ton of reads before I run my SQL to know how to run my SQL. Where’s my data, for example?
[1] https://www.voltdb.com/blog/2015/04/01/foundationdbs-lesson-...
CockroachDB also stores SQL metadata inside of the KV store, but that metadata is also gossiped around the cluster (i.e. it is replicated to every node) so that SQL execution almost never has to read it from the KV store. Handling changes to this SQL metadata is challenging and required a design [1] that is unable to take advantage of the easy correctness of simply doing reads from the KV store on every operation.
[1] https://github.com/cockroachdb/cockroach/blob/master/docs/RF...
(For the record -- in case I seem to be especially pessimistic -- I'm actually extremely hopeful that CRDB will actually make scaling OLTP applications significantly easier for everyone. It's just that this probably shouldn't be FDB all over again.)
This is a really hard question to answer because there are so many different scenarios to test. Even if you restrict the question to KV operations a simple answer isn't forthcoming. For example, if you want a distributed system with the fastest writes, then Cassandra is your best choice. Oh, but you better pay a lot of attention to all of the caveats regarding tombstones and node outages and overwriting data.
Overall, performance has improved a lot over the past year, but we still have a lot of work to do. Some of the high-contention scenarios that caused poor performance in the Jepsen tests have seen a huge improvement (10-100x), though admittedly we were starting from a poor position.
So far we've performed the most significant comparisons against Cassandra and MongoDB. In terms of performance Cassandra > MongoDB > CockroachDB, but note that most (all?) benchmarks test idealized situations. YCSB tests ideal write scenarios where data is never overwritten or deleted and data correctness is easily achieved because there are no node outages.
Unsurprisingly, PostgreSQL will outperform a single-node Cockroach cluster for simple operations. The internal target was to be within 2x of PostgreSQL performance for simple INSERT, UPDATE, DELETE and SELECT statements. We were at that target earlier in the year, but a quick check reveals that we've slipped a bit.
Sorry I don't have anything more concrete to offer at this time. I realize I've been saying "we're working on it" for too long. Really, we are.
Google's Spanner is pretty great too but also missing any json columns and too locked into GCP. We're also looking at using a JanusGraph on top of ScyllaDB to get better querying ability with strong multi-region access.
Slightly limited but distributed SQL queries can run across all nodes to speed up access or scan lots of data - or you can choose something like a customer_id to shard all the tables the same way so that all data for that customer would be on the same node, allowing for full SQL syntax and features on data of that node (because it would just be a regular postgres server).
For a data-warehouse/fast-realtime scenario, I would also recommend MemSQL although it is proprietary.
Why don't they just do what porn companies do, create a professional sounding parent company that can interact with various enterprises while leaving the underlying implementations opaque?
Even if an author thinks a name like AIDS.app, murder.js, syphilis.io, represents the nature of the project well, I'd advise him to change it.
Same thing happened with the "Nimrod" programming language. "But it means a great hunter!" No, people hated the name every time it came up and its author eventually renamed it to Nim.
In reference to radioresistance: While an LD50 has been reported for wild type C. elegans individuals, an upper lethal limit has not been established, rather "nearly all animals were alive with no indication of excess lethality up to 800 Gy, the highest dose... measured."
I don't like equipment that's rusted. I don't like when joints have rusted into position.
I don't like getting rusty at skills.
I don't like rusty pipes or rust getting into drinking water. I don't like turning on a faucet and seeing the color of rust come out.
I know it's not 100% rational, but the feeling of revulsion nudged me away from learning about it until much later.
It also invites easy puns about a system that runs "like it has a lot of rust in it" (already seen it happen).
"Oh, but it's actually this ubergeeky reference to--" I don't care. The lay audience cares even less. The insidious metal oxide is what I think of, double so for the average person.
You can argue that potential managers might be influenced but it should be easy enough to put the merits of the product before the name itself.
If so, then you agree the name matters, "and we're just haggling over the details."
Stuff like this:
https://www.amazon.com/Never-Sleeps-Young-Crazy-Horse/dp/B01....
https://www.pinterest.com/tmdesigns56/rust-is-beautiful/?lp=...
It's an interesting question of how much it's worth trading a memorable name for rubbing some people the wrong way.
At some point, these people who constantly protest the name should realize they've been heard and let the discussion focus on something other than their pet peeve that has nothing to do with the underlying tech. Yes, we realize you have a problem with the name. We realized it a dozen news items ago. But since this is HackerNews, you'd think other aspects -- distributed transactions and the approach they took, the fact that this is in Go (!) and the benchmarks on it, the single executable that incorporates the C++ RocksDB embedded leveldb, etc, etc -- would take top billing.
This is an issue of basic hygiene of marketing. The name is obnoxious.
It is like being a great developer and showing up in a interview looking/ dressed like a homeless person. Yes, we get it, you can be all talented and such, but people will have a visceral response to the lack of basic hygiene and not want to associate with that person.
In this case it is yet another database.
That being said, I tested npgsql a few days ago and it worked perfectly. Nothing too extensive, just created a db, 2 tables, inserted some data and did a join. So please give it a shot and if it doesn't work, file issues and we'll address them.
(disclosure: I work at cockroach labs)