CockroachDB 2.0 released
cockroachlabs.com
cockroachlabs.com
I tried lucene based databases that offered amazing search capability but were riddled with data corruption issues. Then there was RethinkDB which was very promising but ran out of funding.
I am skeptical that a networked database with multiple nodes can match the performance of a single master database such as MySQL, PostgreSQL, or SQL Server. I did a quick benchmark of CockroachDB 1.x and MySQL last year and found that CockroachDB was 5-10x slower on simple CRUD queries: https://github.com/caleblloyd/MySqlCockroachBench/wiki/Concu... Are there any good independent benchmarks of performance?
According to crunchbase, Cockroach Labs has raised $53.5m (RethinkDB had only raised $12m). Is there evidence that Cockroach Labs is on track to make money and is going to survive?
If the company is healthy and the performance is verified as close to a major RDBMS, I would be comfortable trying it out as a primary datastore. If those questions can't be definitely answered right now, I'll probably continue to wait for the company and the tech to mature.
It likely can't unless there is some black magic going on. Single node speed will always be faster. But once you get to that point where a single node chokes on the amount of data or query throughput that you have, you don't really have a choice anymore.
My personal plan is to start with Postgres RDS. Grow until RDS doesn't work anymore, then move to ultra beefy bare metal servers in colocation with AWS Direct Connect. If I ever outgrow an 4x24 core server with 3TB of memory on a RAIDed NVMe disk cluster, I might move to Citus.
For the various distributed database companies out there, I believe the one that will win in the marketplace is the one working or partnering to develop specialized hardware and networking, and then optimizing for it.
You will still need another beefy server in another datacenter with configured replication and failover, which is one of the problem Cockroach is trying to solve.
Also bedrock has plenty of problems and is highly unrecommended.
soo... oracle? :b
However, if you look at the path Cockroach is taking, they're doing all the right things to shave those years down. They already Jepsen themselves, and they've been doing rigorous testing forever now.
Very exciting.
When Amazon did this for Aurora, they showed that with their custom storage backplane, they were able to absolutely murder normal RDS-based MySQL in terms of TPC-C throughput. Take a look: https://www.slideshare.net/AmazonWebServices/dat202getting-s.... The bottom right quadrant of that slide shows what happens to a standard MySQL server when given 10,000 TPC-C warehouses to chew on (about 800 GB of TPC-C dataset). The max throughput is supposed to be ~128,000 tpmC. RDS MySQL in their tests was able to do 69. That is quite simply, not impressive. Even Aurora, which is a full 136x the tpmC throughput, is still more than 10x less than the max throughput that should be achieved at that number of warehouses.
CockroachDB can scale out, so while it may have more latency when it performs a simple transaction or millions of simple transactions over a small set of data, it can also easily scale to handle 10,000 TPC-C warehouses, with 126,000 tpmC. If you're going to talk about performance, you really need a serious benchmark, or you're just kidding yourself.
Let me put this another way. While Amazon frequently talks about how Aurora can be scaled to 64TB of data, in the context of TPC-C, that's a risible claim. A maximal instance of Aurora, tuned by AWS, can handle max TPC-C throughput at something between 80GB and 800GB of data (i.e. 1,000 to 10,000 warehouses). It's like they're selling you a pickup truck with a bed they say can handle 64 tons, only if you load it that way, it can only travel 2mph, and can't turn. Caveat emptor!
https://github.com/cockroachdb/cockroach/issues/17777#issuec...
What distributed databases give you is scalability as data grows, and high-availability for safety. Performance comes from that scale and concurrency when your data and queries fit the model.
However, the latency will typically be higher since it always has to talk to at least 3 nodes.
If you consider CRDB's DistSQL which lets you run queries on multiple machines, it might be possible to beat a traditional RDBMS on read latency for large queries.
I like the automatic replication of CRDB but the vast majority of apps can't afford to take a big performance hit and don't need large horizontal scale-out.
They also only handle data corruption at the cluster level so there's not really any tools to extract data from a partially corrupted DB on a single node.
Even with MySQL, the replication is still asynchronous from binlog? Is it also the case with postgres?
Also has logical replication now, although I wouldnt recommend it because it doesn't support DDL changes yet. Overall it's a solid db choice, but there is still some work involved.
Otherwise I'd recommend SQL Server for a fantastic commercial database that runs cross-platform and has great tooling to make everything easier and some advanced features.
Technically, I don't really need multi-master, a "standard RDBMS" with automatic failover would probably fit the bill, but I could not for the life of me figure out any standard (and free) way of doing that for mysql/mariadb or postgresql.
I haven't gone back to try the full entity framework recently, but if you find any bugs, please let us know. If you want to roll your own, we're here to help too.
It's great if you want to contribute though but I'd recommend just working on the PG EF Core provider which already has some work towards CRDB specific features: https://github.com/npgsql/Npgsql.EntityFrameworkCore.Postgre...
While I keep hearing that the RDMS(Postgres, MySQL...) was built on year of experiences I alwawyas think why we cannot come up with something nicer, more friendlier to the old way. Eventually we will build up knowledge on the new thing and have a system that just as good but also as much friendlier.
Example: when looking at how MongoDB handle replication, it's so easy to just add a new node and have it join cluster without seeding some data and set binlog position (like MySQL).
Or looking at RethinkDB query language, it's just so much easier compare with SQL, at least to me.
Nowsday, I cry whenever I see SQL query. I learn to master it, but life's too short to write those kind of thing.
> According to crunchbase, Cockroach Labs has raised $53.5m (RethinkDB had only raised $12m). Is there evidence that Cockroach Labs is on track to make money and is going to survive?
I'm in a similar situation: new project, CRDB is a good fit on paper, but I'm unable to get traction even for a proof of concept exercise because of the risk associated with a database backed by a startup. This is at a large, conservative, 300k+ employee company, so I guess it's understandable... but regrettable nonetheless.
See you guys in 5 years ;-)
We also run Jepsen tests against the database nightly, ensuring that its isolation guarantees do not regress.
[0]: https://www.cockroachlabs.com/blog/cockroachdb-beta-passes-j...
https://www.cockroachlabs.com/blog/diy-jepsen-testing-cockro...
Ideally you set up 3 master nodes, with distributed etcd cluster for state, and enough machines to run your services with replication.
All of which is not bad at all to get started with if you use something like Kops.
Some useful docs: https://kubernetes.io/docs/admin/high-availability/building/
[1] https://docs.mongodb.com/manual/tutorial/sharding-segmenting...
https://www.cockroachlabs.com/docs/stable/install-client-dri...
That being said, I am really excited to try this out. 2.0 adding JSON support is what converted me to try this out for a project.
[0]: https://github.com/cockroachdb/cockroach/blob/master/docs/RF...
That seems like a scary situation, .. and you're correct. It doesn't change the reality though, heh.
So how does Cockroach compare to Postgres, or Maria, TiDB, etc on ease of management? Any thoughts?
Don't use any distributed relational database unless you actually have a need for it, like horizontal scalability for massive data, multi-regional access, or 100% uptime guarantees.
https://www.cockroachlabs.com/docs/stable/frequently-asked-q...
When is CockroachDB a good choice? CockroachDB is well suited for applications that require reliable, available, and correct data regardless of scale. It is built to automatically replicate, rebalance, and recover with minimal configuration and operational overhead. Specific use cases include:
Distributed or replicated OLTP Multi-datacenter deployments Multi-region deployments Cloud migrations Cloud-native infrastructure initiatives
When is CockroachDB not a good choice? CockroachDB is not a good choice when very low latency reads and writes are critical; use an in-memory database instead.
Also, CockroachDB is not yet suitable for:
Heavy analytics / OLAP
What level of reliability, how, what tradeoff in my table design will I need to make, same thing for availability, correctness and scale.
None of the info in this FAQ allow me to know that CockroachDB is the right choice for me against the competition which advertises the same generic DB marketing terms.
Edit: And yes, I can go and read the documentation and deep dive into the internals myself, but I don't care enough to do so, because I already have DBs that fulfill those use cases that I know off, and I would hope therefore that CockroachDB would make it very quick, easy and in my face to find the info that will make me go: Ah Ha, this is the distinguishing factor and the reason why I might want to favor and care to use CockroachDB the next time I've got such a use case.
The name was chosen in 2012, two years before the open source project was started. I had just gotten done with an exhausting and ultimately frustrating survey of OSS database products for the backend of a new private photo sharing service called Viewfinder. I'd tried and found wanting MySQL, Postgres, AWS SimpleDB, Hbase, Cassandra, and Riak.
I was annoyed. Why wasn't there a scalable, survivable, consistent database with transactions? I was even willing to drop transactions as a requirement – a terrible sacrifice. The frustration led me to write a manifesto. What would the "right" database look like?
I imagined it would be composed of symmetric nodes, require no external dependencies, spread itself naturally across availability zones for survival. Each node would autonomously replicate and repair data. These were the capabilities that led me to the name "cockroach", because they'll colonize the available resources and are nearly impossible to kill.
- Spencer Kimball
I don't think we should be censoring a discussion just because you've already had that particular discussion.
Obviously they didn't anticipate constant naming discussions.
People hate them, but they’re one of the animals most similar to us. We hate them because we don’t like to admit our nature.
Do you really think annoying comments on HN are the anomaly where people care about the name? Or is it in itself evidence that there is some percentage of the population for whom the name is an obstacle?
Initially I thought I could get used to it, but after many years watching in HN I haven't succeeded yet.
Perhaps you cut the word in half, but it just creates another two strong words and shows the survival power of the word :(
Calling it CDB doesn't click to me either.
Congratulations to the CockroachDB team. I have been using this on a few projects. I find the ease of scalability and redundancy very nice.
Also, the ability to use SQL commands on a transactional DB is REALLY helpful.
Keep it up.
https://www.huffingtonpost.com/molly-reynolds/why-brand-name...
PRAGMA hissing_noise();Are they really going to? As someone with many years of marketing/branding/advertising experience, CockroachDB is a bold, powerful and meaningful name, that, in my opinion, is one of the best branding cases I've ever seen for tech infra products.
Your point, which is that other people may use the same justification, doesn't validate the point. It just means that there may be more than one person who would use the same, poor, reason.
I think someone should fork it and use Palmetto Bug in the name instead in some creative way. It'd be keeping the spirit of the original name while also not having the same immediate disgust some people have with the name. In South Carolina, a common name to refer to cockroaches is the Palmetto bug.
Back to the awesome creation of God[3] commonly known as the cockroach. According to Wikipedia[4], there are some 4600 known species of cockroaches, and only 30 of them are associated with human habitats. The rest of them are basically outdoor roaches, so if you find one in your house they are probably there by accident and trying to find a way out.
In fact, the only species you really need to worry about is the so-called "German" cockroach, which Germans call the Russian roach, and Russians call the Polish roach (and I'm sure the Poles blame the Turks or some such group).
If you see this roach in your house then you are in trouble. They like dirty houses so clean your house. And they have evolved to avoid sweets so you have to buy the correct kind of bait[5]. I like to spin long yarns about my battles against these bugs to anyone who cares to listen. In the end you need to deal with the humans who are too friendly to these roaches. Look for college students who live 3-4 to an apartment and go home for the weekend, people who work long hours and leave bags of half-eaten fast food in their open trash cans, and other unfortunate neighbors. Knock on their door and give them packs of the aforementioned roach bait to leave around their dwellings (this requires some tact). In time you can certainly eliminate them, thanks to modern chemistry. It should cost less than $50 and a few months of vigilance.
Frankly the German roach doesn't really fit the profile of your average cockroach. It's slow and seems to be asking you, almost daring you, to kill it, knowing there are 20 others that will take its place. A standard cockroach is the kind that runs away really fast, and would be just as happy living outdoors. They are far more widespread and just as survivable as the pest species, and lend their name to an awesome database product that will probably outlive discussions about its name.
I for one hope the creators keep the name and use the human revulsion about it as a marketing benefit. No one will forget their encounter with the mighty cockroach, and business people will come to associate it with sustainable competitive advantage which cockroaches have in spades.
[0]: https://www.etymonline.com/word/cockroach
[1]: http://liquipedia.net/starcraft2/Roach_(Legacy_of_the_Void)
[2]: https://www.rockpapershotgun.com/2010/11/02/genetic-algorith...
[3]: Genesis 1:24-25
[4]: https://en.wikipedia.org/wiki/Cockroach
[5]: https://www.amazon.com/Combat-Month-Roach-Killing-Station/dp...
The good news is that DBs are a niche of their own and as such don't need to appeal to the masses, making the name choice somewhat less important. But I still think, as someone else already pointed out here, that selling an exec (or non-tech decision maker) on this will be just slightly more difficult because of the name, as irrational as that might seem.
Also, I wonder why Apache doesn't change their name. It would be a good opportunity for rebranding.
So. For me, personally, I don't care about the name. I generally care that it's great tech, and it clearly has a great team behind it. However....
If I worked at CockroachDB, and I saw the negative feedback around the name, I'd take it to heart. At the end of the day, the name is marketing for the hard work of their engineers, and marketing for the engineers that want to use this DB (remember, they need to sell it to their managers who may not be technical).
This issue can show up in unexpected ways. For example, for cloud providers like Compose (IBM company), would they be comfortable with putting "CockroachDB" on the front page? They might if it's good enough, but it's at least a consideration (i.e. another meeting, another stakeholder to convince).
Or how about an enterprise company that's going through due diligence, and when their client asks them about their tech stack do they say "CockroachDB" or do they obfuscate the name by saying "It's a high-performance distributed database". That's a crucial moment to market CockroachDB, and it could get lost. As sad as it is, saying that you're using MySQL "because Oracle" is a point of leverage for some sales people.
Is the name worth it? Asking honestly.
You shouldn't, marketing is not about your personal feelings or feedback on your marketing. Cockroach name is clearly superior to every other database name, look how memorable it is and how much buzz it generates.