Jepsen: Crate 0.54.9 version divergence
aphyr.com
aphyr.com
I love these Jepsen posts, and Kyle Kingsbury's work is amazing.
It all shows that distributed systems are really complicated and difficult to get right. These posts on the open source ones are really great. But how do you go about evaluating a proprietary one?
Let's say my company is looking for a columnar store. We look at the evaluation of Cassandra and think, "Nope."
But then someone says, "Hey! Let's go with Redshift!" How do you know that's any better?
We all know that DB makers claim certain things. Why would Amazon be better at this tricky subject than the people behind ElasticSearch?
My question is that I would really like to know what people think about this?
Is the decision that it's better to go with the unreliable demon that you know? Or is it better to pass the responsibility off to a company like Amazon, hope for the best, and lawyer up if unexpected things happen?
http://www.datastax.com/dev/blog/testing-apache-cassandra-wi...
I meant a general hypothetical situation where I'm the decision maker and also the DBA.
I have my own experiences with different products. Likely a lot of other people do too. I don't really think that the weaknesses exposed in the Jepsen tests would be covered by your everyday person on IRC.
I mean, I'm one of those everyday people on IRC talking about the things I do and don't know about. I could tell you from personal experience that I've had a bad run with MySQL and MongoDB in production. But I couldn't tell you exactly what went wrong or where or why.
I could talk to you about the symptoms, but I couldn't address the root causes off the top of my head.
And that's exactly the point that I'm trying to ask about here.
Let's say that I know OSS database is weak in certain situations. Then I know exactly what those are and what to do to protect myself or at the very least what use cases not to deploy them for.
The same can't be said for proprietary products like Redshift. I might know, anecdotally that it does or doesn't fail, but I don't know exactly how or why or what to do to protect myself.
I was asking how other people in this situation make their decisions. Obviously, a lot of people use a lot of things that don't get CAP theorem right. And perhaps a lot of people never experience any problems. But if I'm going to be extremely thorough about this, I wonder what other people think about when making these same decisions.
The choice between proprietary data systems seems more likely to be driven by the reputation of the company than by the reputation of the tech. The whole point of paying when OSS alternatives are available is that when something breaks or goes wrong you actually have a proverbial "neck to choke" (one that isn't your own ;) ).
RedShift technically is a service for which you pay not a database (I did not use it but I believe it is Postgres - don't quote me on that), so if it doesn't provide what's promised you take it with Amazon.
Anyway the summarize databases. Currently we have two types:
- relational databases - old and proven, always consistent, but typically not distributed
- so called NoSQL which I think only worth attention are AP (from the CAP). These databases scale, but are eventually consistent. This means they are consistent most of the time, but not always. It's good for storing things that while important, it's acceptable when some individual entries are wrong or lost (user sessions, shopping carts, tracked information about users etc)
There are also databases like Mongo which are snake oil.
Generally these two narrow it down quite a bit. Consider how your app(s) would deal these possibly relaxed consistency guarantees.
Hoping for the best and thinking about lawyers shouldn't be part of any evaluation strategy.
; For each [version, reads] pair, discard those with one value
multis (remove (fn <a href="/data/posts/332/k vs">k vs</a>
(= 1 (count (set (map :value vs)))))