Hazelcast 3.8.3
jepsen.io
jepsen.io
Wow. I wonder if the documentation and naming schemes are bad because they're aspirational, or because the developers don't know what they're doing?
A developer can only test up to the point that their code works as they thought it should work. QA takes over from that point.
At my last job we held off running a chaos test suite against our platform in the certification environment until _after_ the QA department came back with their report and gave the platform all green checkmarks. This was an exercise in both validating the value of the chaos engineering/testing work as well as partially discrediting the idea that QA is capable of exercising enough of the system in enough ways to expose non-obvious/nuanced edge cases.
I mean I guess there's an argument to be made that QA could be engineering and building their own rich tooling like a chaos testing framework, but that seems extremely unlikely to happen in the current typical formulation of QA.
And of course they knew what they were doing since that was the whole point of Hazelcast. You could use the same Java constructs you knew and loved but on a distributed grid. Problem is of course that it wasn't well implemented and likely not such a good idea to begin with.
Also, his Twitter is a great thing for confusing people who want to read about lock schemes and relationship management etc...
Why use a lock that isn't safe when there are other alternatives?
"When there is a network partition we favour availability over consistency. Which is what the Jepsen test shows. [...] It is also the PACELC contract of the entire category of IMDG systems. IMDGs are used for maximum speed and low latency. They represent a set of design trade-offs to achieve this primary purpose."
Based on what data? Even hazelcast had that same claim before they got exposed?
I've only seen 2 systems that did rigorous in-house fault-tolerant testing (foundationdb and cockroachdb) and later when "jepsen-verified", they actually backed their claim or had few modifications to uphold their claim.
Talk of Hazelcast reminded me of Coherence (now Oracle, before Tangosol): https://www.javalobby.org//java/forums/t78008.html
Oracle docs sensibly call it a distributed cache: http://www.oracle.com/technetwork/middleware/coherence/distr...
When reading that the distributed map implements the ConcurrentHashMap interface we were immediately skeptical, but it took more work to prove our skepticism was founded.
Hazelcast has a fantastic foundation for building clustered applications, but some of the design choices they made both internally and API wise are not choices we would have made.
We are moving to a custom merge policy because of some of these choices.
"We wish to thank Jordan Halterman for his discussion of Hazelcast use cases. Luigi Dell’Aquila & Luca Garulli from OrientDB, and Denis Sukhoroslov from BagriDB, were instrumental in understanding those systems’ use of Hazelcast. Thanks also to Julia Evans, Sarah Huffman, Camille Fournier, Moishe Lettvin, Tim Kordas, André Arko, Allison Kaptur, Coda Hale, and Peter Alvaro for reading and offering comments on initial drafts. This research was performed independently by Jepsen, without compensation, and conducted in accordance with the Jepsen ethics policy."
I'd like to know who stands behind this research
completely off-topic: have you ever thought about giving your training (https://jepsen.io/training) in an online way (mooc, udemy, whatever)? I would love to learn about distributed systems from you, but today I think it would be almost impossible since you only seems to give in organization trainings.
It's a good course; better, IMO, at what it teaches than equivalent Udemy courses. And I'm a pretty good teacher, I can be pretty engaging while talking about this stuff and it's fun. But at the race-to-the-bottom prices of the MOOC economy, it's a nonstarter. The Udemy "every class is ten bucks" disease discourages really capable, competent people from sharing what they know.
(And Packt et al. finding somebody to read some slides is not a good counterexample. I said "really capable, competent" for a reason. I was approached by one of their competitors--a bigger company than they are--to write a book on Mesos on the back of two blog posts...)
For context, Jepsen started as a series of volunteer nights-and-weekends blog posts and conference talks. About three years in I bootstrapped the business as a consultancy, with one client lined up and ~15K USD in [available] savings. Scraped through the first year by dipping into credit cards, learned a lot about pricing and pipelines, and am doing pretty well now. Having money in the bank lets me do more volunteer work, like this analysis. :)
Jepsen makes money through consulting services (usually paid analysis work), training classes, speaking engagements (I charge at for-profit orgs and speak for free at nonprofit events), and Amazon Marketplace subscriptions.
Disclaimer: I work on Geode.
You can imagine GemFire/Gridgain as an apples-to-apples comparison. Both are "enterprise" in-memory data grids originally intended for managing data in low-latency OLTP applications which later added analytics/OLAP features. Geode/Ignite are the open source options for these two IMDGs and also a good apples-to-apples comparison. (Hazelcast also has enterprise/OSS verisons I would compare accordingly)
I can't speak to the current comparison between these systems, but I can compare them to SnappyData. SnappyData deeply integrates GemFire with Spark to bring high concurrency, high availability and mutability to Spark applications. In the world of combining Spark with a datastore over a connector (cassandra, hive, mysql, mongo etc) to enable "database-like" features in Spark, SnappyData has taken the next step of integration. In Snappy, the database (GemFire) and the Spark executors share the same block manager and VM so the systems no longer communicate over a "connector." This, along with our database optimizations, provides the best performance for Spark applciations in what I like to call the "Spark Database Ecosystem."
As such, comparing SnappyData to GemFire/Hazelcast/Gridgain does not make much sense unless you are trying to use Spark in conjunction with these systems. In that case, the main difference I would point out is that SnappyData will necessarily perform better as any of them would need to use a connector to interact with Spark. The better comparison would be between SnappyData and Ignite, as Ignite contains a direct Spark abstraction called "IgniteRDD." That said, the majority of the comparisons/benchmarks we've run have been against MemSQL+Spark and Cassandra+Spark, so I don't have much to say about Ignite vs SnappyData.
User manigandham mentions SnappyData's Approximate Query Processing features (called Synopses Data Engine) which is unique within this space, but a discussion of which would take this too far afield.
However, I'd really like to credit the author on the design of the site, a pleasure to read.
What fascinates me about the now-known-to-be-untrue claims in the Hazelcast documentation is that the claims were made despite the developers having know way of knowing if they were true because they’d not conducted these kinds of tests themselves.
Documentation that reflects what the developers wish were true rather than what is actually true is not a new phenomenon, but is potentially fatal for this kind of software.