Call Me Maybe: Chronos
aphyr.com
aphyr.com
The general take-away from most of them is -- do not believe the vendors of distributed systems. Even systems that were meant to solve consensus (etcd, consul) failed:
https://aphyr.com/posts/316-call-me-maybe-etcd-and-consul
Elasticsearch didn't do well, MongoDB, Kafka, etc.
The only ones that did ok that I remember were: Zookeeper, Riak (with siblings), and maybe Cassandra.
If the network recovers while the last WAL log your replica read still exists on the server, your replica probably get to catch up, as long as your replica can read the logs as fast as they're being rotated out.
Postgres doesn't really try much else in the realm of distributed systems, so there's not much else you can test.
Skytools would be interesting, though, I've never felt particularly comfortable with replication schemes that trust an external process to ship rows and other changes, compared to the binary replication that will absolutely fail if you don't have every bit.
Rather than wal_keep_segments, you may (should) use replication slots (http://www.postgresql.org/docs/9.4/static/warm-standby.html#...), by setting the primary_slot_name parameter in recovery.conf. With replication slots, WAL files will be kept for every replica, regardless of how far the replica lags behind. So you can very directly control the ability to recover from network partitions (although in exchange you will need to monitor the disk usage for pg_xlog at the master and, if necessary, forcibly drop the slot).
In practice, the last time I ran pg, I was fond of wal-e which, esp if you're on EC2, is handy because you can store basically all of the WAL segments ever, with snapshots, in S3, and you can bring new replicas online without a read load on any of the existing nodes. It will also bring a replica online that's been shut off for days or weeks for service.
This is really just an S3 version implementation of the wal archiver, which was originally designed afaict for storing infinite history on NFS. Come to think of it, the new place has NFS. rubs chin
Anyway, thanks for the pointer! Def something I would look into in the future. Maybe it would be fun to write a jepsen test for recovering pg replicas.
As some comment below, the reason for this observed behavior is due to the fact that the server, together with the client, form a kind of distributed system. It is not classified as a failure to perform a write and fail to ACK that to the client. The reverse (not performing an ACKed write) is a failure. However, this suggests that every non ACKed write operation should be retried with care, because it may have gone through. It's better if you can use idempotency semantics, like 9.5's upcoming http://www.postgresql.org/docs/9.5/static/sql-insert.html#SQ... (ON CONFLICT, aka UPSERT).
It's worthwhile to note that this is a fairly general problem in many, many kinds of systems.
If that's actually a problem, you can use 2PC. If you didn't get back an ack for the PREPARE TRANSACTION it failed (and you better abort it explicitly until you get an error that it doesnt exist!). Once the PREPARE succeeded you send COMMIT PREPARED until you suceed or get an error that the repared transaction doesn't exist anymore (in which case you suceeded but didn't get the memo about it).
Some are very thankful for the analysis, others don't seem to appreciate how important correctness is in this type of distributed systems. I find it's a good indicator of the maturity of the teams working on them, and how much I'm willing to trust each system.
I hate to nitpick but I'm generally curious. How do you know it's a good indicator? It can definitely be _an_ indicator but how do you know it was good? Did you have a follow up experience that validated what the indicator indicated?
One way is to note whether or not the response to the report dismisses it or hand-waves it away. Another response is to neglect the report for 18 months with no reply.
The tests that aphyr conducts are brutal, which you can see by casual inspection of his Call Me Maybe site or his talks. If the response of a software team tries to divert his attention or your attention from the issue, it does not leave a good impression.
The criticism is that last-write-wins is the default, and that's a fair criticism--that's not an honest default for a robust eventually consistent system like Riak. My guess is they default to LWW to ease onboarding of new users. Everyone is worried about rivaling MongoDB's out-of-the-box experience so you don't scare people away...
TBH, though, if you're using a distributed, eventually consistent storage system without understanding the consequences, you're probably using the wrong tool.
Also, an old blog entry from yammer engg. blog: http://old.eng.yammer.com/call-me-maybe-hbase/
"Like I suggested before, you should stop by our offices @ 88 Stevenson for lunch some day, and you can chat with our best and brightest about this (and other things) until the cows come home. Then you can save yourself a lot of back and forth, and nerdraging about computers."
Aphyr spending his time finding bugs and unspecified behavior in their software is dismissed as "nerdraging". And instead of discussing it openly, he should spend a day and go to "88 Stevenson", wherever that is...
I'm going to assume he means Stevenson Avenue, Polmont, Falkirk, Stirlingshire, United Kingdom -- since that's the first Google result I got for "Stevenson" as an address.
The point was that it's not the kind of reply you should make in a public issue tracker. It comes off as: "We're only interested in talking to people who are in San Francisco and are cool enough to hang out with us."
Except he's in SF.
Aphyr tests distributed systems to destruction. He seems to be really good at it. You've had an account here for nearly three years, and you write distributed software; it's fairly shocking to me that you're not familiar with his work. You should browse https://aphyr.com/tags/jepsen for a bit, to get a feel for what he does.
As an anecdotal aside, I tend to use aphyr's interactions with a development team as a bit of an indicator about the team's quality. A team that acknowledges the problems that he's finding and hurries to address them is probably a good one. A team that outright denies that the problems exist is a pretty poor one. Your responses to #513 (especially the nerd raging bit) aren't encouraging.
I read it less as that and more "you are embarrassing us by discussing our product's flaws in public. please bring this behind closed doors"
There's an 88 Stevenson in SoMa in San Francisco, so that's pretty canonical.
Also, you have to give Air credit for being civil, facilitating discussion, and taking responsibility for the issues.
Only for somebody living there. As I point out in the parallel post (which is possibly too gray to be seen) Google is almost certainly really giving Stevenson in UK as the answer for people who search from there.
Perfect point!
Of course, somebody can't imagine the people exist who aren't in the same city as he is.
https://mesosphere.com/contact/
"88 Stevenson St, San Francisco, CA 94105"
And it's often forgotten that Google is not the same for everybody, customizing their result based on their own heuristics of who is asking and from where he's asking.
But maybe the author of that post actually knows the person he communicates to is actually in the same city? I don't know.
I like face to face discussions. It would have been nice to talk to him about what he's building, and what he's trying to accomplish. Google hangouts would have worked, too.
I'm pretty sure he is merely (!) building tests that demonstrate defects in distributed systems software.
I wonder if you think that makes the issues he files against your software somehow less valid.
It seems the context is "if it isn't tested, it is broken". One might surmise that the kind of testing he did on your software is a kind that you did't do.
I would disagree with face-to-face being a useful way to confront a painful bug report.
Honestly, that is the money shot. Its not worth running an off the shelf cluster system that isn't self-healing from a network partition and/or a node loss->node restored cycle.
"This isn’t the end of the world–it does illustrate the fragility of a system with three distinct quorums, all of which must be available and connected to one another, but there will always be certain classes of network failure that can break a distributed scheduler."
As we build increasingly complex distributed from what are essentially modular distributed systems themselves, this is something to be aware of. The failure characteristics of all the underlying systems will bubble up in surprising ways.
A network partition like Jepsen tests for is a once-every-other-month problem due to network maintenance, hardware failures, etc. So yeah, it is the end of the world for me. I like not having to wake up in the middle of the night every few months.
Note that I briefly evaluated Chronos, but tossed it aside as a toy when realizing it didn't support any form of constraints. Aurora is a very nice scheduler for HA / distributed cron (with constraints!) and long running services.
It would be fantastic if Kyle would test out Aurora, if not only to break it so that upstream fixes it. He is generally successful in breaking ALL THE THINGS.
If you can replace ZK with a native Mesos construct, seems like it allows you to remove ZK entirely. I meant to say "optionally allow removing ZK as a dependency" in the original post. You're totally correct in that regard.
The annoying problem that I also encountered with Chronos was the lack of error reporting from the REST API. Invalid jobs would just return "400 Bad Request" with no error message. The error sometimes wouldn't even be reported in the logs.
If vendors fix things after the test, there are usually 2 types of fixes -- docs or code. Docs mean tell people how system really behaves and how data might be corrupted or actually fix the problem if possible.
So if Chronos just says in bold red letters on the front -- "You'll lose data in a partition" or "Our system is neither C or A if you use these options" that's ok too. The users can then at least make an informed choice.
I've tried out Aurora (and helped improve the setup docs a small bit) and found it quite nice. If you're going to try it out use vagrant, and to deploy it in production I suggest working off the Vagrant setup: https://github.com/apache/aurora/blob/master/examples/vagran...
> Instrument your jobs to identify whether they ran or not
Prometheus.io let's you alert on when the job last succeeded, which is generally what you want. http://prometheus.io/docs/instrumenting/pushing/ has a full example for Java.
"really wish I had more time to dig into chronos because it has so many fascinating failure modes but I've got 2 other systems before talks!"
I'm not saying the idea is bad, but comparing the developer community of Marathon with Chronos you see some obvious differences in commits, developers, issues (and responses), releases etc.
We've been using it in conjunction with Marathon and Mesos, and my impression is that it's a now half-dead project riddled with bugs (especially in the web UI) and I'm unsure whether we should invest more resources and infrastructure around this project.
Aurora does look interesting. But Apache doesn't have a much better reputation, and I'm not particularly keen on going away from Marathon. Any DevOps engineers with container/mesos infrastructures wanting to chime in?
Disclaimer: I work for Pivotal (in Labs), which donates the largest chunk of engineering effort to Cloud Foundry, including the Lattice team.
The differences (and advantages) between Lattice and Mesos could've been made clearer too I guess, although your docs and FAQs are quite nice.
And it doesn't require you to build your own snowflake PaaS, which you'll be married to forever at your own expense.
I've worked on CF so I'm biased in its favour. But my actual paying job is helping clients to deliver user value. Tinkering around with various tools is fun, but it's not delivering user value.
Typing
cf push app-name
Is.And it works. It just plain old works. That makes my life 100x easier.
Currently in use at HubSpot, OpenTable, Groupon (last I heard) and EverTrue.
I can say that Mesosphere is 100% committed to both Marathon and Chronos. You're right that Marathon gets more attention for the moment, but we have grand plans for both.
A strength of Mesos is that it's an open platform. If Chronos isn't right for you, there are solid alternatives. I'd appreciate if you raise issues against Chronos though so we can improve it : )
Excuse the plug, but if you'd like to help us improve these widely-used frameworks we'd love to hear from you: https://mesosphere.com/careers/
Wat? Why is that necessary to mention? Every process is fragile and needs to be supervised!
hahaha, well naming things is one of the hardest problems in computer science after all.
http://doc.akka.io/japi/akka/2.3.9/akka/actor/dungeon/Childr...
"We need more slaves!"
It's a bit of nomenclature that would never be used if people weren't mimicking choices from decades ago.What name to use other than slave might be a bikeshed issue, but the bikeshed is currently painted with a confederate flag. ;)
Incidentally, filicide is also ancient and ongoing.
Slavery is a bad practice and I would definitely not work with someone who felt otherwise, but other than that, your argument is just worthless escapism. Have fun talking to HR in the future.
I will have fun talking to HR, thank you.
If you have a DB "slave", the "master" really doesn't control it. The master doesn't enforce it's functionality. The master won't typically supervise it, the master doesn't own it. The master didn't buy it, and there's really no economics in the situation at all.
When you view systems in terms of autonomy and commitments (ie, promise theory), the metaphor of slavery becomes grossly inaccurate when trying to describe what a system is really doing.
I'd call it a replica instead.
On a master/slave bus, nothing works if the master is shut down. The master literally supervises everything by arbitrating communication.
In a replica system like elastic search, failure of one node will usually just cause some other node to take over its task.
If you want to get the word changed do it out of respect for those living in slavery today, and not because some of the most privileged people on earth (first world software developers regardless their minority) could get their feelings hurt.
I got the impression that the people most vocal about these terms don't really care about the issues. They just want the thrill of changing the world, without actually doing or changing anything.
If we ban all the bad words from our vocabulary, it only serves to pretend the bad things they name don't exist.