Jepsen: Etcd 3.4.3
jepsen.io
jepsen.io
>We would love to see someone in the etcd community integrate the etcd Jepsen tests directly into the existing etcd release pipeline.
I consider this to a be an issue of higher priority than any of the bugs they just found, because this will ensure preventable bugs don't crop up in the future. It's shocking to me that Jepson goes through all the effort and than very few projects build a permanent pipeline for it. It's debatable these bugs would've existed if a Jepson pipeline had been consistently in use from the 0.4.x days. I'm sure it's no simple task, but neither is a lot of the existing testing infrastructure for etcd.
I don't think it would have helped: the Jepsen tests I wrote in 2014 only checked single-key gets, puts, and CaS operations; the problems we found in this report were in watches and locks.
[0]: https://www.cockroachlabs.com/blog/diy-jepsen-testing-cockro...
It would be nice to get follow-ups too, personally for Cassandra, Kafka, couch, and elastic.
I miss the unique and informative style of the old reports.
I had worked on an alternative etcd impl and had to workaround this assumption as well. It is technically documented in the proto[0], and numeric 0 is of course "unset" or "default" in proto3 land.
One thing I would like to see tested is nested transactions where one txn child mutates something then the second sibling txn child uses that something. I've found that implementation is lacking.
0 - https://github.com/etcd-io/etcd/blob/53f15caf73b9285d6043009...
https://aphyr.com/posts/284-call-me-maybe-mongodb ("Mostly, it just throws those writes away entirely: no rollback files, no nothing. I don’t really know why.")
https://aphyr.com/posts/317-call-me-maybe-elasticsearch ("When the cluster comes back together, one primary blithely overwrites the other’s state.")
etc etc
Whoops, looks suspiciously like someone tested the revision integer for truthiness to see if something was passed.
But it can require a bit more more thought in your protocol design. So for example in this case, the options would be
- design things such that 0 is ok to mean "most current" (so start revisions at 1) (this is hard after the fact, but if you know from day 0 that missing values for int types will be 0, you can design everything to start at 1) (Edit: maybe this is how etcd works?)
- explicitly break out revision into a message type (so you can notice if it's not provided)
- use something like "-1" to mean "now", so that 0 isn't overloaded
etc...
(You could argue maybe the right call for proto3 was to have a flag on a field saying if you want to be able to notice if it was provided or not. Best of all worlds, at cost of a bit of complexity.)
But now you can't figure this out at all without adding another, boolean field, _and setting it separately_, which I'm pretty sure nobody is going to do unless they really have to, leading to the type of issue we're seeing here.
Google themselves provide https://github.com/protocolbuffers/protobuf/blob/master/src/... to deal with this situation.
Quoted from the documentation:
> Wrappers for primitive (non-message) types.
> These types are useful for places where we need to distinguish between the absence of a primitive typed field and its default value.
It should probably be advertised more, as we've experienced that default values of optional fields are a surprising feature for smart developers who are new to protobufs. Maybe it's seen as a wart in the design? Getting rid of Null is hard.
I kinda wish they went with the approach that all fields are required, unless they are explicitly declared as optional. This is how Rust does it, and people seem to like it.
It’s the distinction between an optional type and an optional value (default value). With optional values but no optional types, you can’t be certain about the caller’s intentions. It’s a distinction that’s subtle but important, therefore a “gotcha”.
Getting rid of null is a noble idea, because of the headaches that null tends to induce. Optional types (like Rust’s) is a neat way to get the behavior of null without the value of null. Proto3 doesn’t have null, it’s replaced with arcane wrapping that’s arguably less straightforward.
Please let me know if I’m talking past you, it isn’t intentional, just late in the day. :)
Yes, I would love if there were set/unset bits available. Even if it was awkward.
and sometimes defining your own, e.g. nullable lists.
It's not elegant, but it is simple.
I do no work at all in this area, but i love these reports. They're examples of well-written, clear, "engineer-mind" reports that we would all do well to emulate.
Can etcd be also used as a general distributed database like FoundationDB or ScyllaDB? If so how does it compare to those other optiions?
https://etcd.io/docs/v3.3.12/dev-guide/limit/ https://github.com/etcd-io/etcd/blob/master/Documentation/op...
TiKV, a distributed kv store that can store many terabytes of data, actually uses etcd internally for metadata storage.
There are different databases like CockroachDB or Dgraph that use etcd's raft libraries but build more application focused (SQL, graphql, etc) APIs[2].
[1]: https://github.com/etcd-io/etcd/blob/master/Documentation/in...
[2]: https://github.com/etcd-io/etcd/tree/master/raft#notable-use...
edit: seems like my recollection of the mongodb one is the 2.4.x one [1] and the later ones are much better. Also mongo included the Jepsen test suite into their CI.
AWS doesn't guarantee that failures within an AZ are independent of each other, so it's not clear how you would estimate what availability you'd gain with this. Losing everything in an AZ + 1 instance sounds like a very unusual and specific scenario to design for.
An AZ going down doesn’t make all other hardware reliable, and equally a machine going down from a cluster doesn’t mean that all AZs are going to be reliable.
Many products have uptime requirements above what Amazon can provide at the AZ or machine level.