HNHacker News
TopNewBestAskShowJobs

aphyr

9,565 karma · joined January 15, 2010

https://aphyr.com

Hacker, physics geek, clojurer, rubyist, aikidoist, photographer.

submissionscomments
aphyr··on When flat rate movers won't answer your calls
This is a whole story unto itself. I (and a bunch of other small-site operators) actually got the chance to ask Ofcom directly about how small a site would have to be to declare itself out-of-scope of part 3/5 services, and in short: the law doesn't specify, Ofcom declines to say, and it could be as small as "a one-man band". Ofcom has indicated that they're interested in pursuing aggressive enforcement against small service providers (read: web sites). I wound up concluding that the forum I help run could be defensibly out-of-scope, but aphyr.com gets significantly more traffic, and I'm trying to limit risks.

I think in total I spent something like 80 hours reading the legislation, working through thousands of pages of Ofcom guidance, and corresponding with Ofcom trying to figure out whether and how compliance was feasible.

https://blog.woof.group/announcements/updates-on-the-osa

aphyr··on When flat rate movers won't answer your calls
I made a lot of my furniture--it's hard to replace.
aphyr··on Jepsen: Amazon RDS for PostgreSQL 17.4
Thanks to you both! I've updated the article to discuss this, and we've got an update on the AWS blog too. :-)

https://jepsen.io/analyses/amazon-rds-for-postgresql-17.4

aphyr··on Jepsen: Amazon RDS for PostgreSQL 17.4
Thank you matashii--this would definitely explain it. I've also received another email suggesting this anomaly is due to the difference in commit/visibility order between primary and secondary. Is there by chance a writeup of this available anywhere that I can link to? It looks like https://postgrespro.com/list/thread-id/1827129 miiight be related, but I'm not certain. If so, I'd like to update the report.

My email is aphyr@jepsen.io, if you'd like to drop me a line. :-)

aphyr··on Jepsen: Amazon RDS for PostgreSQL 17.4
Yeah, that's right. It may be that the (apparent) order of transactions differs between primary and secondary.
aphyr··on Jepsen: Amazon RDS for PostgreSQL 17.4
This is a very good question! I do not understand AWS's replication architecture well enough to reimplement it with standard Postgres yet. This behavior doesn't happen in single-node Postgres, as far as I can tell, but it might happen in some replication setups!

I also understand there are lots of ways to do Postgres replication in general, with varying results. For instance, here's Bin Wang's report on Patroni: https://www.binwang.me/2024-12-02-PostgreSQL-High-Availabili...

aphyr··on Jepsen: Amazon RDS for PostgreSQL 17.4
I have actually been working on this (very slowly, in occasional nights and weekends!) Peter Alvaro and I reported on a safety issue in RDS for MySQL here too: https://jepsen.io/analyses/mysql-8.0.34#fractured-read-like-...
aphyr··on Jepsen: Amazon RDS for PostgreSQL 17.4
Folks on HN are often upset with the titles of Jepsen reports, so perhaps a little more context is in order. Jepsen reports are usually the product of a long collaboration with a client. Clients often have strong feelings about how the report is titled--is it too harsh on the system, or too favorable? Does it capture the most meaningful of the dozen-odd issues we found? Is it fair, in the sense that Jepsen aims to be an honest broker of database safety findings? How will it be interpreted in ten years when people link to it routinely, but the findings no longer apply to recent versions? The resulting discussions can be, ah, vigorous.

The way I've threaded this needle, after several frustrating attempts, is to have a policy of titling all reports "Jepsen: <system> <version>". HN is of course welcome to choose their own link text if they prefer a more descriptive, or colorful, phrase. :-)

aphyr··on Jepsen: Amazon RDS for PostgreSQL 17.4
This isn't just stale data, in the sense of "a point-in-time consistent snapshot which does not reflect some recent transactions". I think what's going on here is that a read-only transaction against a secondary can observe some transaction T, but also miss transactions which must have logically executed before T.
aphyr··on Did we miss P In CAP? Partial Progress Conjecture under Asynchrony
The fundamental definition of consensus already requires that some proposal will eventually (given sufficient communicating, non-faulty, non-malicious nodes) win, so that doesn't seem particularly novel: https://lamport.azurewebsites.net/pubs/lower-bound.pdf

The core problem here that it's impossible for any replica to tell whether its speculatively-executed operations were legal or not; it might have to admit "sorry, I lied to you", and back up at any point. That seems to be what they mean by "partial progress". You can guess that things might happen, but you will often be wrong.

This idea's been around for a while--Fekete et al outlined speculative execution at local replicas with eventual convergence on a Serializable history in their 1996 PODC paper "Eventually Serializable Data Services": https://groups.csail.mit.edu/tds/papers/Lynch/podc96-esds.pd.... These systems have essentially two tiers of operations: weak operations, which the system might re-order (which could cause totally different results, including turning successes to failures and vice-versa!), and strong operations, which are Serializable. I assume Cassandra does the same. Assuming it's correct, the strong operations work like any other consensus system: they must block (or abort) when there isn't sufficient communication with other replicas. The weak ones might give arbitrarily weird results. In CAP's terms, the strong operations are CP, the weak ones are AP.

You see similar dynamics at play in probabilistic consensus systems, like Bitcoin, by the way. In Bitcoin technically all operations are weak ones, but the probability of non-Serializable outcomes should decrease quickly over time.

Having a consensus system that merges conflicting proposals is a nice idea, but I don't think Cassandra is novel here either. I don't have a citation handy, but I recall a conversation with Heidi Howard at HPTS (maybe 2017?) where she explained that one of the advantages of leaderless Paxos is that when you're building a replicated state machine, you can treat what would normally be conflicting proposals from multiple replicas as sets of transitions. Instead of rejecting all but one proposal, you can union them, then apply them to the state machine in some (deterministic) order--say, by lexicographically sorting the proposals.

aphyr··on Jepsen: Bufstream 0.1
I (and apparently the Confluent docs?) may be wrong about this. I've added an update to the report.
aphyr··on Jepsen: Bufstream 0.1
Apple is positively swimming in money! They could pay me! (Hi, Apple ;-))
aphyr··on Jepsen: Bufstream 0.1
Kafka actually does call these transactions! However (and this is a loooong discussion I can't really dig into right now) there's sort of two ways to look at "exactly once". One is in the sense that DB transactions are "exactly once"; a transaction's effects shouldn't be duplicated or lost. But in another sense "exactly once" is a sort of dataflow graph property that relates messages across topic-partitions. That's a little more akin to ACID "consistency".

You can use transactions to get to that dataflow property, in the same sort of way that Serializable transaction systems guarantee certain kinds of domain-level consistency. For example, Serializability guarantees that any invariant preserved by a set of transactions, considered purely in isolation, is also preserved by concurrent histories of those transactions. I think you can argue Kafka intends to reach "exactly-once semantics" through transactions in that way.

aphyr··on Jepsen: Bufstream 0.1
I'm not really sure how to answer this question, but even a few chapters worth of clear prose would go a long way. We lay out a bunch of questions in the discussion section that would be really helpful in firming up intended txn semantics.
aphyr··on Jepsen: Bufstream 0.1
Both are true, but we use "transactions" for clarity, since the semantics of consumers outside transactions is even murkier. Every read in this workload takes place in the context of a transaction, and goes through the transactional offset commit path.
aphyr··on Jepsen: Bufstream 0.1
I don't think so, but I've said a lot about databases in the last fifteen years haha.

Sometimes I look at what people say about FDB and it feels like... folks are putting words in my mouth that I don't recognize. I was very impressed by a short phone conversation with their engineers ~12 years ago. That's good, but that's not, like, a substantive experimental evaluation. That's "I focus my unpaid efforts on databases which seem more likely to yield fun, interesting results".

aphyr··on Jepsen: Bufstream 0.1
Nope. You'll find a full list of analyses here: https://jepsen.io/analyses
aphyr··on Jepsen: Bufstream 0.1
I'm not sure. I've worked on a few projects now which employed simulation testing and passed, only to discover serious bugs using Jepsen. State space exploration and oracle design are hard problems, and I'm not convinced there's a single, ideal path for DB testing that subsumes all others. I prefer more of a "complete breakfast" approach.

On another axis: Jepsen isn't "trying to get there [to FDB's testing]" because Jepsen and FDB's tests are solving different problems. Jepsen exists to test arbitrary, third-party databases without their cooperation, or even access to the source. FoundationDB's test suite is designed to test FoundationDB, and they have political and engineering buy-in to design the database from the ground up to cooperate with a deterministic (and, I suspect, protocol-aware) simulation framework.

To some extent Antithesis may be able to bridge the gap by rendering arbitrary distributed binaries deterministic. Something I'd like to explore!

aphyr··on Jepsen: Bufstream 0.1
You'll find lots about the Jepsen analysis process here: https://jepsen.io/services/analysis
aphyr··on Jepsen: Bufstream 0.1
I have not yet, though you're not the first to ask. Some folks have suggested it might be... how do you say... fun? :-)
aphyr··on Jepsen: Bufstream 0.1
I would love to do a Kafka analysis. :-)
aphyr··on Jepsen: Bufstream 0.1
Here's a good article from New Relic on the problem, if you'd like more detail: https://newrelic.com/blog/best-practices/kafka-consumer-conf...

Or here, you can reproduce it yourself using the Bufstream or Redpanda/Kafka test suite. Here's a real quick run I just dashed off. You can watch it skip over writes: https://gist.github.com/aphyr/1af2c4eef9aacde7f08f1582304908...

lein run test --enable-auto-commit --bin bufstream-0.1.3-rc.12 --time-limit 30 --txn --final-time-limit 1/10000

aphyr··on Jepsen: Bufstream 0.1
Ack, pardon me. That should be fixed now!
aphyr··on Jepsen: Bufstream 0.1
It is a little surprising, and I agree, the docs here are not doing a particularly good job of explaining it. It might help to ask: if you don't explicitly commit, how does Kafka know when you've processed the messages it gave you? It doesn't! It assumes any message it hands you is instantaneously processed.

Auto-commit is a bit like handing someone an ice cream cone, then immediately walking away and assuming they ate it. Sometimes people drop their ice cream immediately after you hand it to them, and never get a bite.

aphyr··on Jepsen: Jetcd 0.8.2
For this test, you can do it on pretty much any reasonable Linux machine. Longer histories can churn through more CPU and RAM--some of the more aggressive tests I ran for this work involved 20 GB heaps and 50 cores--but you can tune all that lower.
aphyr··on Jepsen: Jetcd 0.8.2
Yeah, that's a good way of phrasing it! :-)
aphyr··on Jepsen: Jetcd 0.8.2
> strict serializability doesn't imply idempotency

I think we're probably getting at the same thing, but I do want to clarify a bit. A Strict Serializable history, like a Serializable one, requires equivalence to a total order of transactions. That's clearly not true for etcd+jetcd: no possible order of transactions can allow (e.g.) a transaction to read from its own future. It's totally fine to submit non-idempotent transactions against a Serializable system: systems which actually provide Serializable will execute known-committed transactions exactly once.

Plenty of other databases pass this test; etcd+jetcd does not. This system is simply not Serializable.

aphyr··on Jepsen: Datomic Pro 1.0.7075
This is also good to hear! I'm not sure whether I'd call it a "footgun" per se--that's really an empirical question about how Datomic's users understand its model. I can say that as someone with some database experience and a few weeks of reading the Datomic docs, this issue actually "broke" several of the tests I wrote for Datomic. It was especially tricky because the transactions mostly worked as expected, but would occasionally "lose updates" or cause updates intended for one entity to wind up assigned to another.

Things looked fine in my manual testing, but when I ran the full test suite Elle kept catching what looked like serious Serializability violations. Took me quite a while to figure out I was holding the database wrong!

aphyr··on Jepsen: Datomic Pro 1.0.7075
This is good to hear! Nubank has also argued that in their extensive use of Datomic, this kind of issue doesn't really show up. They suggest custom transaction functions are infrequently written, not often composed, and don't usually perform the kind of precondition validation that would lead to this sort of mistake.
aphyr··on Jepsen: Datomic Pro 1.0.7075
Yeah, I think this is next to Zookeeper as one of the most positive Jepsen reports. :-)
← PreviousPage 2 of 26Next →