Jepsen: Radix DLT 1.0-Beta.35.1
jepsen.io
jepsen.io
> When asked, RDX Works executives informed Jepsen that blockchain/DLT readers would normally understand present-tense English statements like these to be referring to potential future behavior, rather than the present.
> Jepsen is no stranger to ambitious claims, and aims to understand, analyze, and report on systems as they presently behave—in the context of their documentation, marketing, and community understanding. Jepsen encourages vendors and community members alike to take care with this sort of linguistic ambiguity.
(sorry for relatively low effort comment.)
Making marketing material truer will lower the amount of people trapped into the vicious cycle of repeating false claims for the sake of propping up a risky investment, and is good for everyone involved long term.
But short-term, making outrageously optimistic statements will create a large amount of proponents which is what probably financed this report in the first place.
The difficulty is in ramping up truth without destroying too much the franchisees who are the ones who put up the actual money but valued their investment in light of a false/exaggerated statement.
Let's appreciate this for what it is. Blockchains are, at their heart, a type of database (or at the very least a ledger which can be the foundation on which some subset of database semantics can be layered). Performance and reliability are empirical claims which can be tested empirically, using the kind of methodology that Jepsen has been innovating for many years. It is very much to the credit of RDX Works that they subjected their product to this type of testing. I'm not saying anything about the way the use the test results in their blog post and marketing materials, though.
What I'd like to see going forward is that it's routine for blockchain-based databases to be tested the same way as real databases, based on actual shipping product rather than speculative goals. Whether you think this would be validating or devastating reveals quite a bit about your preconceptions, but either way would be a win for truth and progress.
For example, because of the blockchain trilemma, increasing transaction speed isn't always a good thing since it could sacrifice decentralization.
As someone who's been in the blockchain space for 8 years I've seen numerous scamcoins become popular because they lure people in with claims of high TPS. High TPS isn't actually that hard of a problem. High TPS with decentralization and security IS hard.
Case-in-point, blockchains like Bitcoin and Ethereum limit TPS on purpose specifically because they care about decentralization. After all, that's the whole point of a blockchain.
I would encourage you and the people downvoting to read up on a few things. 1) The blockchain trilemma. 2) Why Bitcoin and Ethereum limit transaction speed and block size. 3) Layer 2 Rollups (zkRollups) that offer an actual solution to transaction speeds without sacrificing decentralization.
In summary, my point is that judging a blockchain by the standards of a traditional database is like judging a traditional database by the standards of a filesystem. Sure they both store data and yes a filesystem might store data faster but there are other constraints that make them fundamentally different (like ACID transactions).
We understand crypto very well: the goalpost is always changing for a solution looking for a problem. Yday you were probably telling everyone your investment was the best because it could handle a million per second, today you tell us we're idiot: TPS don't matter. So what matters, what is this thing good for ? It's no better than Ethereum and at least Ethereum can launch a tree of related scams so it has half a utility.
That certainly does not inspire any confidence.
> Jepsen respectfully declines to do so.
Thank you for sticking to that.
In private DBs, reads from the DB node are considered transactions and need to follow the same rules as writes. But on public blockchains(ledgers) only state manipulation is what matters. For example, Metamask obtaining an address balance would be a transaction, but no one calls it that way because it doesn't modify the state.
/s
In my experience it is not. Rather, aggressive fsyncing/O_DIRECT usage are common. The rationale for this is usually partition risk: better to durably log a write before propagating it than to potentially fail in propagating it and then be left in a position of having to either reactively fsync or hope that automatic flush-to-disk will persist your unexpectedly-sole possession of that update.
That's quite a stretch. The report states explicitly a) the usual proviso that they can only prove the presence of bugs, not their absence, but more pertinently b) that their methodology is more usually applied to lower-latency databases, with the implication that they are less confident of their conclusions in this new regime:
> Radix’s low throughput and high latency may have masked safety violations. In particular, our tests required several hours to reproduce e.g. aborted read (#13).
Note also that they didn't even attempt to test what happens in the presence of malicious nodes!
One of your recent-ish blog article for instance, that contrast (commercial vs technical) with the Jepsen report: https://www.radstakes.com/post/airdrops-incoming-radstakes-p...
And your statement of the 2% fee you seem to take on every staking: https://www.radstakes.com/post/radstakes-year-end-report
This report will not help your business, but you should pressure RDX Works into doing a new one on the next version rather than convince us we didn't read properly some quite shocking things in the first one :s
Admit it, they exaggerated and removed the statement to instead put it in a timeline now saying: Unlimited scalability and composability to carry DeFi into the global mainstream future with sharded, linearly scalable Cerberus consensus.
A much more reasonable statement /s
You can see a version here of the statement: https://web.archive.org/web/20210711084824/http://www.radixd...
1.4m TPS on a DLT : Radix' last consensus algorithm 'Tempo' publicly achieved 1.4m TPS in 2018, the current world record. The new algorithm 'Cerberus' is theoretically infinitely scalable
For me that means they pretended to have this throughput in 2018, before this "Cerberus" miracle that is infinite. So one could repeat on twitter at every opportunity that Radix is 1.4m TPS, world record, and "later" be "infinite" with Cerberus.
As they say in https://www.radixdlt.com/post/replaying-bitcoin, use the #1MTPS on Twitter to share your opinions. lol.
Aphyr is the best. No matter your fave distributed tech and all it's CAP-don't-matter stuff, he shows that... CAP does very much matter and it's really hard.
Cassandra? Kafka? MongoDB (bwahahahah)? It's all got edge cases.
He should be getting paid a million bucks a year by various auditing/accounting firms and the FTC/SEC for validating crypto claims. It would be a massive public service.
I like that its being treated like a database, and that the safety/correctness of any blockchain has to be viewed from a distributed database standpoint.
From the very first paragraph:
>"This work was funded by Radix Tokens (Jersey) Limited, and conducted in collaboration with RDX Works Ltd ..."
The post also links to RDX Works Ltd's blog post on this collaboration. Also in the first paragraph.
Basically an astute reader of Jepsen testing may deduce "I need to use a single node system" which is one option but without the availability characteristics modern users usually want
That's what a successful test looks like. Sure, they found correctness issues with the lock api, but they were able to lend some confidence to etcd's core api. There might be bugs, but jepsen didn't observe any.
Sure, that's not proof that systems don't have tradeoffs, but it's proof that jepsen doesn't literally always find consistency issues or other core problems. Said another way, Jepsen is testing for correctness issues in systems. They either find some or they don't. These results are interesting, even if they are not a panacea.
There are other tradeoffs, but for most systems, correctness is important enough it merits consideration on its own right.
I don't think readers of jepsen misunderstand what's being tested or what it means, nor do they misunderstand that there are other tradeoffs to consider, so I think your comment is off the mark.
By analogy, if we were reading a post about "I load-tested this bridge which claims to support 50 tons of weight, and it broke at 5 tons", no one would be saying "yeah, but there are tradeoffs for bridges. This post just makes it clear that there's tradeoffs. If you make the bridge stronger, it would be more expensive, and to imply there's not tradeoffs is misleading. An astute reader might deduce that they should never drive on bridges again".
I don't think that would be a reasonable interpretation of such a post, nor do I think the interpretation you espouse here portrays an accurate sentiment.
PS: There is also another trend, that is System claiming they fixed the issues Jepsen finds, without submitting themselves again for analysis...but I digress now...
Here's an analogy that may help me communicate how I feel since I realize my message is not landing: Let's say I'm buying a condo in the San Francisco Bay area. And let's say the building that I'm looking to buy in advertisers that it is historic but seismically retrofitted. Then say Jepsen-earthquake-test comes in and shakes the ground beneath the building and shows that the building indeed collapses with enough of an earthquake: would that or would that not be enough information for me to decide whether or not the seismically retrofitted building is good enough for my needs? There's a lot of ambiguity in answer that question.
We've been doing distributed systems for over 50 years now. There's nothing ambiguous in either the language Jepsen uses or the claims he is examining. The crypto shills chose to pretend the language is ambiguous and invent their own definitions on purpose.
Including the ridiculous "nah, everyone understands that when we speak in present tense it means we mean some unspecified point in a nebulous future".
When he demonstrated that Riak was dropping 30-70% of writes, even with the strongest consistency settings, or that Mongo had multiple scenarios of data loss, we are not talking about the subtleties of the English idiom
A Jepsen style test optimized not for bending to show where things break but instead for showing likely real world style situations with an eval of which are most likely to arise would be far more valuable for people
- If we're sticking with your example further up-thread, you'd buy a house that was advertised as "can withstand a 4.0 magnitude earthquake", that then failed when subjected to a 4.0 magnitude earthquake.
- 4.0 magnitude earthquakes happen all of the time [0].
More or less what I'm saying is, engineers generally don't think DBs lose data, and when they start coming up with ways that might happen (node failure, network failure, clock desync), distributed DBs assure them with algorithms and configuration knobs. Aphyr puts those assurances to the test, which is so, so valuable to us all.
It's also worth saying that this space is pretty technically complicated. All the DB engineers I know use some form of Jepsen-style testing (or Jepsen itself) because it's amazingly great.
> what most users need to understand about said systems
The target audience of these reports are the system builders and the software engineers building services on top of them; not end-users consuming higher level services.
The real world is significantly more "contrived" than anything Jepsen can come up with.
Transactions timing out, Nodes losing transactions even after they've been acknowledged etc.? None of this is contrived.
Rather, it would be "Builders claim their buildings are earthquake proof, and jepson was able to show a subset collapse in earthquakes. The rest may or may not be earthquake proof".
That's still very valuable. It's really valuable to know when something is wrong. It would be more valuable to know that something is definitely right (correct / "earthquake proof"), but jepsen cannot prove that.
FWIW, I think you're raising really good questions in this thread. Qualitative safety is highly contextual, depending on fault model, significance of individual operations, concurrency, throughput, latency demands, operational characteristics, data volume, etc etc., and I try to touch on that in the "Toward a Culture of Safety" section in this report. Hopefully that resonated for you.
In my opinion the Rethink DB 2.1.5 report went fairly well for that project. If the claims are aligned with the reality of the product, it's clear the Jepsen report will highlight that.