HNHacker News
TopNewBestAskShowJobs

andras_gerlits

32 karma · joined May 15, 2018

submissionscomments
andras_gerlits··on [dead]
(...) What this means is that we lose many efficiencies when we talk about configuration as different from data. The fact is, no matter how much we’re trying to separate the two, configuration is data. All major outages experienced by (generally) well-designed high-availability systems are because this truth is missed by almost our entire industry. (...)
andras_gerlits··on Jon Pretty wins in court against sexual harassment claims by Scala community
"I stand by spreading hearsay, even if a court happens to disagree"
andras_gerlits··on Using Promise Theory to solve the distributed consensus problem
I can only repeat what I told you earlier. Our distributed consistency model meets the SQL-standard's requirements for consistency and tolerates such outages. This is a fact.

CAP is a bad model for more reasons than the ones listed in that article. My favourite one is that it requires Linearizability, which nobody does in SQL. The disconnect when saying "SQL is not consistent" to me is just too much. CAP is based on a badly defined idea that comes from a presentation that was wrong in what it said.

That you need to tolerate outages of entire regions is a good argument to make in itself, there's no need to point at CAP. My answer to that is that as there's a way to define consistency in a way that allows for it to manage partition problems more gracefully, and that is the model we show. If you require communication to establish consistency and stream the changes associated with the specific timeslot at the same time, partition means that the global state will move on without the changes from the partitioned areas and that they will show up once the region rejoins the global network. While separated, they can still do (SQL-) consistent reads and writes for (intra-region) records they can modify.

andras_gerlits··on Using Promise Theory to solve the distributed consensus problem
CAP means many different things. If you took the time to read what I have to say about it, you would know that I'm saying that we're beating the requirements Brewer sets out in his original presentation, where he introduces the concept of the C-A-P tradeoff. He's clearly wrong in what he says in the presentation, which is what we say we're beating. We can say this, because we're meeting the requirements for "C" there (DBMS-consistency) and because we don't suffer the trade-offs mentioned there. In fact, our system can be both available and partition-tolerant with a definition of "C" that matches the ones laid out in the SQL-spec, as the reads are always local. The SQL-standard doesn't mandate time-related availability guarantees for writes.
andras_gerlits··on Using Promise Theory to solve the distributed consensus problem
There's a mental leap you need to make before it clicks into place and I can't do it for you. I have several people who understand what I say and why I say it, but I get that this isn't an easy step to make. If you read the essay, you understand how we turn time into data. You also understand how we can construct partial and global hierarchy of orders, how these are both strongly ordered and still move independently and how observers can construct their own consistent view from the data available to them. There's formal proof in the science-paper that we can do consistency all the way to snapshot.

To summarise: we have a consistent system that works via one-way streaming, via redundant channels. Best possible cache-invalidation. Reads are both consistent and locally available. Upper-bound time of writes is predictable and doesn't suffer from blocking. I'm not sure what other improvements anyone could want from such a system, this is the best such systems can ever be. These improvements go well beyond the limits CAP sets out and it basically makes the whole argument moot.

andras_gerlits··on Using Promise Theory to solve the distributed consensus problem
The reason you don't understand why my claims are different from being just another solution with clocks and mutexes is because you haven't actually engaged with the essence of it. I'll give you a hint: the consistency mechanism is decoupled from collision resolution and that makes the consistency ledger both deterministic and inherently non-blocking. I don't want to go into more specific questions around the implementation, but I can tell you with certainty that we provide much better promises around progressing global state and sustained write-latency than anything else and our client-centric consistency guarantees local-only reads.

But this is about the solution, not the science. The science is basically unlimited, its only real restriction is around our budget.

If you want to understand how the client-centric consistency mechanism takes care of these things, I write about it on medium and Twitter all the time.

But again, I don't feel like you owe me your time or attention

andras_gerlits··on Using Promise Theory to solve the distributed consensus problem
We need to clarify first if we're talking about CAP or real life. CAP requirements are absurd. To quote Dominick Tornow from here:

https://blog.dtornow.com/the-cap-theorem.-the-bad-the-bad-th...

"Note that Gilbert and Lynch’s definition requires any non-failed nodenot just some non-failed nodeto be able to generate valid responses, even if that node is isolated from other nodes. That is an unnecessary and unrealistic restriction. In essence, that restriction renders consensus protocols not highly available, because only some nodes, the nodes in the majority, are able to respond."

Partition with a capital P doesn't really move me as our experiences contradict its assertions, as the model it bases its statements on is fundamentally broken.

So let's talk about partition in the real world. Practically speaking, partitions mean that some nodes experience high latency when communicating. If you need to update some record in a specific data-centre and you can only to talk to that place via a single channel and that channel is disrupted; yes, you're going to experience a slowdown or even halt. Nowadays however, even widely used consensus-algorithms can mitigate that, and the delinquent node will eventually be dropped from the group if it causes enough problems. We don't do anything very different in this regard. Since our deterministic ledger can be replicated across multiple data-centres (as no nodes in it create original information), and since observers will only rely on records the ledger already sent out and since the only reason you would need to talk to this ledger is if you need to modify a shared record with other observers, you can always pick a different ledger-instance, there are no "master copies" anywhere. Remember that we can stream time-information with the data, so the client can always calculate its "point in time" and reconstruct a (even globally) consistent view of the data it has.

Sure, if you choose to centralise all your data behind a single flaky connection, you're gonna have a bad time. The point is that the setup we built allows you to not need to do that and it does that transparently, behind SQL semantics.

andras_gerlits··on Using Promise Theory to solve the distributed consensus problem
Okay, so we can move beyond CAP. So we don't talk about implementation in either the science-paper or in the intro for it (which is what this thread is about). I mostly write about implementation, Mark mostly writes about the science. So yes, redundancy will only make latency-spikes less likely. Notice however, that the way we establish strong consistency is based on communication also. There's a part in the essay I keep linking where I talk about what happens to nodes with flaky connections, it's called "Stability.

https://medium.com/p/5e397cb12e63#373c

There are three separate mitigation-solutions that go into how total-order strong consistency can keep marching on even if a specific child-group is intermittently isolated.

The first one is redundancy by deterministic replication, so there will always be many replicas which aren't just shards, but full copies of the _consistency ledger_. It's not a database, not a cache, just the thing that establishes the consistency between nodes. These instances all "race each other" to provide the first value of the outputs to other nodes.

The second one is the latency-mitigation we talked about earlier, I don't think we need to waste more breath on that.

The third one is that since the consistency-mechanism requires an explicit association-instruction to interleave the children's versions into its parent's (so that these versions can be exposed to nodes observing it from afar). If the child goes AWOL, it won't be able to associate its versions to its parent, so it won't keep up everyone else either. In this case the total-order won't be affected by the child-group's local-order, which is still allowed to make progress, as long as it's not trying to update any record that is distant to it.

andras_gerlits··on Using Promise Theory to solve the distributed consensus problem
Well no, by making the probability of latency-events extremely unlikely by establishing redundant channels.

https://medium.com/p/5e397cb12e63#7df1

Considering that this is maybe the 12th time I'm linking the explanation, I now think you're not looking for a discussion, you're here for a fight. Please let me know if you find any problems with the article in its reasoning. It was vetted by a lot of people before, so I really need to notify them too if you do.

andras_gerlits··on Using Promise Theory to solve the distributed consensus problem
It's clear for everyone that if you define Consistency via Linearizability, CAP-like problems will apply, as you're necessarily creating original information on a potentially remote node. That's not the issue.

The issue is that people in practice almost never use Linearizability for their Consistency, for example I don't know of a single SQL-implementation that does full linearizability (or Strict Serializable, same difference). So the industry already means something totally different from CAP's definition, and for those consistency-levels CAP doesn't apply. In fact, I present a technical series of arguments here how you can overcome these limits in practice: https://medium.com/p/5e397cb12e63

There's a specific section about CAP in the end, but I talk about node-loss. replication, strong consistency and all the others also.

andras_gerlits··on Using Promise Theory to solve the distributed consensus problem
Yes, that's exactly the problem with the existing models and why CAP was formalised.
andras_gerlits··on Using Promise Theory to solve the distributed consensus problem
Can I replicate people deterministically and control all their sensory inputs in this scenario?
andras_gerlits··on Using Promise Theory to solve the distributed consensus problem
No. CAP requires linearizability for its definition. If your consistency-model moves with the network's ability to communicate, your strong consistency can progress even if you somehow manage to lose your redundant replicas. This is what the CAP-section is about:

https://medium.com/p/5e397cb12e63#04a5

This is the summary: "In other words: any distributed solution that fits the SQL standard can rightly claim that it scales SQL databases, and Brewer’s model can certainly accommodate a framework for that. His model however, is not the only kind of distributed SQL database that can exist, therefore his assertion that all distributed consistent systems must pick where they position themselves on his famous triangle is wrong. The system we explain here for example, is an exception. Formally: because our consistency model stays within the bounds of what the SQL standard allows and includes network communication; and informally because we can fine-tune latency variability according to the use-case of the specific datastore within the system and can even be reduced to only be a theoretical concern."

andras_gerlits··on Using Promise Theory to solve the distributed consensus problem
Any modification to existing data must be "haggled for" somewhere, you're right about that. When you say "partition event", what is being partitioned here? A specific communication-link between two nodes. It's entirely possible (no, extremely likely) that your EU node would have access to a different US node, but not the one that's having the partition-event right now.

That's exactly the point of this section in the essay: https://medium.com/p/5e397cb12e63#7df1

Networks will have latency-spikes, but if you can stream time-information the same way you can stream others, you can use redundancies to mitigate the disruption of any single channel.

andras_gerlits··on Using Promise Theory to solve the distributed consensus problem
It's fine to not understand things. Being condescending is a natural reaction to feeling challenged, I really do understand where you're coming from. It's not even your fault, really. Anyway, when you're ready, feel free to pick it up again, the material doesn't take things personally and neither do I.
andras_gerlits··on Using Promise Theory to solve the distributed consensus problem
Almost there Dustin, thank you. This article is about the mental model behind scaling consistency across arbitrary geographical distances and how this model allows us to communicate time-information the same way we now communicate data. It's an intro to the science-paper, not the implementation.

With regards to the implementation (which is omniledger.io): We basically make SQL scale via Kafka by piggybacking on the semantics of both technologies.

There's no central authority for any of the tables. The version-ledger is totally schema-agnostic, only the clients understand it. In fact, it federates schemas the same way it does records. Tables are not topics either, in fact, we scale by importing namespaces via a command-line interface, which specifies a schema. Any database can register to this namespace after which they become as much of the "master" to the namespace as any other. There's no hierarchy between them, the ledger does everything for them. Since the ledger itself is deterministically replicated, a separate instance of it can be co-located with each instance that runs the JDBC-connections to the database, raising availability of it to the availability of the whole system. Each component in the setup scales with the number of partitions created for it.

andras_gerlits··on Using Promise Theory to solve the distributed consensus problem
This section discusses some failure modes (the first one is about the failure of a specific node): https://itnext.io/how-simple-can-scale-your-sql-beat-cap-and...

This section discusses latency-spike mitigation (which is how Brewer defines CAP colloquially): https://itnext.io/how-simple-can-scale-your-sql-beat-cap-and...

This section dissects the problem when trying to apply CAP to non-linearizable systems like SQL: https://itnext.io/how-simple-can-scale-your-sql-beat-cap-and...

Again, if you're not happy with the lack of scientific rigour in this technical article (ie: not science-paper), you can connect the dots in this one: https://www.researchgate.net/publication/359578461_Continuou...

andras_gerlits··on Using Promise Theory to solve the distributed consensus problem
Sure, there are other ways of doing this, the demos don't prove what I say, they only show that the theory works on a practical level. The science-paper however, does show how this mechanism can scale consistency.

Think about it this way: Our system simplifies running inter-system consistency and makes it much faster. Considering that Spanner exists, what proof would you accept to validate our claims?

I understand that reading all the material and putting them together isn't a small ask, but no new tech is easy to understand at first. Anyway, the material is there for anyone who cares to look and the system does what we claim it does.

I don't think anyone owes us their time to check our claims, but I don't think the fact that not lot of people will do this changes anything either.

andras_gerlits··on Using Promise Theory to solve the distributed consensus problem
Sure. Here: https://medium.com/p/5e397cb12e63#04a5
andras_gerlits··on Using Promise Theory to solve the distributed consensus problem
We do you one better. We show how all information can be made redundant via determinism and how that means you can supply multiple copies of them across parallel, redundant channels.

My essay talks about this in detail around the latency-mitigation section and the failure-modes part, but these are questions much closer to the actual implementation. Mark discusses the new mental model behind it, I talk about technology.

https://itnext.io/how-simple-can-scale-your-sql-beat-cap-and...

andras_gerlits··on Using Promise Theory to solve the distributed consensus problem
PACELC is an extension of CAP, so it suffers from the same problem of trying to apply a universal clock to the whole system. If you do that, you will suffer these limitations. With client-centric consistency, you can work around these problems and global order "unfolds" in a "just in time" manner.

The consistency-levels can go all the way to SNAPSHOT.

andras_gerlits··on Using Promise Theory to solve the distributed consensus problem
I don't think I ever heard people accuse our paper of not being rigorous enough, but more than happy to listen to specific problems with it: https://www.researchgate.net/publication/359578461_Continuou...
andras_gerlits··on Using Promise Theory to solve the distributed consensus problem
The CAP argument falls apart as soon as you decouple consistency from wall-clocks. Consistent systems only suffer from CAP limitations if they need strict serializability.

The point is that if you only look at the order of the data and not the wall-clock of some actor in the system, you can "calibrate" these separate ordering mechanisms together into a coherent whole and from that, you can build up all of the SQL guarantees.

Linearizability makes systems suffer because it tries to enforce a Newtonian model in a relativistic world. Order will naturally emerge faster when you're closer to the data than if you're farther. Measuring these with the same clock is what causes CAP, not some inherent property of distributed consistency.

andras_gerlits··on Using Promise Theory to solve the distributed consensus problem
I also have a demo here where I show transactionality between two SQL-databases, a MySQL and a Postgres instance: https://www.youtube.com/watch?v=XJSSjY4szZE

And another, where I show loose coupling, ie: that the system continues even when an instance goes down and that it catches up once restarted. https://www.youtube.com/watch?v=R4_phLs4d_M

andras_gerlits··on Using Promise Theory to solve the distributed consensus problem
It's ironic that so many of you are missing the rigour, as that is exactly the thing that "undoes" the CAP arguments.

Anyway, if math is what you guys are missing, it's in this science-paper linked in the article: https://www.researchgate.net/publication/359578461_Continuou...

This is a gentler intro to the concepts. You can also read my essay on why this setup works better than the often used semantics: https://medium.com/p/5e397cb12e63

There's a specific section at the end on why CAP only applies to a very specific subset of SQL databases.

andras_gerlits··on Using Promise Theory to solve the distributed consensus problem
How to decentralise and scale distributed consistency beyond the commonly accepted limits
andras_gerlits··on [dead]
How we built the world's first loosely coupled, ACID, master-master platform that can accommodate any mainstream SQL-database.
andras_gerlits··on Dora the Infra Exploder
The EU is (implicitly) mandating EU financial institutions to implement a hybrid-cloud infra. This should bring about the era of the commodity cloud and shared, scaled consistency
andras_gerlits··on How the microservice vs. monolith debate became meaningless
This is the first data-platform based on our original research.

For paper, see here: https://www.researchgate.net/publication/359578461_Continuou...

andras_gerlits··on How To Solve Cache-Invalidation
There's also another thread linked from there, which talks about the inherent limitations around distributed clocks.

https://twitter.com/AndrasGerlits/status/1690782596622381056

Page 1 of 2Next →