HNHacker News
TopNewBestAskShowJobs

_benedict

684 karma · joined June 13, 2014

submissionscomments
_benedict··on Why don’t we just open the windows? Covid-19 prevention lost in translation
> my understanding is that passive house is 3-4x better ACH than a typical recently built house in the UK

Part L1A of UK building regulations stipulate an air permeability of no more than 10m3/hr/m2 of building envelope at 50pa of pressure. My home has approximately a 1:1 ratio of building envelope to volume, though this will vary a great deal. But for such a home this works out to 10ACH at 50pa. There are various studies on how to convert these measurements to natural ventilation rates, but the conversion rate is around 10-30, with the typical house being adjusted by a factor of 20, i.e. 0.5ACH, however this can be as high as 1ACH.

Some homes will be constructed to higher standards, but will not likely achieve less than half this rating.

This also ignores the presence of trickle vents inside windows in homes without mechanical ventilation, and that the majority of homes in the UK were built to lower standards - typical values for homes built prior to 2010 are 15-20m3/hr/m2.

Therefore, 0.5-1ACH is a pretty standard background ventilation rate for homes in the UK. Your typical Passivhaus home with 0.4ACH from mechanical ventilation has less ventilation than a typical home, not more.

Only very large homes built to strict modern standards and with a high interior volume to building envelope ratio could possibly have 3-4x less ventilation.

_benedict··on Why don’t we just open the windows? Covid-19 prevention lost in translation
This isn't true. Most MVHRs have a summer bypass mode in which they admit the exterior air without retaining any interior heat, but they do not cool the air entering the house.
_benedict··on Why don’t we just open the windows? Covid-19 prevention lost in translation
Many non-passive house buildings have 0.4ACH or more without mechanical ventilation. The purpose of an MVHR is to provide similar trickle rates of ventilation without the high energy cost, but the goal of "opening a window" is to provide much higher airflow rates, in which case the cost would be much higher with mechanical ventilation.

It looks like 3-6ACH[1] or more is desirable to control the spread of infection. This would require very expensive MVHR systems to be installed. Realistically much larger units, probably more than one unit, as well as large ducting would need to be installed.

Such high flow rates would come with other problems: These flow rates would lead to unacceptably dry interior environments with typical heat exchangers, though there exist heat exchangers that retain some moisture. Higher flow rates also achieve lower heat exchanger efficiencies, and the power required to move the air would not be insignificant either.

[1] https://qz.com/1907977/how-to-check-air-ventilation-to-preve...

_benedict··on Raft Consensus Protocol
Thanks, glad to hear it is helpful. Do let me know if you have any feedback, particularly that might help clarify things.
_benedict··on Raft Consensus Protocol
It's still under development, this paper proposes it as an enhancement for Cassandra, so benchmarks are a way off. Given its properties it is likely to have similar performance to Tempo and Caesar for consensus (perhaps lower throughput than Tempo but better latency than Caesar). It is not dissimilar to Calvin for multi-shard execution, trading away the NxN sequencing layer (and poor failure properties) for dependency tracking.
_benedict··on Raft Consensus Protocol
Most of the trade-offs inherent to EPaxos are resolved by Accord[1], the protocol being developed by the Apache Cassandra community (myself included). There's also Tempo[2] and Caesar[3], amongst others, but they retain some of the trade-offs (and introduce their own). All of these protocols are also capable of supporting multi-shard operations, making them a significant advance over MultiPaxos and Raft.

[1] https://cwiki.apache.org/confluence/download/attachments/188...

[2] https://arxiv.org/abs/2104.01142

[3] https://arxiv.org/abs/1704.03319

_benedict··on Raft Consensus Protocol
Lamport intended MultiPaxos to be called Paxos, and it was described in his original Part Time Parliament paper which was written in 89 and published in 98.

These days I think Paxos refers both to what Lamport thought of as single decree Paxos, or Paxos leader election, and also the lineage of protocols that derive from it, as well as sometimes specifically MultiPaxos. But it’s OK, words can have multiple meanings.

For what it’s worth even EPaxos is long in the tooth now. There have been several advancements in the intervening years, the latest being Accord[1], a protocol the Cassandra community has developed for cross shard distributed transactions.

[1] https://cwiki.apache.org/confluence/download/attachments/188...

(Full disclosure: I’m one of the authors)

_benedict··on Paxos automatically determined safe and secure
> Let me know if you ever want to catch up to chat consensus and storage.

Likewise!

> working through figure 4 carefully is key to understanding the paper's contribution here, especially (c)(i) which is a simple scenario that can lead to indefinite unavailability of the cluster

Unless I am badly misreading figure 4, example (c)(i) seems impossible to encounter - with a correct implementation of Paxos, at least. For a record to be appended to the distributed log, it must have first been proposed to a majority of participants. So the log record must be recoverable from these other replicas by performing normal Paxos recovery, or may be invalidated if no other quorum has witnessed it. By fail-stopping in this situation, a quorum of the remaining replicas will either agree a different log record to append if it had not reached a quorum, or restore the uncommitted entry if it had. (c)(ii) seems similar to me.

All of (b) and (d) seem to exceed the failure tolerance of the cluster if you count corruption failures as fail-stop, and (a) is not a problem for protocols that verify their data with a quorum before answering, or for those that ensure a faulty process cannot be the leader.

> Discarding entries due to such conflation introduces the possibility of a global data loss (as shown earlier in Figure 2)."

Again, Figure 2 does not list data loss as a possibility for either Crash (fail-stop) or Reconfigure approaches.

> We actually found some interesting cases where PAR was able to help us maximize the durability of data we had already replicated

This is interesting. I’m not entirely clear what the distinction is that’s offered by PAR here, as I think all distributed consensus protocols must separate uncommitted, maybe-committed and definitely-committed records for correctness. At least in Cassandra these three states are expressly separated into promise, proposal and commit registers.

Either way, I and colleagues will be developing the first real-world leaderless transaction system for Cassandra over the coming year, and you’ve convinced me to expand Cassandra’s deterministic cluster simulations to include a variety of disk faults, as this isn’t a particularly large amount of work and I’m sure there will anyway be some unexpected surprises in old and new code.

_benedict··on Paxos automatically determined safe and secure
I’d just like to say in preface that I appreciate this back and forth, as text can feel more combative than intended. I have learned in the process.

> Again, the PAR paper describes the specific examples where a single disk sector failure on a single replica can be a showstopper event

Again, I don’t believe it does, but I may have missed it. If you could specify the particular scenario you imagine that would be great. The paper explicitly only rejects recovery in the scenario that there are already f failed processes, so we are talking about f+1 faults.

In essence, AFAICT the paper says that you can obtain additional failure tolerance by treating corruption as a non-fail-stop failure, and my contention is that the necessary replication factor for surviving other faults confers sufficient protection against corruption, and that as a result you most likely cannot reduce your replication factor with this improvement.

> Sure, just grep the 2012 VSR Revisited paper

Thanks. I read Section 4.3 before writing my prior reply, and I do not believe it addresses the problem.

_benedict··on Paxos automatically determined safe and secure
> However, PAR would mean that a cluster can at least recover itself and then replace the disk in the background, without this being a showstopper event.

This isn’t a showstopper event, it’s a process stopper event. I also outlined non-PAR approaches to recovering just fine without replacing the node.

> Theoretically, but in practice

I also have extensive practical confirmation that this approach works well.

> This aspect of recovery is often simply not tested.

I agree that test suites do not cover this well enough, and I commend you for working this into Tiger Beetle. This is something I hope to expand Cassandra's new deterministic simulation framework to incorporate in future as well, but Cassandra does benefit from a great deal of real world exposure to this kind of fault.

> PAR also shows cases where even a single local disk sector is enough to bring down a cluster that doesn't implement protocol-aware recovery for consensus storage.

I’m not sure I agree. The paper discounts reconfiguration because there could only be f+1 live processes and one could have corruption in a relevant sector, but under normal models this is simply f live processes. It seems that to accommodate this scenario we must duplicate the record identifier, and store both separately from the record itself.

It's not clear to me this scenario warrants the additional complexity, storage and bandwidth, as I’m not sure guarding against this is enough to reduce your replication factor. But it's worth considering, and this particular recovery enhancement is quite simple (we just need to duplicate the ballot in Paxos, so we can arbitrate between a split decision of other replicas that retain an intact record), so thanks for highlighting it more clearly for me.

> this kind of fault is addressed by Viewstamped Replication's 2012 revision

Could you point me to the place in the paper? I cannot see how VR solves this problem without introducing a risk of inconsistency, without necessitating an additional round-trip before responding to a client. Specifically, I think there are only three ways to address this particular problem, and I don’t see them discussed in VR:

1. The coordinator may record the responses of each replica before acknowledging to the client, so that the loss of any a write to any one disk may be recoverable. 2. The coordinator may require k>f+1 responses before answering a client, so that we may tolerate the loss of k-(f+1) disks losing a write 3. The coordinator may require an additional round-trip to record the consensus decision before answering a client

I don’t see any of these approaches discussed in VR. I see discussion of recovering the log from other replicas, but this cannot be done safely if the replica is not itself aware that it has lost data.

_benedict··on Paxos automatically determined safe and secure
> This is expensive and wasteful.

Recoverable disk faults of course are handled at the storage layer, so we're talking here about unrecoverable faults. My risk model prefers replacing any disk that is encountering unrecoverable errors, to avoid the disk causing additional problems as it degrades (no matter how good your error detection is, it is preferable not to lean on it more often than necessary).

To ensure only the affected data is concerned, it is possible to relocate only that disk's data to a new replica, to minimise the burden.

I'll concede that a better model might permit a replica some number of unrecoverable errors (few, but more than one in some time interval) before a total replacement is performed, but in this case there are still repair mechanisms to bring the replica up-to-date, and simply performing normal recovery for consensus operations whose state it cannot recover before it participates in further decisions is sufficient for ensuring correctness.

> could potentially lead to cascading failure

Trying to correct faults on broadly failing disks can lead to cascading failure, as the cluster becomes preoccupied with correcting faults that cannot be resolved, and user queries become unserviceable. Bootstrapping a new replica is also a very modest CPU burden.

> Worse, what if all nodes have a single disk sector error in different places on disk?

This isn't a concern in practice from experience, and theoretically the risk is also minimal. Unrecoverable disk errors occur at a rate of fewer than one per many petabytes. This provides plenty of time to replace a replica - and if this isn't sufficient, you can increase your replication factor until you can survive the requisite number of coincident faults.

> Could you explain how the paper addresses this particular kind of fault?

I'm genuinely curious to hear your answer here if you have a moment. I can imagine some mechanisms to help recover from a single fault like this in a distributed fashion, but would love to understand what I might be missing from the paper.

_benedict··on Paxos automatically determined safe and secure
> For example, how can a single replica even detect that it had a disk fault if the last I/O it acknowledged externally was simply never written to disk at all, despite an fsync barrier?

Could you explain how the paper addresses this particular kind of fault, as I cannot find it? This isn't one of the fault categories listed in Table 2, and the paper assumes that the storage layer reliably detects faults. I cannot see how a local system by itself can ever reliably detect this class of fault?

I can imagine coordinators detecting that a response from a replica is not up-to-date, but that happens anyway in most consensus protocols and this doesn't help correctness if only f+1 replicas had persisted the record (and another f have persisted a record with a conflicting outcome)

>I would highly recommend the PAR paper to you, if you have not incorporated it already into Cassandra.

I have read (well, skimmed) the PAR paper a few times, and while it is a nice paper I am unconvinced it is particularly helpful, beyond reiterating the fact that you cannot trust your disks - but this problem has to anyway be addressed by databases.

I think that the kind of active recovery it outlines could even be counter-productive in some cases. Once a disk fault is detected, it is unclear that it is safe or helpful to try to replicate correct data to the node. It may have less storage available to it now, if it has sensibly isolated the faulty disk, or may be unable to persist the data and so the recovery process may simply result in a lot of recurring work for the system.

Probably the most general purpose course of action is to replace the node, since nodes are plentiful. The faulty node can then have its disk replaced before being returned to the pool of available nodes.

_benedict··on Paxos automatically determined safe and secure
> I would say that the majority of leader-based state machine replication protocols today are based on the more popular RAFT, which is itself a derivative of Viewstamped Replication.

In production systems, most software appears to use Raft. In academia much more significant advances have been made for Paxos, and most of the research I am aware of is invested into Paxos. I am not really aware of much protocol advancement in academia that is built upon Raft?

> Implementing Paxos correctly in the presence of a realistic storage fault model is certainly non-trivial.

I disagree. The failure model for Paxos as specified assumes that processes fail-stop. So all that is required for correctness is that a process fail-stops in the presence of a detectable disk fault. To detect the kind of fault you report here, we only require checksums - either that are in part computed upon data that is not stored in the sector in question, or that are computed over multiple sectors (or both). This latter criterion holds pretty commonly, I think.

Of course, all distributed consensus protocols must handle reconfiguration in order to replace faulty processes, which would need to be invoked here. So this does not really seem any more challenging?

Certainly the software I maintain (Apache Cassandra), that currently employs Paxos is protected against this kind of fault by virtue of the chunk-level checksums that span many sectors, and the property that processes will not respond to operations that encounter any kind of local fault.

_benedict··on Paxos automatically determined safe and secure
What people think of as Paxos was not originally intended to be called Paxos by Lamport, but the leader election phase for what is now called Multi-Paxos (and he intended to call Paxos). So I don't think there's a difference for algorithmic heritage between Paxos and Multi-Paxos (perhaps you know this, but for those reading your comment it might help clarify).

I think it is anyway reasonable to say that most consensus algorithms derive concepts from Paxos, since (despite how terribly it was presented by Lamport) it is the consensus protocol that captured most attention. Most recent advances in distributed consensus derive from Paxos, not Viewstamped Replication. As far as I know all leaderless distributed consensus protocols derive from Paxos, and most protocol optimisations that have been developed apply to Paxos or one of its derivatives.

I also happen to think Paxos is pretty easy to implement correctly, particularly by comparison to the other protocols, in large part due to its active replication semantics, permitting that commands may be processed by replicas in any order. This means failover is much less complicated to negotiate, as the new leader does not expect to have a complete view of the log. Though of course membership changes remain complicated, and it may be beneficial for the leader to be able to assume it has a complete view of the log - but this is an optimisation rather than an inherent property for correct implementation.

> correctness unfortunately rapidly breaks down when physical disks are used, since these may misdirect reads and writes

What do you mean by "misdirect" here?

_benedict··on Netflix's new player breaks the ability to modify the seeking of a playing video
Why do you suppose people will select sites with good moderation? Is there any more reason to expect this, than that they would join a non-conspiracy Facebook group today?
_benedict··on Brexit is ‘going badly,’ say Brits in new poll
This was not a Brexit dividend, as participation in the procurement was optional, and it was undertaken while still a member of the single market. The same strategy could have been pursued while a member of the bloc.
_benedict··on Benchmark results: Cassandra 4.0, 3.11, Scylla 4.4 [video]
This is not what I was referring to, no.

When operating a database at huge scale surprising things happen, because everything that can happen will happen. So operators are interested in ensuring the database behaves well in these extreme circumstances. This isn’t specifically about horizontal scalability, though that is a necessary component.

To your point about vertical scalability, no doubt Scylla performs better here. However the details of your mentioned comparison are perhaps misleading, as Cassandra can happily exploit more than 16vCPUs before its performance materially plateaus.

While it’s true that the JVM imposes some restrictions, they do not translate to a difference in performance on the order of that claimed in this post. JVMs have also been NUMA-aware for some time. The main explanatory factor is relative investment and focus.

_benedict··on Benchmark results: Cassandra 4.0, 3.11, Scylla 4.4 [video]
In part it's also simply a demonstration of different priorities. Scylla's USP is performance, so a lot of elbow grease is spent there. The Apache Cassandra community is focused primarily on operating at scale, as that's its USP.

Performance is adequate for Cassandra, so the community has (for several years) primarily focused elsewhere. It will be a priority again in future, but in the meantime with many huge scale users out there the community has focused on guaranteeing correctness and stability at scale. For example, the Harry[1] toolkit for validating huge databases, and an adversarial cluster simulator[2] for exposing distributed and other complex bugs. Also a huge amount of behind-the-scenes work that isn't so easy to call out.

The community is now focusing on expanding the utility of the database for these use cases. For example the recently proposed enhancement to bring state-of-the-art general purpose transactions[3] to Apache Cassandra.

[1] https://github.com/apache/cassandra-harry

[2] https://cwiki.apache.org/confluence/display/CASSANDRA/CEP-10...

[3] https://cwiki.apache.org/confluence/download/attachments/188...

[edit] disclaimer: I’m an Apache Cassandra contributor involved with some of the above work.

_benedict··on The database ruins all good ideas
I believe it only guarantees serializable isolation. You may get strict serializable, but it doesn’t appear to be the case that you will know for sure if any transactions were not linearizable.
_benedict··on The Health Hazards of Plug-In Air Fresheners (2017)
I think it depends on the level of fibre release you expect. I don't think the issue is that the fibres bypass an N95 mask, but that if there are many fibres in the air you want to filter out a much higher proportion than 95%.

In the UK for non-licensed work (removing things like asbestos floor tiles which have low levels of asbestos that are bound in another substrate) you only need to wear the equivalent of an N95 mask.

_benedict··on EU vaccine rollout severely lags behind
The U.K. is also experiencing a significant shortfall in doses compared to those that were projected to be delivered by AZ. Production yields have been lower across the board according to Oxford researchers involved in the scale up, and several delivery milestones were missed for the U.K.
_benedict··on EU vaccine rollout severely lags behind
This is a very narrow view. The U.K. has made sure there are international supply chains to distribute the Oxford vaccine, and far more of that is being supplied internationally than the entirety of doses exported from the EU.

I think your 50% figure is also too high, though it is a high proportion, but these doses are primarily going to wealthy countries.

If the EU were asking for U.K. doses to be routed to other (particularly poorer) countries I would be sympathetic to the position, but this spat is about getting the EU doses, not about equitable access.

_benedict··on EU vaccine rollout severely lags behind
The UK is not hoarding, every vaccine is being used.

The EU so far is doing very little to vaccinate the world, for all its current moralising focus on shipping of a handful of doses to richer countries.

_benedict··on EU vaccine rollout severely lags behind
All of the mentioned actors are egotistical sociopaths in your example, as the EU is not talking about equitable access, only access for its own citizens.

If you care about equitable access you should be praising the UK, and in particular Oxford, for making its vaccine available at cost, and for having dedicated supply chains already established for poorer nations.

There's still plenty to criticise, but the current framing demonstrates the egocentrism of all of the actors.

_benedict··on EU vaccine rollout severely lags behind
It's worth remembering the UK bought 3x as many doses per capita than the EU (of AZ), and its domestic manufacturing capacity anyway has to serve a much smaller population. The luck in large part is the Ox/AZ was successful and fast, and the UK had placed a very large order for it.
_benedict··on EU vaccine rollout severely lags behind
It does not say there is no competing contract, it says no contract would prevent fulfilment of its contract. The EU contract stipulates that the delivery schedule is only an estimate, so another contract that was estimated to permit that delivery schedule would not conflict. It also seems to stipulate that the initial 300M doses will be manufactured within the EU, but this is incidental to this claim.
_benedict··on EU vaccine rollout severely lags behind
I don’t know why you presuppose the U.K. is where it would have had to be manufactured. Manufacture in the U.K. appears to have been made to happen at the insistence and investment of the U.K. government, from my recollection of a thread by an Oxford researcher involved.

There appears to be a similar deal with Novavax to setup similar domestic capacity in the UK for its vaccine.

_benedict··on EU vaccine rollout severely lags behind
That’s not been demonstrated. It has been claimed that a separate supply chain was initially purchased for the U.K. within the EU. No evidence has been presented that AZ doses from EU supply chain were diverted.
_benedict··on EU vaccine rollout severely lags behind
This is not what the contract says. The contract says there is no other contract that would prevent them fulfilling the EU contract. Since the contract makes no promises about delivery schedule it’s pretty clear the U.K. contract does not prevent its fulfilment, as the U.K. contract can only delay it.

However the EU contract seems to insist the initial 300M doses are made within the EU itself, and the EU has its own dedicated supply chain regardless that can eventually deliver, so this also seems to exclude conflict with the U.K. contract.

_benedict··on How to Efficiently Choose the Right Database for Your Applications
It doesn't fall under my definition of true* OSS as

1) it's owned by a private company, so the long term direction of the project is privately controlled;

2) it is released under only AGPL, by far the least business-friendly OSS license - I assume specifically to encourage direct licensing

There's nothing wrong with this, but it does constrain the freedoms of people who use the software more than the Apache license, and offers fewer opportunities to influence the project direction than the Apache Software Foundation (for all its many flaws).

Full disclosure: I'm a committer to Apache Cassandra, which is very much not a dead project, though it has been quiet for a while - focusing on not very visible aspects of the database.

* perhaps that's poor phrasing from my original post, or perhaps it is conveniently defined, but OSS isn't a scalar and we lack sufficient labels to express the relative freedoms associated with certain models

← PreviousPage 7 of 11Next →