HNHacker News
TopNewBestAskShowJobs

to_ziegler

54 karma · joined January 6, 2026

submissionscomments
to_ziegler··on TigerBeetle Core System Architecture: Deconstructing Performance Engineering
It's an insight from Heidi Howard et al. that came out after VSR: https://arxiv.org/abs/1608.06696 and can be applied to VSR (and others).

The basic idea is pretty simple. In VSR, there are two main phases:

1. Leader election

2. Normal replication / request processing

Before Heidi Howard’s insight, these two phases typically used the same quorum size - for example, 4 out of 6 replicas.

The key observation was that the two phases can actually use different quorum sizes, as long as the relevant quorums still intersect.

With 6 replicas, we could use a quorum of 4 for view change and a quorum of 3 for normal processing, because 4+3>6. This guarantees that every view-change quorum intersects every processing quorum. Therefore, if an operation was committed by a processing quorum, at least one replica participating in the subsequent view change knows about that operation. Combined with the protocol's view-change/log-selection rules, this ensures that committed operations are preserved when the new leader takes over.

If this interests you, Heidi gave a talk about this at systems distributed: https://youtu.be/P0cAG-RM1_c which will be released soon.

to_ziegler··on TigerBeetle Core System Architecture: Deconstructing Performance Engineering
Tobi here from TB. Great question, and it's important to be nuanced here.

It really depends on the problem space. For example, many OLAP workloads (analytical) contain large amounts of parallelizable work, then multi-core execution is absolutely the way to go. That also aligns well with the direction CPU technology is taking, with core counts continuing to increase.

For the transactional workloads we see at TigerBeetle, and in other transactional systems I've worked with, the picture is quite different. We see a lot of read-modify-write operations combined with a power-law distribution of the data.

Take a simple banking example: some accounts, such as those belonging to large online retailers, see much more activity than the average individual account. You might have 80 - 90% of transfers touching a relatively small number of these hot accounts.

Operations on the same account must be serialized to preserve correctness. That means this part of the workload cannot be meaningfully parallelized. In fact, attempting to parallelize it can make performance worse because of lock contention and coordination overhead, something the "Universal Scalability Law" captures quite well (but is also easy to test out yourself with a simple experiment).

Instead, we focus on batched execution. We carefully structure execution to make effective use of CPU caches and efficient algorithms, so that a single batch can be processed extremely efficiently without any coordination. Batch execution also allows to amortize I/O and replication.

That being said, there are areas where we could use multi-threading (e.g. compaction) that are not on the hot execution path.

to_ziegler··on TigerBeetle Core System Architecture: Deconstructing Performance Engineering
Hi, Tobi here from TB. Great question! Generally, 130 ms of network latency is challenging, and there often isn't an easy way around it as you're ultimately constrained by the speed of light (e.g. cross region deployments).

That said, network latency usually follows a distribution. For example, the median might be 130 ms while p99 is 200 ms. So one important goal is to avoid being affected by the high-latency tail.

In consensus and replication systems such as TigerBeetle, you can reduce the impact quite a bit by taking advantage of the fact that you only need a quorum. We have six replicas, and under normal operation we only need acknowledgements from three (including the primary, since we use flexible quorums). That means the primary only has to wait for the two fastest replicas to respond. This is very effective at reducing tail latency.

Then, to get as close as possible to speed-of-light latency, you want to avoid adding unnecessary latency inside the system itself. We've done quite a few algorithmic optimizations there over the past year. For example, introducing radix sort and tournament trees to make CPU processing more efficient.

to_ziegler··on High-Performance DBMSs with io_uring: When and How to use it
Pretty nice library! Personally, I think having a language-native library is great. The only downside is that it becomes a long-term commitment if it needs to stay in sync with the development of io_uring. I wonder whether it would be a good idea to provide C bindings from the language-specific library, so it can be more easily integrated and tested into the liburing test suite with minimal effort.
to_ziegler··on High-Performance DBMSs with io_uring: When and How to use it
There is some interesting ongoing research on eBPF and uring that you might find interesting, e.g., RingGuard: Guarding io_uring with eBPF (https://dl.acm.org/doi/10.1145/3609021.3609304 ).
to_ziegler··on High-Performance DBMSs with io_uring: When and How to use it
Great point. It is indeed the case that most io_uring libraries are lagging behind liburing. There are two main reasons for this: (1) io_uring development is moving very quickly (see the linked figure in another comment), and (2) liburing is maintained by Axboe himself and is therefore tightly coupled to io_uring’s ongoing development. One pragmatic "solution" is to wrap liburing with FFI bindings and maintain those yourself.
to_ziegler··on High-Performance DBMSs with io_uring: When and How to use it
The post is updated now to reflect this.
to_ziegler··on High-Performance DBMSs with io_uring: When and How to use it
Good catch! We will fix this in the next version and change it to brk/sbrk or mmap
to_ziegler··on High-Performance DBMSs with io_uring: When and How to use it
That’s great to hear! We are happy it helped.
to_ziegler··on High-Performance DBMSs with io_uring: When and How to use it
We also wrote up a very concise, high-level summary here, if you want the short version: https://toziegler.github.io/2025-12-08-io-uring/