Slog: Cheating the low-latency vs. strict serializability tradeoff
dbmsmusings.blogspot.com
dbmsmusings.blogspot.com
Both systems operate very similarly for local transactions that only touch data "owned" by a single master region; they just relay the transaction to be executed by the master. For multi-region transactions, Spanner uses a coordinator to perform two-phase commit, which acquires locks on all regions before allowing the transaction to proceed.
Slog does something similar, but effectively pipelines the locking to achieve higher throughput. First there's a global coordination step that globally-orders the transaction, without any locking (which means this step can use batching for high throughput). Then, each region's master independently acquires local locks in that global order, and replicates those locks as transactions so that replicas can deterministically apply them in the same order. Finally, each replica independently executes the transaction once it sees that all of the locks have been acquired. So a lock blocks the execution of conflicting transactions, but it doesn't block their replication. Once the replication is done, the locking overhead of actually executing the transactions should be comparable to a non-distributed DB.
All of this communication has a latency penalty, of course; there's no avoiding that for a consistent distributed DB. But the point is that it provides better throughput for transactions with conflicts. For transactions that only touch one region, the latency is still just a single round-trip to the master region, and that can be very fast if your client locality is high.
The benchmark results are heavily normalized, since it wasn't possible to do an apples-to-apples comparison on the same replication topology. So they don't demonstrate convincingly that Slog is faster than Spanner, in numerical terms. However, they do show that Spanner's throughput drops off much more quickly with increasing contention, compared to Slog.
IOW you start to suspect that the client may execute a multi-region transaction. Then you can prepare by syncing data across regions.
SLOG is CP from CAP, so indeed suffers from unavailability in the event of a network partition.