The blog post has some numbers. We use Verneuil at backtrace to replicate small to medium sized databases (one per user), with up to a few hundred replicated databases per process. We see a few million write transactions a day (so not a
super heavy write load), and our periodic polls for replication lag only notes lag > 5 seconds ~100 times per day, and > 1 minute <= 1 time a day.
In practice, the main source of data staleness is often the period at which readers can poll for changes (a blind S3 GET of one blob, for each database). With a background thread to refresh data once a second, lag should usually be the order of 2-3 seconds. As to when that makes sense... I think it's good for data that doesn't see changes too often, and for data that's not directly generated and consumed interactively. For Backtrace, that's mostly one of:
1. metadata that's updated programmatically (e.g., after analysing crashes), and displayed to interactive users
2. data that's updated interactively (e.g., analysis configuration), and then propagated to worker processes
It wouldn't make sense to use the replication capability to display the current crash analysis configuration back to the user: when I change some configuration and hit save, I expect to see the changes I made. We want to service both interactive reads and writes from the local (source of truth) sqlite db.
Depending on the domain, I may or may not be OK with a propagation delay for the changes to impact behaviour; for example, we can let the analysis code fetch its configuration from a read replica, in order to improve isolation and scalability. When propagation delays are acceptable, read replicas help build more reliable and scalable systems.