One of rqlite's big limitations is that it resyncs the entire DB at startup time. Being able to start with a "snapshot" and then incrementally replicate changes would be a big help.
One of rqlite's big limitations is that it resyncs the entire DB at startup time. Being able to start with a "snapshot" and then incrementally replicate changes would be a big help.
I'm not sure I follow why it's a "big limitation"? Is it causing you long start-up times? I'm definitely interested in improving this, if it's an issue. What are you actually seeing?
Also, rqlite does do log truncation (as per Raft spec), so after a certain amount of log entries (8192 by default) node restarts work exactly like you suggested. The SQLite database is restored from a snapshot, and any remaining Raft Log entries are applied to the database.
We're storing a few GB of data in the sqlite DB. Rebuilding those when rqlite restarts is slow and intensive process compared to just using the file on disk over again.
Our particular use case means we'll end up restarting 100+ replica nodes all at once, so the way we're doing things makes it more painful than necessary.
Try setting "-raft-snap" to a lower number, maybe 1024, and see if it helps. You'll have much fewer log entries to apply on startup. However the node will perform a snapshot more often, and writes are blocked during the snapshotting. It's a trade-off.
It might be possible to always restart using some sort of snapshot, independent of Raft, but that would add significant complexity to rqlite. The fact the SQLite database is built from scratch on startup, from the data in Raft log, means rqlite is much more robust.
We need that sqlite file to never go away. Even a few seconds is bad. And since our replicas are spread all over the world, it's not feasible to move 1GB+ data from the "servers" fast enough.
Is there a way for us to use that sqlite file without it ever going away? We've thought about hardlinking it elsewhere and replacing the hardlink when rqlite is up, but haven't built any tooling to do that.
Today the rqlite code deletes the SQLite database (if present) and then rebuilds it from the Raft log. It makes things so simple, and ensures the node can always recover, regardless of the prior state of the SQLite database -- basically the Raft log is the only thing that matters and that is guaranteed to be the same under each node.
The fundamental issue here is that Raft can only guarantee that the Raft log is in consensus, so rqlite can rely on that. It's always possible the one of the copies of SQLite under a single node gets a different state that all other nodes. This is because the change to the Raft log, and corresponding change to SQLite, are not atomic. Blowing away the SQLite database means a restart would fix this.
If this is important -- and what you ask sounds reasonable for the read-only case that rqlite can support -- I guess the code could rebuild the SQLite database in a temporary place, wait until that's done, and then quickly swap any existing SQLite file with the rebuilt copy. That would minimize the time the file is not present. But the file has to go away at some point.
Alternatively rqlite could open any existing SQLite file and DROP all data first. At least that way the file wouldn't disappear, but the data in the database would wink out of existence and then come back. WDYT?
Ie:
log1: empty db
log2..N: changes
log3: snapshot at N
log4..M: changes
You start with empty db, and replay up to N; or if you have a db that matches log3 - start with that and only replay log4+?Something to think about -- thanks.
The db is snapshotted at a preset time (easy way: lock db for writes whilst snap is in progress, preferably on read replicas) and then restored again (preferably, elsewhere) to see if restore did indeed work as expected (and isn't corrupt).
The latest log sequence number could also be stored in the same db (in a separate table) so that it is captured along with the db-snapshot, atomically.
Neat! This is reminiscent of the replay of event log in Smalltalk.