Mycelite: SQLite extension to synchronize changes across SQLite instances
github.com
github.com
db.commit()
# Give Mycelite time to sync before script exits
import time
time.sleep(5)
This really needs a way to wait for the sync to complete. Sleep won't cut it.Languages Rust 100.0%
did the link change??
https://mycelial.com/docs/get-started/python#observing-synch...
If I were to write a tutorial on how to set up PG replication and "observe the replication", I think that would probably involve "write to master, sleep for a bit, then see the data is in the secondary."
I asked this to ChatGPT 4 and it suggested something like this that seems reasonable:
Get it from the main server: SELECT * FROM pg_stat_replication, and the write_ls column has the log transaction location; Then get the location of the replication log and how much has been received: SELECT pg_last_wal_receive_lsn(); Then you can figure out the delta from any of these: SELECT pg_wal_lsn_diff(write_ls, pg_last_wal_receive_lsn()); or similar.
That seems like it isn't a hack.
This is just demo code, run these two short scripts to see a synchronising example, at least that's my reading.
//! Replicator prototype
//!
//! ** For demo use only!
And the configuration suggests it sends some data to https://us-east-1.mycelial.com [2] And that's not really explained anywhere I can see.I don't think the authors would currently suggest using it for anything real.
[1] https://github.com/mycelial/mycelite/blob/main/mycelite/src/...
[2] https://github.com/mycelial/mycelite/blob/main/mycelite/src/...
Every blocking operation needs an async variant. Everything should be async-first.
[0] https://github.com/vlcn-io/cr-sqlite/
[1] https://github.com/rhashimoto/wa-sqlite/
[2] https://github.com/rhashimoto/wa-sqlite/discussions/63
[3] https://github.com/vlcn-io/cr-sqlite/#how-does-it-work
Litestream & LiteFS are both physical, async replication tools so they actually work at the page level. When a transaction commits, the changed pages are shipped off to a replication destination. By replaying those changes in the same order, you can rebuild the current state of the database (or any point-in-time state in between). Both of these systems require a single primary node at a time so multi-master doesn't work with them.
Litestream only replicates to "dumb storage" (e.g. S3, SFTP) whereas LiteFS replicates to other live nodes so you can have live replicas.
The issue that CRDTs are solving isn’t simply getting the bytes from one device to another, it’s ensuring that conflicts are merged in a sane way, which this project doesn’t clearly address. Simply syncing a DB across multiple devices won’t help if your app isn’t designed to handle data conflicts.
CRDTs merge data in a predictable way but not necessarily in a consistent or correct way according to the requirements of any particular application.
I find CRDTs interesting and they are certainly useful for some types of applications with relatively lax consistency rules. But I don't see how they help with the sort of consistency requirements that finance apps are bound to have.
How likely is it that the right one exists (or that you can combine a few of them) to express some arbitrary, application defined consistency rule or merge strategy?
On top of that, any solution would have to agree with all your other requirements, such as your chosen degree of normalisation, your performance requirements, etc.
The question OP is asking is, why is it comparing itself to CRDTs and Actual Budget? If Mycelite has special sauce, what is that special sauce? If it doesn't have special sauce, do they understand they're attempting to solve a problem which requires special sauce?
It is just unclear to me how this project does so, and if it doesn’t, then why it even references CRDTs in the first place
It works differently - there's a persistence library called Core Data on each client (most recently SwiftData). You can choose which database to use - most likely SQLite, but you could use XML files to store your data.
And then Core Data simply makes sure everything is in sync. It's an Apple system library, which means your data will automatically sync even if the app isn't running.
This feels like better approach than OP's Mycelite or better-known cr-sqlite.
Any reason why syncing on the database level would be better? It feels like something the library should do, not the database.
This year at WWDC we got CKSyncEngine which gives you the primitives to add a CloudKit syncing to your own persistence layer.
Then we got personal computers.
Then we got mainframes again with HTML5 dumb terminals, rebranded as "cloud."
Now we're starting to see renewed interest in personal computers.
'cept this time round, those slick black rectangles are so locked down you can hardly program on them. And not to mention that people using them are mostly just looking at tiktok, twitter, instagram, snap, youtube.
Personally I don't consider the black rectangles real computers. They're consumer content consumption devices, basically like televisions. They are to computers what a TV is to a production studio.
People seem to forget how hard it is. So lets say you're pushing changes from one system to another. How are you detecting changes? timestamps or a direct diff? Or checksums? Have you remembered to think about deletions? Can you even detect deletions? How do you keep track of which bits have successfully synced and which havent? If something unusual goes wrong (network goes down) halfway through a sync can you recover from that? Incremental sync inevitably drifts over time due to bugs and downtime, so are you also building a full-scale wipe-and-refresh or compare-and-reconcile to accompany the incremental sync? And dont even get me started on two-way syncs and conflict resolution and how hard that is to do without some sort of review-by-human component.
Furthermore decisions about which db is used are rarely driven by the synchronization model. Except maybe sometimes the question "eventual consistency yes/no/a bit only".
And in academia (wrt. common use-cases) we do have a bunch of well analyzed ways to archive synchronization the main question often isn't one of academic research needing to be done but of which model with which parameters and implementation details on top to use for a specific kind of db. Which mainly is discussed by devs when writing new dbs, i.e. not a everyday thing. While for the many new sqlite sync approaches the answers is nearly always "handwaving eventual consistency", cause nothing else would work on the edge anyway.
Which, for some systems translates to: It will be in sync when all the entropy in the universe is gone.
sure some systems under correct operation will never ever be fully in sync as long as they are used
but most systems also do not need to be ever fully in sync
it's good enough that for a specific context they will be in sync in not to much time if that context stops changing
and that is something they do provide
e.g. after updating a JSON blob stored under a specific id that update will be eventually available in the not too distant future and if no future changes to that document happen then in the context of that document the system will be fully in sync. But because you have very man documents there will always be a document which isn't yet in sync and in turn the system as a whole will never be fully in sync. But in the end that doesn't really matter.
Also: no computer system will live that long, and if they aren't put in sync before they stop working they will literally never be fully in sync.
it's still your system
but a system as a whole never being fully in sync is a red herring argument which misses the point and focuses on an aspect which might sound like a problem but hardly ever is a problem at all
I haven't had a chance to compare yet.
This wouldn’t work ofc as fs access in Vercel is limited, plus syncing the db would (probably) delay the network response. But
[1] https://rqlite.io/docs/faq/#can-i-use-rqlite-to-replicate-my...
On the write-side, you can load the verneuil VFS as a runtime extension. There's a similar gotcha with reboots, where the replication state is persisted but not fsync-ed, so you get a full refresh on reboots. The S3 storage is content addressed though, so you "only" pay for useless bandwidth and API calls, not storage.
Can you talk more about your use case? I'm the author of LiteFS which does real-time replication of databases. It doesn't sound like it's a fit for your use case but I'm always interested in hearing about what people are doing.
How "real time" is the synchronization?
What are limitations?
E.g. for an sqlite db that lives on a Network Attached Storage (NAS)?
There's probably an extension for permissions in sqlite but out of the box the db files are plaintext (you need an extension for that).
Good starting point for extensions: https://antonz.org/sqlean/
SQLite only supports a single writer at a time, but does support concurrent reads. There is also an experimental branch that supports concurrent writes.