Kopia: Fast and secure open-source backup software
kopia.io
kopia.io
> Yes, there's a whole bunch of things currently not captured at the filesystem level, including setuid/gid, hardlinks, mount points, sockets, xattr, ACLs, etc.
That was three years ago and it sounds like things are only slightly better.
These people are incompetent at making backup software.
I can’t be bothered to fix the names (it’s not realistically a problem), but if backup software can’t handle them, then it can’t handle backup up my data.
True, but one should get into the practice of creating persistent policies rather than ad-hoc `chcon`'s all the time, so that a `restorecon -FR /` always restores all SELinux labels to the intended state, then the biggest hassle is just booting into initramfs shell to mount rootfs and `touch /.autorelabel`
We have many files (millions) and lots of churn over ~80Tb total.
Kopia has exhibited some issues:
- takes about 120GB (!) of ram to perform regular maintenance & takes about 5hrs to do so. There are ideas floating around to cherry pick the large inefficiencies in the GC code but it’s yet to be worked on. I’ll try to have a internship accepted to work on this in my company.
- there’s a good activity on the repository but the releases are not quick to come and the PRs are not very fast to be examined
- the local cache gets enormous and if we try to saddle it, we have huge download spikes (>10% of repo size) during maintenance. Same as above: pb is acknowledged but yet to be solved
- the documentation is very S3 centric, and we discovered too late that the tiered backup (long term files go into cold storage on s3) is only supported on S3, while we use azure. We contributed a PR to implement it in June, yet to be merged (see point 2)
So, not too bad, especially for a small-ish project maintained by mainly one person (from the looks of my interactions on slack and seeing the commit log). The maintainer is easy to reach and will answer, but external prs are slow. If I could use zfs cheaply on azure via s3, I’d use it over kopia, but as of now, it works.
Am I dumb for just doing some rclone+rsync.net?
I am using rsync to rsync.net from multiple different hosts with different configurations. I run the same command on every host running variations of *nix, no messing about with different tools needed.
I found Borgmatic ( https://torsion.org/borgmatic/ ) to be the best way to run my backups. It takes care of everything from pruning to verifying the checksum etc... and it integrates with some monitoring (like cronitor).
So Borgmatic + rsync.net is the best combo
Also, storing the ZFS snapshots on Blob storage would still require us to retrieve the entirety of the 80TB before being able to use it. I need native ZFS at Blobstore-competitive prices
Can't you?
Either with basically:
zfs send|rclone rcat
Or more directly via s3 tooling?https://github.com/presslabs/z3
I am obviously biased but it's pretty amazing. AMA.
Do you (sorry, but just checking) repeatedly test backups? Eg pull monthly and bit verify that they're correct? Are you aware of anyone testing in this way?
Thanks so much!
It still has a lot of potential, IMHO. You e.g. find some hints how to use it with AWS storage tiering in the docs.
I am just a very happy user!
And, while not directly, I know a number of companies, including mine, do test restores all the time.
Duplicati has a web interface, so with a proper authentication in place, you can use it to remotely monitor and manage backups.
Duplicati doesn't keep a local cache. It uses SQLite files for file meta data, but not for the content themselves.
I like Duplicati's snapshotting mechanism. You can specify how long or how many snapshots to keep, and my anecdotal evidence is that it's archival storage-friendly. I imagine S3 and it's lifetime management rules can bring a decent and cost effective backup solution.
I'm using Google Drive 2TB plan, and I didn't see Kopia supporting Google drive out of the box.
So anyway, I'm looking for alternatives.
Duplicati also, somewhat annoyingly, nails 100% cpu for a while during backup which spins up the fans and gets my laptop very hot. I've been meaning to see if there's a simple way to modify the code to prevent this, but I'm very unfamiliar with C#.
- Restic (https://restic.net/)
- Borg backup (https://www.borgbackup.org/)
- Duplicati (https://www.duplicati.com/)
- Kopia (https://kopia.io/)
- Duplicay (https://duplicacy.com/)
- Duplicity (https://duplicity.us/)
I've tried all of these, and none are as reliable or powerful as Bacula.
It's way more complex at first, but you will have peace of mind. And backup/restore speed is way faster.
You can even easily setup automatic restore jobs to prove your backups work!
Do you have any odd requirements that one might serve better than the rest? If you just want bog-standard backups, any of them will probably do.
It might be worth adding an asterisk there.
https://borgbackup.readthedocs.io/en/stable/faq.html#can-i-b...
Faster than any other backup software (because it knows what's changed from the last snapshot being the filesystem itself but external backup tools always have to scan the entire directories to know what's changed), battle tested reliability with added benefit like transparent compression.
A bit of explanation on how fast it can be than external tools. (I don't work for the said service in the article or promote it.)
https://arstechnica.com/information-technology/2015/12/rsync...
Then you'll realize Borg is the one with least data corruption complaint on the internet which is good as your secondary backup.
Easily checked with, "[app name] data corruption" on Google.
And see who else lists vulnerability and corruption bugs upfront like Borg does and know the developers are forthcoming about these important issues.
https://borgbackup.readthedocs.io/en/stable/changes.html
The term "best" apparently means reliable for backup and also they don't start choking on large data sets taking huge amount of memories and roundtrip times.
They don't work against your favorite S3 compatible targets but there are services that can be targeted for those tools or just roll your own dedicated backup $5 Linux instance to avoid crying in the future.
With those 2, I don't care what other tools exist anymore.
Initially I thought this was a corporate project and was looking for the monetization model, but then I found https://github.com/kopia/kopia/blob/master/GOVERNANCE.md
I feel like the project might benefit from making their governance model more prominent on the website.
Too bad no one besides kopia does ecc, which is the reason I switched, but when I checked out why restic didn't do it, it was because they saw what other people did and apparently it's way too easy to make a bad implementation.
I tried to restore a ~200 GB file (stored remotely on a Hetzner Storage Box), and it failed (or at least did not finish after being left for ~20 hours; there was also no progress indicator or status I could find in the UI).
I also tried to restore a folder with about ~32 GB of data in it, and that also failed (the UI did report an error, but I don't recall it being useful).
Also, in general use, the UI would get disconnected from the repository every few days, and sometimes the backup overview list would show folders as being size 0 (which maybe indicated they failed; they showed up with an "incomplete" [or similar] tag in the UI).
Just for fun, since I still had it installed and haven't gotten around to cleaning up the remote data, I updated to latest (v0.14.1) and tried the restore tasks again.
Both the single file restore and the folder restore worked, though the single file restore still didn't have any progress indicator I could see.
Looking through the changelog, nothing really stood out to me as something which would have fixed this. Not really sure what went wrong the first time around, perhaps it was network issues with Hetzner?
Edit: Found this very ad-hoc "benchmark" from over a year ago claiming that Kopia managed significantly better deduplication than Restic after several backups (what took Restic 2.2GB, Kopia did in <700MB). No idea if the advantage falls off outside of this particular benchmark, but if it doesn't then that's a pretty big improvement. https://github.com/kopia/kopia/issues/1809#issuecomment-1067...
Edit Edit: Never mind, this benchmark was from before Restic supported compression, which is the why its size is so much larger. Feels like that should have been mentioned.
My dealbreaker with Restic was near-realtime backup - the discussion has been open for 5 years now: https://forum.restic.net/t/continuous-backup/593); this is also a UX problem. I haven't checked if Kopia supports it (or has better support, anyway), though.
https://relicabackup.com/features
https://github.com/netinvent/npbackup
Which is my current short-list for cross platform backup...
> ... was from before Restic supported compression, which is the why its size is so much larger
deduplication != compression...
So we don't actually know which has better deduplication? Compression algorithms are well-established and you can find a million people who benchmarked them all for different purposes, but deduplication algorithms I never saw a comparison of. I don't even know if these things have proper names or if people just refer to them as "rsync-like"
The presence of a WebUI is so nice compared to CLI-only tools.
kopia also cleanly support non Linux platforms like MacOS and Windows. It has a UI too but borg is getting some now too.
The UI didn't seem like a good fit for those who are less technical. I don't remember the specifics.
Does anyone have recommendations for backup services for the average user?
> The UI didn't seem like a good fit
Didn't seem like, or wasn't based on your trying it? Did you end up trying it? If not, what do your relatives use currently, is that better (even if not ideal, since you're still asking for recommendations)?
These days, from the minimal time I've spent looking at this, it seems that Borg and Restic offer basically the same feature set with similar performance. I'm curious if you (or anyone else) considered Borg and what set Restic apart for you.
Ditto for Kopia, I guess. I've never heard of it before.
"Kopia" means "copy" in Swedish and probably more Nordic languages, too. Very hard to pronounce in English so it would be interesting to hear it said.
As a Pole, I actually greatly appreciate these Slavic names in tech.
Tangentially, as far as OSS names of Polish go, kopia is pretty tame. A popular deduplicating app is named czkawka (hiccup). Now that choice is just mean towards non-Polish speakers. :)
Czkawka is a Polish word which means hiccup.
I chose this name because I wanted to hear people speaking other languages pronounce it, so feel free to spell it the way you want.
This name is not as bad as it seems, because I was also thinking about using words like żółć, gżegżółka or żołądźWell indeed.
There's a project on GitHub with 1.7k stars called GitKurwa[1].
Now that's proper untame Polish. ;-)
Does anyone know how Borg 1.x and 2 would compare to Kopia?
Sigh… and unfortunately all too common for there to be no cold storage support.
My use cases are very basic, but it just quite works all the time. My total backup is roughly 1TB with more than million files at multiple locations like local, remote SFTP and Amazon S3.
I like being able to extract the selected data(directory/file) from snapshot without restoring the whole snapshot.
de-duplcaition is icing on top.
The developer Jarek is very responsive on Kopia Slack as well.
* Can run in a "I haven't done a backup for a while, so I'll do one now" mode when the laptop is awake.
* Both laptops can be writing to the same repository at the same time, sharing common files, dedup, etc.
Missing:
* Only supports VSS via a couple of scripts I couldn't get working. (Restic is nice with that.)
Is https://kopia.io/docs/advanced/actions/#windows-shadow-copy not working for you?
But I probably ought to double check that restoring works :)
Such as ?
I see no mention of any of that anywhere obvious. Do you get to just make up core properties of your software because it feels good?
One feature that I wished these tools would have provided is support for public key encryption. The secret key then need not be exposed.
https://github.com/andrewchambers/bupstash/blob/master/doc/g...
I got a Synology in my house that could be utilized
Note that Veeam contributes to Kopia - https://www.veeam.com/sys451
[1] https://www.veeam.com/agent-for-windows-community-edition.ht...
Veeam is an enterprise solution, and popular among sysadmins.
Also, it's unclear to me what happens if you attempt a snapshot in the middle of something like a database transaction or even a basic file write. Seems likely that the snapshot would still be corrupted. So for databases you're stuck using db-specific methods like pg_dump.
All this complexity makes it very difficult to make self-hosting realistic and safe by default for non-experts, which is the problem I'm having.
[0]: https://forum.restic.net/t/what-happens-if-file-changes-duri...
[1]: https://learn.microsoft.com/en-us/windows-server/storage/fil...
I personally use compose for all my services now and back up my compose.yaml by stopping the entire stack and running a restic container that mounts all volumes in the compose.yaml.[1] It's not zero downtime, but it's good enough, and it's extremely portable since it can restore itself.
[1]: https://gist.github.com/acuteaura/61f221ada67f49193bc1f93955...
Perhaps you could be more specific, because the former is exactly what a filesystem snapshot is meant to do, and the latter is exactly what an ACID database is meant to allow assuming the former.
> Look at what Kanister does with its recipes to get consistent DB snapshots
I looked at a few examples and they mostly seemed to involve running the usual database dump commands.
You just quiesce the database first. Any decent backup engine has support to talk to a DB and pause / flush everything.
Now of course it's all about ZFS, so there's at least snapshots paired with replication - but the story for anything else is still pretty bad, with you having to put all the fiddly pieces together. I'm sure some people taught their backup tool about their special named backup snapshots sprinkled about in `.zfs/snapshot` directories, but given the fiddly nature of it I'm also sure most people just ended up YOLOing raw directories, temporal-smearing be damned.
I know I did!
I finally got around to fixing that last year with zfsnapr[1]. `zfsnapr mount /mnt/backup` and there's a snapshot of the system - all datasets, mounted recursively - ready for whatever backup tool of the year is.
I'm kind of disappointed in mentioning it over on the Practical ZFS forum that the response was not "why didn't you just use <existing solution everyone uses>", but "I can see why that might be useful".
Well, yes, it makes backups actually work.
> Also, it's unclear to me what happens if you attempt a snapshot in the middle of something like a database transaction or even a basic file write. Seems likely that the snapshot would still be corrupted
A snapshot is a point-in-time image of the filesystem at a given point. Any ACID database worth the name will roll back the in-flight transaction just like they would if you issued it a `kill -9`.
For other file writes, that's really down to whether or not such interruptions were considered by the writer. You may well have half-written files in your snapshot, with the file contents as they were in between two write() calls. Ideally this will only be in the form of temporary files, prior to their rename() over the data they're replacing.
For everything else - well, you have more than one snapshot backed up, right?
I know that Tape Backups are not hip and sexy, but CloudNordic showed us just last month why they still matter even in 2023 and beyond, so you'd definitely want to look at an additional solution for your large servers, with a proper rotation/retention strategy (e.g., GFS). You _need_ offline backups, if you think you don't, you just got lucky for now - or have data that can be recreated from other sources.
For an online hot/warm solution, I'd use sending ZFS Snapshots into a backup server to then compress and encrypt them there, though keep in mind that for running systems, it may still not be enough (e.g., backing up a running Postgresql server through a file system snapshot may not be enough - there's an entire section in the documentation about backup options).
That said, it's good to have more options, and you really want to use something for your personal stuff as well, so the more options there are, and the more user-friendly/turnkey they are, the better!
Just be aware that backup solutions in a corporate/network environment are more complicated than just copying some files across. And also remember: Good companies test their backups - but great companies test their restores.
https://www.postgresql.org/docs/16/backup-file.html
> An alternative file-system backup approach is to make a “consistent snapshot” of the data directory, if the file system supports that functionality (and you are willing to trust that it is implemented correctly). The typical procedure is to make a “frozen snapshot” of the volume containing the database, then copy the whole data directory (not just parts, see above) from the snapshot to a backup device, then release the frozen snapshot. This will work even while the database server is running. However, a backup created in this way saves the database files in a state as if the database server was not properly shut down; therefore, when you start the database server on the backed-up data, it will think the previous server instance crashed and will replay the WAL log. This is not a problem; just be aware of it (and be sure to include the WAL files in your backup). You can perform a CHECKPOINT before taking the snapshot to reduce recovery time.
It sounds like enough to me.
Can't even claim that this is a recent addition, since it's documented like that since Postgresql 8, which was released in 2005.