Deduplicating Archiver with Compression and Encryption
borgbackup.org
borgbackup.org
I've been expecting more and more from it, and at this point, there are only two things I wish would be better supported: - deduplication across several machines (you can already backup several machines to the same repository, but it's not efficient); - builtin redundancy to deal with bitrot.
I'm still using rsnapshot for secondary backups, as I can't afford a bug in backup software, but I'm considering switching to restic for that, as it provides deduplication as well and doesn't share code with borg.
1. no realtime incremental backups (that is, use filesystem observers in order to be able to perform incremental backups at short intervals
2. no multithreading
I didn't find any paired tool which could do the task #1, so for large systems, I think backups are very resource-intensive, and can't be scheduled frequently.
2. for many users not that important for the daily backups, because they are quick anyway. for first backup or for users with huge daily changes, it would be nice to have though, that's why it is on the long term TODO.
if you want to put more load on the CPU, you can partition your input data, feed each partition to a separate borg process (and separate borg repo). multiprocessing instead of multithreading.
- Rsync.net: https://www.rsync.net/
- BorgBase: https://www.borgbase.com/
- Lima-Labs: https://storage.lima-labs.com/
- others?
(I'm not affiliated in any way, I'm just a client of one of them.)
It is based in Malta, but why? For "tax efficiency"? They share the office address with [1] and the CEO of the parent company PeakFord[2] seems to be a CEO-for-hire[3].
Their GDPR page[4] mentions that they are not based in the EU and want to use the German privacy regulator. Malta is very much part of the EU.
[1]: https://opes.com.mt/corporate-services/ [2]: https://www.peakford.com/about/ [3]: https://offshoreleaks.icij.org/nodes/56060344 [4]: https://www.borgbase.com/gdpr
Hope this helps to clear your reservations. I should update our GDPR page with regard to the regulator. That's outdated by now.
Thanks for being a class act, based just on your presence here I might go Borg for a backup solution in the near future
If you only need incorporation, accounting and taxes solved, you can just pay ~hourly. Price will depend on time spent. Budget more in the beginning to find a good setup that will save time later. Accountants generally know very little about automation opportunities, so you need to find them yourself.
For sales positions, I'd always go for results-based compensation (like awarding equity as results are reached).
If the person does more than pure sales, you need a non-tech cofounder with some equity and the task to build commercial processes, do sales, decide features.
What you need will depend very much on your own skills and the product. Some stories: For my first startup, a physical consumer product, I had an equal partner with complementary skills who was very involved in product dev and sales. This worked well.
Later, in two other startups, I had technical cofounders and there was nobody to do sales and collect feedback. Worked less well obviously.
Hope this helps a bit and good luck with your idea!
I don't think it is - it has been around for several years and appears to be completely legit.
CEO is based in Hong Kong (I think) and appears to be more "international" than perhaps you or I are. He is relatively active here and on reddit.
I hope to call him up and meet for a drink next time I am building things in HK ...
More international for sure. But with Covid last year and offspring on the way this year, we are looking to spend more time in Europe closer to family. Hence the change of incorporation earlier this year.
https://www.time4vps.com/?affid=1881&_ga=2.61032364.12812044... - (referral link)
https://www.time4vps.com no affiliate link for time4vps
I previously used duplicity. It would do a full backup once every week or month, resulting in GBs of data being send over the line.
Borg chunks the data and uses deduplication. It result is just one huge backup at the beginning and incremental changes afterwards.
The abstractions used are also much friendlier. I never took the time to retrieve specific point in time backups from duplicity. With Borg one can mount the whole backup repository as a fuse mount, which each different backup being a distinct directory.
rsync.net offers cheap online storage for borg backups.
One caveat, borg is not yet multiprocess.
We also maintain an Ansible role to set it up quickly on new servers. https://github.com/borgbase/ansible-role-borgbackup
And last, if you need a place to put your backups, try my https://borgbase.com offering. It's purpose-built for Borg and offers handy features, like separating each repository and monitoring for stale backups.
It only deduplicates though. Does not compress.
Yeah, restic is a great backup tool also.
OTOH it is a pity the we could not join forces (Python vs. Go and also some other differences), but otoh it means more choice for everybody wanting a nice deduplicating backup tool.
I have a variety of machines I'd like to backup (desktops/servers/laptops/phones) running a variety of OSes.
I have a NAS ZFS machine that I'd like to host a copy of all my data from each machine.
From the NAS, I'd like to backup to a cloud host (e.g. rsync.net).
What I'm unsure about is where to introduce Borg in this scheme. I see a few permutations: 1) Borg from each machine to the NAS, then rsync to rsync.net 2) Rsync/ZFS send/etc. from each machine to the NAS, then Borg to a different location on the NAS, then rsync the Borg repo to rsync.net 3) Rsync/ZFS send/etc. from each machine to the NAS, then Borg directly to rsync.net
I'm leaning towards #3 personally. Thoughts?
I'd lean towards 1, so each machine has a standalone encrypted backup, but 3 would provide easier access without the borg client for local backups, and better de-duplication if files are shared across machines.
I guess I just don't see what value borg provides if you're already using ZFS throughout.
If you control the NAS, physically, then you can reduce some complexity by having unencrypted (and easily browsable) backups there ... and save the encryption for the rsync.net side of things ...
ALSO, if the NAS side of things is unencrypted, then you can establish a nice zfs snapshot schedule on the NAS and have those quickly and easily browsable as well. If you have borg backups on the NAS then even the simplest of restores becomes a full blown "restore" operation with decryption and keys, etc.
https://borgbackup.readthedocs.io/en/stable/faq.html#can-i-c...
2) It's proprietary with source-available but not open-source. This makes me hesitate to rely on it in the long-term as I don't know if it'll remain supported, especially considering that:
3) Development speed was very slow when I dropped Duplicacy about 6 months ago, and seeing Borg have many contributors makes me think it's more likely to stick around in the long-term.
4) GUI doesn't give much insight into the status of the backup other than progress %.
5) Restore operations are easier and quicker with Borg/Vorta because you can mount the backup via FUSE.
I also like how it is possible to navigate into archives by just mounting them in the filesystem: https://borgbackup.readthedocs.io/en/stable/usage/mount.html
Even wrote a blog post about it: https://simon-frey.com/blog/borgvorta-is-finally-a-usable-ba...
One thing that’s confusing me is the prune strategy. It seems prune removes the whole backup set at once. Let’s say I have daily backups and weekly backups. At some point I am pruning the daily backups. It seems this means that if the weekly backup ran on Sunday but a file was created on Monday and deleted on Friday, the file would be completely deleted during prune and there would be no trace of it.
I would much prefer if the pruning was done on per file basis and not by pruning whole backup sets.
When you create a backup ARCHIVE in a borg REPOSITORY, the archive contains all input files you gave to borg.
Each archive has a name and usually the name is something like machinename-setname-date-time.
You can create such backups rather often (that is cheap, due to deduplication), but you do not want to keep tons of archives long term (some borg commands take O(archive count) time).
Thus, one runs borg prune to "thin out" archives and only keep some following some prune policy, like keeping 60 daily, 12 monthly, 10 yearly or so.
It is important to use --prefix if one creates multiple different archive sequences in the same repo, e.g. if you create pc-home-date-time as well as pc-system-date-time - then you need to run prune with --prefix=pc-home- and another prune with --prefix=pc-system- .
Usually, archives are not modified after their creation (the only way to do that is to carefully use borg recreate command).
Then just don't prune! Unchanged files take no space. There's no reason you can't keep many months of daily backups, and set your prune to only do: --keep-within=365d
For those who aren't familiar with Borg, there is no such thing as a weekly backup, daily backup, monthly backup, etc. Every backup is a quick full backup. It's only the prune (which deletes old backups) option that gives you the choice of bracketing your backups into monthly/weekly/daily sets, so it'll save the last backup of each month for the last X months, the last backup of each week for X weeks, etc.
attic (python) - https://github.com/jborg/attic
borg (c) - https://github.com/borgbackup/borg
bupstash (rust) - https://github.com/andrewchambers/bupstash
duplicacy (go) - https://github.com/gilbertchen/duplicacy
duplicati (c#) - https://github.com/duplicati/duplicati
duplicity (python) - https://github.com/henrysher/duplicity
kopia (go) - https://github.com/kopia/kopia
nfreezer (python) - https://github.com/josephernest/nfreezer
rdedup (rust) - https://github.com/dpc/rdedup
restic (go) - https://github.com/restic/restic
rclone (go) - https://github.com/rclone/rclone
rsnapshot (perl) - https://github.com/rsnapshot/rsnapshot
snebu (c) - https://github.com/derekp7/snebu
tarsnap (c) - https://github.com/Tarsnap/tarsnap
I think there are many more out there (https://github.com/restic/others) - I personally use
restic
while technology wise (speed, only restore needs password) i would prefer rdedup
which is an impressive piece of software but unfortunately without file iterator... :-)attic should not be on the list, it is unmaintained since 2015.
borg came into life as a fork of attic for related reasons, so it should be borg (Python+Cython+C).
Borg has a "lackluster" encryption scheme. https://borgbackup.readthedocs.io/en/stable/internals/securi...
Restic can't pull backups. https://github.com/restic/restic/issues/299
Borg can't either, but it has "borg import-tar", which might be good enough.
> When the above attack model is extended to include multiple clients independently updating the same repository, then Borg fails to provide confidentiality (i.e. guarantees 3) and 4) do not apply any more).
There are some ideas on the issue tracker for fixing this long term (like random nonces, session keys, ...), but that stuff will have to wait until after borg 1.2 (which soon goes into release candidate phase).
There's ways to pull backups. You'd use ssh forwarding and/or the REST backend for append-only backups. See this comment:
https://github.com/restic/restic/issues/299#issuecomment-456...
Restic can take up a lot of RAM, unlike Borg.
Borg offers the option of no encryption. This is useful when backing up to a local drive that is already encrypted, for example with LUKS.
But, they are generally similar.
However having server side support reduces backup times.
Which is to say, you don't need the borg executable server-side if you are happy to run it over SFTP transport which limits some functionality.
Yes, of course you could talk over an sshfs mount point but I can't think of why you would do that over the more basic SFTP transport.
The accepted, and full-featured, way to run borg is with the executable on the server side and I will point out that the borg project distributes a "frozen" version of the tool that allows you to run it without having python in your environment. I believe they use py2exe to pack it up.
This is important for us[1] as we have no interpreters of any kind in our environment (no shell, no python, no perl) so we can only run binary executables ...
[1] rsync.net
“When Borg is writing to a repo on a locally mounted remote file system, e.g. SSHFS, the Borg client only can do file system operations and has no agent running on the remote side, so every operation needs to go over the network, which is slower.”
SSHFS would be tool agnostic/transparent though of course would result in operations going over the network as it’s not a local repo, but pseudo-local.
I’m curious, how do you use plain SFTP transport with borg using rsync.net?
Either way, having the borg executable on the server side is the "correct" way to do things.
no sftp inside borg (nor anything else).
Filippo Valsorda thinks[1] the encryption implementation looks sane.
Borg's "encryption" doesn't instill confidence in me.
Restic ships as a static binary so you can shove the version you used to create the backup and will be able to restore it forever (assuming we still run x64 machines).
Borg requires a server side component so cannot natively backup to cloud object stores. Restic can.
Borg OTOH is very lean and doesn't take a lot of RAM compared to Restic.
You cannot go wrong with either. I tried both and found Restic much easier to run and manage and the memory usage wasn't an issue for me (~4TB backups).
Borg's story is actually similar but the opposite. Attic only used to support zlib, other methods were added later. This was possible because the zlib header uses only a few values for the first two bytes, so there is enough room to indicate various compression formats.
With borg, every backup is logically a full backup (it just is faster due to deduplication). Much easier and less time consuming to deal with.