Show HN: Baxx – Unix-friendly backup service
txt.black
txt.black
For Unix/Linux backups, may I suggest Borg Backup? It encrypts and does dedupe astonishingly well. It also works over SSH incredibly fast, and restores are via a mounted FUSE filesystem so they're easy to pick and choose what you need. It prunes really well too, and is a single executable so it's easy to distribute via Ansible/Puppet/etc. I've been using it for several years and it hasn't failed me yet.
If y'all want to laugh at my bash scripting skills, I have a backup script that sends backup status to a Zabbix monitor server at https://gist.github.com/anthonyclarka2/cef41d201dd5b890dae67... (I'd appreciate improvements to that script if anyone wants to critique)
Simple and explicit is better in shell scripts than being fancy.
It is also very easy to setup a pruning policy for old backups, so that you can say that you want to keep one backup per day for 90 days and after that only one backup per month for 2 years.
Borg runs a binary on the server to do this, that's why it can't easily back up to S3 et al.
Restic also starts uploading changes as soon as it sees them, whereas borg has to scan the entire list of files before it starts uploading changes. Unfortunately for multi-tb datasets with millions of files, borg would take hours just to scan all of the files.
Pruning old archives also worked a lot better in Restic compared to Borg.
If you're happy with Restic, I'd keep sticking with it.
I don't know if it's an option for you, but have you considered something like rsync from each server to a central backup server and then running restic/borg/whatever on that server? Sometimes that works out nicely too, and would be less impactful to other services on the machine.
What did you switch to instead of restic, if you don't mind me asking?
I was using restic and switched to borg because of that and could save hundreds of gigabytes.
Most large files are already in compressed formats: mp4, zip, tgz, jpg, etc.
What you _really_ want is deduplication (which restic does), so that chunks are only stored once, even if multiple machines sharing some data back up to the same repo. Dedup saves me about 30% per backup repo in my own experience.
Not saying compression wouldn't be nice (to have done properly, it is somewhat at odds with encryption), but I'm just saying I don't think it's worth making it a #1 deciding factor.
If I was backing up a single computer, compression would matter more to me. But if I'm backing up 10s or 100s of servers, there's going to be a lot of duplicated data that would make a much bigger difference.
Borg is a fantastic tool but the time it takes to start and parse the files req on a large system negates the benefits of compression for me unfortunately.
It just lacks terminal-based registration. :-) You could use the GraphQL API to manage everything from the command line if you really wanted.
Borg needs a "smart" backend with a filesystem, so it can't use only object-storage.
> When the above attack model is extended to include multiple clients independently updating the same repository, then Borg fails to provide confidentiality (i.e. guarantees 3) and 4) do not apply any more).
The first entry on the Borg FAQ [1] is “can I backup from multiple servers into a single repository” and the answer is basically “yes, but it might be slow”. Seems to me like it should say “not really, it's insecure and might be slow”, plus the tool should make it clear you're doing something unwise if you try it anyway.
I've always shied away from shared backup repositories anyway, because it usually means data from host A can be read by host B, and that's just not what I usually want. The deployment instructions appear sane to me. [2]
[1] https://borgbackup.readthedocs.io/en/stable/faq.html#can-i-b...
[2] https://borgbackup.readthedocs.io/en/stable/deployment/centr...
1. Borg's authenticated encryption is a composition of AES-256 in CTR mode and HMAC-SHA256 (or alternatively, Blake2b256). This consists of two distinct constructions. Most cryptographic vulnerabilities are introduced in the combination of distinct primitives and constructions. While combining AES-CTR and HMAC is a way of doing authenticated encryption, it's not the safest way. As a specific example - in order to avoid introducing a nonce-misuse vulnerability, Borg needs to implement specific logic beyond the construction to ensure that the CTR counter (containing a randomized nonce) is not reused. This kind of overhead is dangerous because it adds a lot of room for a developer to make a fatal mistake.
It would be much better to use a dedicated authenticated encryption scheme, like AES-GCM. AES-GCM is a more complex construction, but it also has very convenient interfaces and language bindings which obviate the need for developers to interact with raw primitives and constructions entirely. You don't need to implement AES-GCM so much as call its seal and unseal operations from your language of choice. There's far less room to footgun yourself here. The even more modern method would be to use something in the ChaPoly family. For example, XSalsa-Poly1305 goes one step further than AES-GCM by using an extended nonce. ChaPoly constructions also have ubiquitous language bindings and easily used interfaces to safe implementations.
2. Borg uses OpenSSL directly. This is not great because libcrypto exposes raw primitives directly to the developer. To its credit, Borg's documentation does provide a few reasons why its use of OpenSSL is trustworthy. But those reasons are more to do with TLS implementations than cryptographic constructions. As with above, working directly with raw primitives provides a lot of room for implementation errors and avoidable security vulnerabilities. It's generally better to abstract away your cryptography implementation to a known-good, trustworthy source which you can just call via a straightforward (and opinionated) interface. For example Borg could use NaCL instead, which would provide XSalsa-Poly1305 encryption out of the box. I hypothesize they'd eliminate a few hundred lines of code with NaCL instead of OpenSSL.
That being said, I don't know if I'd call Borg's encryption "weak." It's not ideal, and I don't personally trust it. But it's not like they're using AES in ECB mode or Mac-Then-Encrypt CBC mode.
> It would be much better to use a dedicated authenticated encryption scheme, like AES-GCM.
Actually, in the scheme Borg uses, which is vulnerable to repeating IVs/nonces, using AES-GCM or Chapoly (or indeed any Wegman-Carter authenticator) would not only delete confidentiality, but also authenticity, i.e. it wouldn't just disclose plaintext to the attacker, but allow an attacker to potentially change your backups [1].
> That being said, I don't know if I'd call Borg's encryption "weak." It's not ideal, and I don't personally trust it. But it's not like they're using AES in ECB mode or Mac-Then-Encrypt CBC mode.
I'm calling it out as weak because it is an insufficient and poor design. If you say encrypted today it should better be up to me throwing my encrypted data up on a random cloud server and be sure that it's actually encrypted. This is not the case with Borg; partial credit for getting some aspects of the construction right (e.g. EtM, separate keys, properly salted master key encryption key derivation) isn't worth much when it fails to provide confidentiality in not-at-all unreasonable, practical usage.
[1] Probably not because by necessity Borg uses a separate HMAC over the plaintext for deduplication, and the manifest has a third layer of HMACing due to protocol issues, so it should be impossible to just change stuff even if you break the ciphertext authentication.
Moreover, I'm pretty sure Borg's developers are actively considering ChaPoly for the future.
Again: the team is actively considering using AES-GCM and ChaPoly in the future. I don't see anything intrinsic to either that preempts their use in Borg.
IIRC AES-GCM was kinda low on the list with a preference for just using Chapoly, because Chapoly just works and is also secure on any processor, unlike AES-GCM, which is very nasty to implement without hardware support for the arithmetic over GF(2^128).
Sadly, the only thing I've found so far that works at all at those scales is Bacula, and that is file-based -- i.e., if you have a gigabyte file that changes by one byte, it backs up the whole gigabyte again. Not ideal.
But I'm not very happy about it, because it's insanely overcomplicated for no really good reason, and because of it being file-based.
https://www.backblaze.com/blog/lto-versus-cloud-storage/
If you're interested, I'm doing experiments with another site with 500T to backup, where I added sampling and sharding to HashBackup (I'm the author).
Sampling allows you to do faster simulated backups to determine the best backup parameters to use. In his case, we determine that a very large block size - 64M - was the best way to backup his data.
Sharding automatically partitions the filesystem so multiple backup can run simultaneously to get backup speed in the 250-400 MB/s range.
It's more at the proof of concept stage, but having another large site to work with would be fantastic! A couple of the larger sites using HashBackup are EURAC (European Research Center) and HMDC (Harvard MIT Data Center)
Would it possible to split those bigger files before doing backup (log files with log splitters, sql databases with incremental specialised backup tools like e.g. xtrabackup), or are these e.g. image files with different versions?
I really like the concept of byte level deduplication, but have often thought the price you pay for the space reduction might not be worth it considering today's network speed as well as storage sizes and prices.
Would be interesting to hear your experience concerning this!
If the backup service is using file timestamps as the only key to refresh then it would have to store the last modification date and a good hash of each block, and when the date of the file being mirrored is updated scan each block to see if there is a change there and update the stored hash and timestamps accordingly. This would need to be orchestrated to reduce the risk of temporary corruption if there are several updates and the rechecking process coincides with a backup sweep (i.e. make sure you don't present updated dates for any block until you can present them for all needed).
For restoration, you either manually concatenate the parts or have an overlay filesystem that operates in reverse: showing the smaller block files as a single large unit.
You'd have to very thoroughly test the overlays and their interaction with the backup service before risking it on important data, so it might not be something you would genuinely consider...
Restic falls over as soon as you cross 2TB sizes, and prune operations are painfully slow. Backing up 8TB took 4 weeks with Restic, and i gave up waiting for prune to finish. It was crossing the 24 hour mark, making it unsuitable for daily backups.
set -euo Pipefail
if "${DEBUG:-false}”; then set -x; fi
The -o Pipefail is irrelevant in this case, as you're not using any.. but better it's there in case you ever extend the script.As others have pointed out: readability comes right after correctness in shell script, do keeping it simple as you have done is always a good idea.
Do you have a link to a good explanation of why you'd set those? The "set -e" seems to indicate that the script would fail immediately if the borg backup job fails, which would prevent the Zabbix "send" command from alerting the monitoring system.
Thank you for your feedback, it's given me more to learn and think about!
you're right that just putting the set options on top of your file would be a bad idea. you need to write your script to actually account for these errors, making possible exitcodes entirely transparent
1 #!/bin/bash
2
3 set -euo pipefail
4 if ${DEBUG:-false}; then set -x; fi
5
6 function random_exitcode(){
7 return $(( RANDOM % 5))
8 }
9
10 function handle_error(){
11 case "$?" in
12 1) echo "handling exitcode 1!";;
13 2) echo "handling exitcode 2!!";;
14 3) echo "handling exitcode 3!!!";;
15 *) echo "encountered unexpected exitcode: $?"; exit 2;;
16 esac
17 }
18
19 random_exitcode || handle_error
or, if you don't like functions: random_exitcode || LAST_EXITCODE=$?
case "${LAST_EXITCODE:=0}" in ...
[0] https://coderwall.com/p/fkfaqq/safer-bash-scripts-with-set-e...If I need to force a sync over the internet, there is a small shell script that would rsync the backup folder to the home nas.
This allows me to run hourly backups even if there is no network connection (or limited bandwidth), and then just auto sync when I get home, or force a backup when I get a good connection.
I do sort of wish that Visual Studio Code had the "auto-fixes" already built for shellcheck's error messages. I should probably create those and make them available somewhere.
I would not trust my backups to your service yet, just because of the "this is a prototype" language. My immediate thought is, "this seems great, I'll have to come back and check it out once it's more of a real business". But therein lies the rub, I think: what will drive me back to check it out later? There doesn't seem to be a mailing list to sign up for. Maybe you'll hit the front page of HN with a full launch later, and I'll see it that way, which would be great, but maybe not!
In any case, nice work!
slack: https://baxx.dev/join/slack
google groups: https://baxx.dev/join/groups * get notified if the file is too small
* get notified if the file is too oldThe repo on GitHub doesn't have a license AFAICS.
Edit: still love idea and execution, thanks for sharing!
Which is likely still less readable than something whose sole purpose is to present to humans, such as HTML/CSS.
> Subscription: 5E per Month
> ...
> I decided to charge 5$ (0.1$ trial)
So which is it? $ or E? And is E supposed to be €? Also, in English the currency marker goes to the left of the number.
But in the majority of european countries where they use euros, the symbol is placed to the right
http://publications.europa.eu/code/en/en-370303.htm#position
I always thought this depended on the currency; like $5 vs 5€.
> When writing currency amounts, the location of the symbol varies by currency. Many currencies in the English-speaking world and Latin America place it before the amount (e.g., R$50,00). The Cape Verdean escudo places its symbol in the decimal separator position (i.e., 20$00). In many European countries such as France, Germany, Greece, Scandinavian countries, the symbol is usually placed after the amount (e.g., 20,50 €).
Edit:
Excerpt from Wikipedia's article for the Euro symbol [1]:
> Placement of the sign also varies. Countries have generated varying conventions or sustained those of their former currencies. For example, in Ireland and the Netherlands, where previous currency signs (£ and ƒ, respectively) were placed before the figure, the euro sign is universally placed in the same position. In many other countries, including France, Belgium, Germany, Italy, Spain, Latvia and Lithuania, an amount such as €3.50 is usually written as 3,50 € instead, largely in accordance with conventions for previous currencies.
> The European Union did indeed usher a guideline on the use of the euro sign, stating it should be placed in front of the amount without any space in English, but after the amount in most other languages.
Wikipedia is wrong if it says that the placement is based on currency. Wikipedia can also be edited by anyone with a pulse, so there you go.
Here are some official government resources on currency code placement in English.
http://publications.europa.eu/code/en/en-370303.htm#position
https://www.btb.termiumplus.gc.ca/tpv2guides/guides/wrtps/in...
> The European Union did indeed usher a guideline on the use of the euro sign, stating it should be placed in front of the amount without any space in English
So...what you're saying is that I'm right. I'm very confused now.
I'm attacking your pedantic attempt to reduce the matter to an inflexible rule.
Some English speaking countries prefix the currency symbol, so if let's say my audience is from Australia, and I want to make the gesture of adhering to their customs, I will prefix it. However, in other (most) scenarios, I'll write it the way I write it 99% of the time, in the postfix position. And that's ok. Everyone doesn't need to use the same customs, words, grammar, or even language. Regarding these things, the only thing we care about, at the end of the day, is being able to understand the other, and the other being able to understand you.
As you've noticed, I said _some_ English speaking countries. That's because there are a lot of people, who even though use a lot of English, only a fraction of that is with people from US, Canada, Australia or Great Britain, so the customs of English users from those countries become irrelevant.
There's customs and there's language. This is an obvious example where one does not have to entail the other.
also fixed the $ to E in the blog, thanks for the heads up
It does for me?
Too many times some developer will pick a common word to launch their service/product thus making it impossible to search for it.
Your name (a) is a single syllable
(b) almost hints at what it does in the name (Baxxup?)
(c) is not easily confused with other products
(d) Your closest competitor from a google SEO standpoint is a tanning salon (single word google)
Great job! I wish I was creative like that. Usually my crap looks like: sbs (server backup script) Works for me but my co-workers hate me.
others I chose are horrible, such as https://scrambled-eggs.xyz which is effectively unsearchable, though you could say it was by design haha
started playing with https://github.com/rivo/tview last week, but didnt have much time.
Agreed. This:
"The way you register is through ssh, just `ssh register@ui.baxx.dev`"
Is very cool and makes me happy.
[0]https://github.com/jackdoe/baxx/blob/3d2f6f014e5def3b5c35496...
[1]https://github.com/jackdoe/baxx/blob/3d2f6f014e5def3b5c35496...
Another down-to-earth, unix-oriented, fair-deal style "cloud" service worth investigating is rsync.net. I've been using it to sync data between my devices, and I like the way it leaves the user in control.
I have not tried Baxx yet, but it seems to be a project in a similar spirit, and I am happy to see more investment in that kind of future.
Great execution!
I get the point, but the "without a website" claim is kind of a cutesy attention-grabbing prevarication.
I'm usually pretty tolerant of those, but personally find that here it mixes poorly with the idea of a backup service. From such a service, I'd prefer to see total honesty, and a sort of buttoned-down, almost military humorlessness/seriousness.
It roughly akin to me saying "my Golang on Linux program is 100% C code free - no C is used in the execution of my code!" .. well, clearly that statement would be false.
However, it's perfectly legitimate to say my project is C code free - and that's what this project is saying. It's saying the project is web-tech free. There just happens to be an off-the-shelf execution environment which uses web-tech.
But that's not the headline. The headline was instead that it's a service "without a website". But there is absolutely, positively, no-nitpicking-involved an HTTPS-accessible website for the service, set up by the service's own creators.
thanks a lot for the feedback, much appreciated!
What additional value can you add?
* i have actually implemented this kinds of systems multiple times and always failed at the 'who watches the watchers' and usually had broken backups (as in i usually needed them few years after the system is written and by then something is always wrong)
* in the future it will use contextual bandits to probe for broken backups (such as: hey, is file XYZ with size X uploaded at Y weird?) and learn from all the customers triggering the AI flywheel, more backups more data, better contextual bandits, better service, more customers (the idea is to use bandits[or something else with exploration factor] for probability of "good" notification)
so in few words: value = out of the box alerting + nice api + machine learning (not implemented yet) + good price (no cloud)
If you can do it at a price competitive to what I can buy storage (and everything that comes along with housing it) and worry about it then it's one of the few things I'm happy enough to have a managed service for.
edit for a bit more context: I do a lot of video work so my storage needs are constantly growing.
Put another way, do think there is ever any amount of CSS that is not superfluous?
CSS can be non-superfluous in documents, like when formatting text or a table.
I'd like to focus on the fact that this required orders of magnitude less effort than designing a web site, and most importantly – the presentation did not suffer. There were no compromises to be made.
Some have mentioned that this lacks functionality, but I'm having a hard time imagining what that might be. Maybe mobile-readiness, in that it's preformatted. Considering that this is a Unix-y tool for power-users, I think it's ok to expect them to read this on a PC.
I'm all for reducing complexity, but at this extent your are removing pretty huge pieces of functionality for extremely marginal streamlining.
It's a gimmick
HTML isn't a technical requirement, but they do say their goal is to sell the service.
I'm not sure though that a txt document is "simpler" than a bare (no js, not images) html document; on one side you can read it with something as simple as curl, on the other the txt didn't render properly on my phone.
to me the `yet another bootstrap css SaaS with three tier pricing, faded monotone brand logos and circlular portrait photos of the “team”` is just depressing
1. If I would send my backups to some SASS there is no way I would do that without encryption.
2. I like to backup my filesystem and not just the files in it (to make sure I've got everything and make restoring easy). Currently, I just dd my block devices, but I am sure that could be optimized to not upload a complete image every time.
Sure, I could solve those two problems myself, but then I could also use AWS S3/Backblaze ;-)
That's actually not a good idea at all.
It's very fragile, the slightest problem and you may lose the entire backup.
Silent corruption may get backed up for months and you won't notice until it's too late.
Doing a restore means having space for the full file, just to restore a single small file.
If you don't know when your file changed you may have to do that multiple times instead of just being able to check when that single file changed.
In short: don't do that. Copy the full directory structure, sure. But not the disk image.
Fragile seems to be a matter of perspective. I do a dd block-device backup before doing anything "risky" with a machine (i.e. upgrading major OS version, etc), because it is by far the most bullet-proof / fail-proof way to go back in time.
Silent corruption might actually be a problem, but that is something you won't solve by backing up files and directories instead of block devices. After all, a block devices backup just contains more information than the files and directories backup. The only advantage of filesystem backups is that you can easier validate if a file should have changed, but even if you detect that it changed even if it shouldn't have, you still need some kind of checksum or so to find out which version is correct.
On the other hand, if you back up the files and directories you have to care about the filesystem type and if there are special types like links, devices nodes and the like. On that side, I had enough unpleasant experiences that I am trying to avoid that trouble.
To restore single files I can simply mount the image on a loop device, so no problem there.
considering in many cases you want to copy live files, you have to copy them to place you can unmount, so now you have 2 problems
if you dont gzip the dd you risk pure cosmic microwave background bitflips and no checksums, and they can corrupt superblocks and etc
if you gzip it (to have crc sums) you cant restore partial files from it so easily unless you use squashfs and significantly complicate the process
so tar and gzip is more robust (compared to the dd) just because of the crc (of gzip not of tar), size and portability
cat file | encrypt -k .pass | curl --data-binary
cat file | encrypt -k .pass > file.enc && curl -T file.enc https://baxx.dev/io/$BAXX_TOKEN/file
(encrypt: https://github.com/jackdoe/updown/blob/master/cmd/encrypt/ma...)offtopic: it is a bit annoying that cat | curl makes curl buffed the whole standard input, so for big files you have to either you have to do -H "Transfer-Encoding: chunked" or use -T file
2) dd block devices is very fragile, its better if you just tar --exclude/dev ... (https://help.ubuntu.com/community/BackupYourSystem/TAR - Alternate backup)
Cost. Currently I pay $10/month to CrashPlan to store 2TB a month. I think your plan is to simply pass through the glacier price, but even that would be much higher cost I think?
Large files. Because I only have finite monthly bandwidth and a limited upload speed, it would be good to do incremental backup of large binaries, only changing the parts that changed.
Basically, this is a toy. Use it as a toy, not as a service where you'd actually put your company files.
That said, good job!
Yes... until you want to take anything back out - then you're going to have to wait, and you're going to pay a lot more than €5
in the future my plan is to do 5€/m for 1TB which is pretty much glacier price
Don't need a backup solution right now, but I keep it in mind.
[1] https://github.com/jackdoe/baxx/blob/master/infra-and-pricin...
FWIW, corporate taxes in Germany aren't 50%, but 15%. What is close to 50% is personal income tax + health care if you pass a certain threshold. However, even if you're running a just sole tradership (Freier Beruf or Gewerbe), this would only be your personal income - all business expenses are not counted towards your personal income tax base and you get the VAT you paid back.
But corporate taxes in DE are about 30%, not 15. Don’t forget your Gewerbesteuer.
If that doesn’t strike you as unfair I’m not sure what would.
https://en.wikipedia.org/wiki/Taxation_in_Germany#Corporatio...
Machine to Machine payments: Say you deploy a bitcoin miner to stranded gas. It starts to mine coins and can start paying you to back up it's data without ever needing to create a Paypal account.
The instructions on the main page would change from:
ssh register@ui.baxx.dev
to something like curl https://ui.baxx.dev/ssh_keys -o /tmp/baxx_ssh_keys
cat /tmp/baxx_ssh_keys
cat /tmp/baxx_ssh_keys >> ~/.ssh/known_hosts
rm /tmp/baxx_ssh_keys
ssh register@ui.baxx.dev
or in one line: curl https://ssh_key.baxx.dev >> ~/.ssh/known_hosts && ssh register@ui.baxx.devEdit: I thought you meant that, given the command, most would still not do it.
it is super cute though, I still get excited every time I try the registration flow.
the registration api is quite easy (https://baxx.dev/help/register)
curl -d '{"email":"your.email@example.com", "password":"mickey mouse"}' \
https://baxx.dev/register
I will probably do make register.sh endpoint that returns a bash script that runs local dialog like:
curl https://baxx.dev/register.sh | sh
(just idea, not implemented yet)this will be fun, making portable tty dialogs :D
i think the problem needs exploration, i have been bitten by bad anomaly detection and setting up alerts post-moretm after losing data one too many times :)
i imagine creating a bunch of features such as:
* file extension
* time of upload
* delta from previous version
* size difference from other files in the directory
* etc..
and having enough customers we should be able to do good-ish predictionwhen a file version is "weird", so we could send notifications such as:
> hey is this file ok?
with some exploration factor, even using UCB will(should) be better than nothing
I plan to use vowpal wabbit's contextual bandits with only 2 actions, "send notification"/"dont send notification" given all the context
this of course will be just extra on top of the manual alert rules, but hopefully it will save some data :)
If the users agree we could also publish anonymized datasets with labeled data (such as: given context alert was sent: was [good]/[bad])
Which will be awesome.
which uploads: /EAAAAAEQJUAYSU5PR3WO7IKOX2ZPM5NSNIALFYJE5MMDJYNMNUZEA===
1: https://github.com/jackdoe/baxx/blob/master/infra-and-pricin...
If I'm going anywhere with my laptop and want access to my data faster than internet speeds I'll preload everything with 'unison' onto the laptop and temporarily have my own Dropbox clone while mobile.
Using those open source tools lets me do basically the same thing as these services while keeping full control of my data.