Ransomware-resistant backups with duplicity and AWS S3
franzoni.eu
franzoni.eu
Your master access to S3 should never go into your servers. Create an IAM access with authorization to only PUT objects into S3.
> For the purpose we have, Governance mode is OK
Maybe not (?), since Governance mode allows for deletion of previous versions. One careless mistake handling your access key/secret and you're exposed to bye bye backups.
End note: this is still not enough. An attacker could compromise your backup script and wait for 40 days before locking yourself out of your data. When you try to recover a backup, you'll notice you have none.
Perhaps most attackers won't have the patience and will just forget about you, but who knows?
A naive protection from that would be to store at least one version of the backups forever. But we're still not covered, since the attacker could compromise that particular version and boom.
I can't think of a comprehensive, fully closed-loop solution right now...
The ZFS snapshots that you may configure on an rsync.net account are immutable.
There are no credentials that can be presented to destroy them or alter their rotation other than our account manager whose changes always pass by a set of human eyes. Which is to say, no 'zfs' commands are ever automated.
So the short answer is you simply point borg backup to rsync.net and configure some days/weeks/months of snapshots.
The long answer - if you're interested:
https://twitter.com/rsyncnet/status/1470669611200770048
... skip to 52:20 for "how to destroy an rsync.net account":
"... Another thing that lights up big and red on our screen is ... someone's got a big schedule of snapshots ... and then they change it to zero ... you've got seven days and four weeks and six months but we want to change that to zero days of snapshots. We see those things ... and we reach out to people."
The Borg pricing is quite competitive with S3 Glacier, although Deep Archive seems to have you beat. (Of course, Deep Archive loses badly if you actually read the data…)
Full price accounts have 7 days that don’t count against your quota.
It’s a tough market, and HN might include users or sysadmin types who might be interested in ZFS backups.
Whether advertisement violates the site rules, I don’t know.
With such access only accessible to a person with MFA, it should be pretty hard to accidentally leak the secret (MFA-authenticated sessions have tokens that are time-limited). If even those AWS permissions are compromised, yes, you could lose your backups, but I hope you don't use S3 as your sole backup storage.
I keep backups on both a local backup server and also in S3. If my AWS credentials are compromised, there should be no access to the local server. If someone accesses my local server, the access it has won't allow deleting S3 backups.
Since my backups are encrypted, the attacker would have to also have the key to the backups to read the content.
There are many reasons why recovery from backup could fail. Verifying recovery from backup is critical, so the solution here is to not delete you oldest backup until you have verified recovery from a newer version.
> Your master access to S3 should never go into your servers. Create an IAM access with authorization to only PUT objects into S3.
Isn't this precisely the approach I take in the article? But, you need to make sure the master access to S3 is not compromised as well. Probably obvious, but certainly not wrong.
> End note: this is still not enough. An attacker could compromise your backup script and wait for 40 days before locking yourself out of your data. When you try to recover a backup, you'll notice you have none.
I agree that's a reasonable attack scenario. Just as for any backup strategy, you'd need a monitoring strategy (i.e. try restoring a backup at least every X days), and that was left as an exercise at the end of the post.
I am using rdiff-backup over SSH to replace it now. This has been reliable so far but recovery times are extensive.
I believe you are correct and I believe that in my private correspondence with the duplicity maintainer (we sometimes sponsor duplicity development[1]) he sort of conceded that borg backup[2] is a better solution.
If the cloud platform you point your borg backups to can configure immutable snapshots (that is, they create, rotate, and destroy them) then a good solution would be using borg backup over SSH and configuring some of those snapshots[3].
[1] https://www.rsync.net/resources/notices/2007cb.html
So far Rustic has been my best bet. It seems to complete very large jobs with static and relatively low memory consumption. Enumeration takes about 30 minutes and a full backup nearly a month on the 8TB job (limited bandwidth available) but it reliably completes. Rustic also resumes interrupted jobs (e.g. due to reboot) in around 10-15 minutes on the 8TB volume and with it seems minimal rework. I'm sure there are other tools that handle this as well, but I've definitely gotten frustrated with finding them. I wish more backup tools gave you some kind of assurance in the marketing materials that they've been validated on, say, 10TB.
I'd wish that too. Also, I'd like to pay for a backup tool, since it's so critical, so that I can get a sort of support, but I have found issues with many of them.
I must say that, with my small population (N=1) for testing, it was hard for me to settle on a backup system. I tried duplicity, then I tried borg, then I used duplicati. I had considered attic and restic as well, but I don't remember why I didn't choose them right now.
I experienced issues with most of them and I reverted to duplicity. With borg, I had a persistent issues where backups would stop and said something like "backup destination is newer than source" (I forgot the exact message, but happened multiple times across versions). Duplicati seemed a very large and complex codebase, and periodically stopped working, mostly because dotnet runtime issues AFAICR.
Simply using object versioning, and an IAM role with no delete permission is sufficient.
The biggest risk with using AWS for backups, is being compromised via a high-privilege IAM role.
If a role has s3:* permission, it can just remove the locks and delete the file.
I worked for a company that had all their databases in AWS RDS and all their database backups as RDS backups stored in the same AWS account. A compromise of a single high-privilege account could cause that company to no longer have ANY data. I warned the higher-ups repeatedly that they needed more than one backup location. They didn't seem to care too much.
If it takes weeks to perform a restore then that should be a major red flag and be addressed.
Of course, if a company has the kind of security that allows them to be hit by ransomware, then they probably aren't diligent about their disaster recovery scenarios.
For my backups I use restic and sync the restic repository data to S3. Even if the source data is corrupted, I can always roll back to the set of S3 object versions from a particular time.
The downside to using S3 object versioning is that I haven't found any good tools to work with object versions across multiple objects.
For example, I need to prune old backups from my restic repository. To do that I have to delete object versions that are no longer current (i.e. the latest version of an object). To accomplish this I had to write a script that uses Boto3 (AWS SDK for Python) that lets me delete non-current object versions and deletes any delete-marker versions.
The code was pretty straightforward, but I wish there was a tool that made it easier to do that.
I have a feeling I set something like this up but it's been a while since I did it.
https://docs.aws.amazon.com/AmazonS3/latest/userguide/object...
I need to confirm that that my recent backups are valid before deleting old backups. Otherwise I might as well not have used S3 versioning.
At this point, AWS should offer versioned S3 backups accessible only offline or through customer support, enabled with a click of a button.
That said, AWS Backup is the answer to ransomware woes and it can't GA soon enough (3P solutions like Rubrik, Druva notwithstanding)
https://docs.aws.amazon.com/aws-backup/latest/devguide/s3-ba...
I cited rsync.net in my article, and I leverage AWS S3 as an example, because they're well-known players. I would be more hesitant to use a small and unknown (hence potentially unreliable) backup provider.
BorgBase itself doesn't have staff, but runs under Peakford.com, which offers a variety of hosting- and consulting offers. Myself, I mainly take care of our open source offerings and community.
A Customer of mine in the financial sector sent their backups to a third-party for independent verification quarterly. The third-party restored data into a clean system reconciled against the production system.
That would be the kind of auditing that would be more apt to detect the "low and slow" attack.
I'll update the article in the next hours to add some caveats.
Duplicity does NOT encrypt the names of your files.
Some might not care about this much, but for me I really don't like my encrypted backups containing so much metadata.
EDIT: FWIW I don't use the AWS backup options. I have separate offsite backup connections to cloud services as well as NAS.
The file names and content are encrypted.
No decent backup software will leave file names in plaintext.
> Duplicity does NOT encrypt the names of your files.
I, personally, do not care, but you're right. For the purpose of this article I could even leave encryption off, though. "I don't trust the backup destination" wouldn't be part of my threat model here.
If you want to limit data, you can create additional drives and mount them at the appropriate location (or change your application config to save to the auxiliary drives)
You can make them immutable for everyone if you wish and the only way to delete them is to close the AWS account.
I cannot think of a safest place for a backup than a bucket with object lock.