Show HN: PGBackup.com, Postgres backup as a service
pgbackup.com
pgbackup.com
This isn't an offsite backup. This an offsite replica that will faithfully replicate a "DELETE FROM account WHERE true".
Marrying that with a scheduled logical pg_dump and/or physical pg_basebackup that runs against the local copy (for performance and not impacting the master) would create a true offsite backup.
Until then, it's just one more place that will get wiped out when the source data is accidentally destroyed.
From a security standpoint the former would be a lot stronger though it would be a lot more resource intensive on your end.
I have a separate script that manages cleaning up of backups older than a certain age, etc.
Bonus points if the server that runs pg_dump and uploads to tarsnap has write-only access (to tarsnap), uses unique unguessable prefixes for the backups, and is distinct from the server that manages cleaning out old backups.
Course that's probably way overkill for a personal server :D
Edit: Not sure why'd you want unique unguessable prefixes, the data is accessed with the key. I actually encode timestamp information into my backup names for easy display.
The random prefix is to prevent overwriting an existing backup. See the sibling comment from cperciva and my reply to it.
Why? If you're worried about someone overwriting old backups, don't: You can't overwrite or modify archives in tarsnap. Once created they can only be read (if you have the read keys) or deleted (if you have the delete keys).
Nice. Some reason I thought a write key would be able to overwrite as well. If that's not possible then yes it's not necessary to have random prefixes. Fyi, most of my tarsnap usage is in "set it and forget" mode so been a while since I perused the docs.
The random prefixes would apply if you're using something that doesn't provide overwrite protection like pushing blobs directly to S3.
Have you looked at WAL-E (sp?) They transfer the wal files to S3, and works especially work if you have a high load database where pg_dump isn't feasible.
The product you should be building is replication (what you have actually built) and wal-e configured to backup to amazon s3 (where I provide the bucket).
you guarantee that the backups are happening, build in some intelligence to make sure catastrophic "DELETE" statements would trigger a backup first, etc.
Give me the ability to spawn a replica using the exact backup that I choose, etc.
This is something I could totally pay for!
If you do want this more out of the box the real options today are either built it yourself or choose a provider that's delivering it.
The same would happen if the WAL could not be retrieved in time (because wal storage full): PGBackup would start a new base_backup and send out pager alerts.
The base_backup+xlog backup is going to be our next feature. Backing those up to s3: possibly. Encrypted?
where would you backup if not s3 ? your own ? that would be a BIG oops...
It's technically rather challenging, but when enough users would ask for it, I know putting some resources into it might be viable.
Just stream the WAL in an encrypted format, and store it there? Easy enough. But that'll mean that it'll potentially take a long time to re-apply all those changes to a base backup.
Actually apply the changes, even though the database and WAL in encrypted? Unlikely, don't really see how that'd be possible. And even if, I'd bet that it'd be fairly easy to generate compression type attacks.
The backup wouldn't actually be particularly useful though. So I'm not really sure what the use-case here is.
> Why did you build PGBackup?
> We were tired of constantly configuring a postgres replica for each (small) project. And then being pretty unsure if the backup would still be up-to-date by the time the primary server crashes.
> When talking with other postgres users about backups, we also found that primary and backup servers often run in a single data center or at a single (budget) provider. In effect, exposing these users to non-neglible risk of loosing database+backup.
There's some - obvious - challenges in getting new users to trust you enough to send them their database copy. At the same time, people host very sensitive data at (virtual) budget servers without thinking twice. What do you think?
- We will try to add Point-in-time-recovery asap (saving base backups+xlogs). These would run of the replica, not putting load on your database server. As koolba correctly points out, this will make it more "backup" than "replica". Will have to figure out how this fits in the pricing model.
- I would personally love to offer better (guarantees of) encryption of the backups, ideally encrypt the data pages on the primary server. We would have to see how this would technically work.
- First couple of backups are currently replicating :)!!
Thanks, HN!
As far as I know, postgres authentication is rather simple and exposing the port does not easily add huge liability.
This is really a trade-off between ease-of-setup and better security. We're eager to talk to (a lot of) users to see what there current setup is and how PGBackup could add some value for them.
Famous last words of security on the internet.
Edit to add why: You'll have full reachability between pgbackup and your customer's servers without opening ports inbound on either side. Traffic is also encrypted :-)
Disclaimer: I am part of Wormhole Network.