Introducing WAL-G: Faster Disaster Recovery for Postgres
citusdata.com
citusdata.com
We've been using WAL-E for years and this looks like a big improvement. The steady, high throughput is a big deal – our prod base backups take 36 hours to restore, so if the recovery speed improvements are as advertised, that's a big win. In the kind of situation in which we'd be using these, the difference between 9 hours and 36 hours is major.
Also, the quality of life improvements are great. Despite deploying WAL-E for years, we _still_ have problems with python, pip, dependencies, etc, so the switch to go is a welcome one. The backup_label issue has bitten us a half dozen times, and every time it's very scary for whoever is on-call. (The right thing to do is to rm a file in the database's main folder, so it's appropriately terrifying.) So switching to the new non-exclusive backups will also be great.
We're on 9.5 at the moment but will be upgrading to 10 after it comes out. Looking forward to testing this out. Awesome work!
> our prod base backups take 36 hours to restore, so if the recovery speed improvements are as advertised, that's a big win.
Yes, if you attach 16TB of storage to each instance, your back-up restores may take a while. :))
We use a combination of streaming replication and wal-e backups. A separate machine performs multiple restores per hour and verifies restores work ok and that the data is recent.
I'd say the differences between WAL-G and Barman are similar to WAL-E and Barman, which comes up relatively frequently. https://news.ycombinator.com/item?id=13573481
In summary, WAL-E is simpler program all around that focuses on cloud storage, barman does more around inventories of backups and file-based backups and configuring Postgres. There are integrative downsides to its span. WAL-E also happens to predate Barman.
Also, what an impressive project to have on the resume as a college intern. I don't think many interns get to tackle something so meaningful.
is google cloud storage on the roadmap?
Minio has an interesting feature where it can be a "gateway" to other cloud storage. Google Cloud Storage is one of their specific examples:
https://docs.minio.io/docs/minio-gateway-for-gcs
So WAL-G would talk to Minio, and Minio would transparently proxy that to GCS.
I assume I am switching from WAL-E to WAL-G for more perf. But WAL-E speaks GCS. If WAL-G needs an extra hop to do so, may lose some of the point of it..
That being said, the Minio team seem pretty good with writing performance optimised code. Frank Wessels (on Minio team), has been writing articles about Go assembler and other Go optimisation things recently. eg:
• https://blog.minio.io/accelerating-blake2b-by-4x-using-simd-...
• https://blog.minio.io/golang-internals-part-2-nice-benefits-...
So the performance impact might not be such a problem. :)
Disclosure: I work on Google Cloud (so I'd love to see this tool point at GCS).
We're going to start rolling it out for Forks/point-in-time recoveries first, which present less risk to start. Later we'll explore either parallel restores from WAL-E and WAL-G or possibly just flip the switch based on the results.
On restoration there's really no risk to data. Further we page our on call for any issues that happen such as WAL not progressing, or servers not coming online out of restore.
For now, since I'm also on GCP, I'm using PGHoard: https://github.com/ohmu/pghoard
Both back up PG's WAL files (Write Ahead Log) and allow restoring your database state as it was at a specific time or after a specific transaction committed. This is known as point-in-time recovery (PITR) [0]
Users and admins make mistakes, and accidentally delete or overwrite data. With PITR you can restore in a new environment, just before the mistake occurred and recover the data from there.
[0] https://www.postgresql.org/docs/9.6/static/continuous-archiv...
However it might be interesting to stream WAL logs to e.g. AWS Kinesis....
I'm afraid to use a lot of storage for WAL segments that are mostly empty:
16 MB per segment x 60 minutes x 24 hours x 7 days = 161 GB/week
Does WAL-G/WAL-E compression help?
On a staging server with little activity, the compressed WAL-E wal files go as low as 1.9MB per 10 minutes. (~2GB/week)
The production server has files between 4 and 12MB per 1 minute or less. (~220GB/week)
WAL-E has a good `wal-e delete retain` command that removes older base backups and wal files.
Great job to everybody that's working on this. I'm looking forward to trying it out.
i have a huge ETL pipeline at twitch that relies heavily on wal-e, and installs it on worker nodes and things like that.
that being said, if WAL-G is faster, i don't care waht it is written in, and am happy to use it
Put differently to be a "successor" it needs to be a drop in replacement ;)
The RSA keys (or path to them) would be passed as environment variables. It would be a little easier to setup than gpg (especially for automatic backup restoration).
Fwiw, the GCS client for go (import "cloud.google.com/go/storage") is very straightforward. Though as others have pointed out, it might be worthwhile to just try to use minio-go if you want to gain Ceph as well.
Disclosure: I work on Google Cloud (and if the way is paved, we will contribute here; seems like a great project)
Good to see people sticking to the unix philosophy of doing one thing well and delegating other concerns - cat and lzop are both fine choices!