HNHacker News
TopNewBestAskShowJobs

melor

68 karma · joined April 19, 2016

submissionscomments
melor··on PGHoard – PostgreSQL backup and restore service for cloud object storages
CPU usage varies based the selected compression algorithm and level used. Snappy and LZMA area available now. Compression is native code. There are some newer interesting algorithms (zstd/lz4) that we are looking into adding.
melor··on PGHoard – PostgreSQL backup and restore service for cloud object storages
One of the pghoard developers here. We developed pghoard for our use case (https://aiven.io):

* Optimizing for roll forward upgrades in a fully automated cloud environment * Streaming: encryption and compression on the fly for the backup streams without creating temp files on disk * Solid object storage support (AWS/GCP/Azure) * Survive over various glitches like faulty networks, processes getting restarted, etc.

Restore speed is very important for us and pghoard is pretty nice in that respect, e.g. 2.5 terabytes restored from an S3 bucket to an AWS i3.8xlarge in half an hour (1.5 gigabytes per second avg). This means hitting all of cpu/disk/network very hard, but at restore time there's not typically much else to do with them.

melor··on Kafka Benchmark Results on AWS, Azure, Google Cloud and UpCloud
The next part in the series will include read-write benchmarking. Taking suggestions for other benchmark scenarios!
melor··on Ask HN: Were you able to mitigate the impact of the AWS us-east-1 incident? How?
We host a number of our customers' database systems on us-east-1.

What worked well for us (https://aiven.io):

- Architecturally relying only to a few cloud provider services (only need VMs, disk, object storage)

- Upfront investment on being able to move services from one region to another without downtime

- Pre-existing tooling for easily (manually) reconfiguring backup destinations on the fly

- Not running everything on just AWS

What did not work so well:

- Backups should automatically reroute to a secondary backup site on N consecutive failures

- Alert spam, need more aggregation

- New failure mode: extremely slow EBS access, some affected VMs were kinda working, but very slowly: need to create a separate alert trigger for this

melor··on Ask HN: Is S3 down?
Only limited impact to Aiven services due to service migration capability http://help.aiven.io/announcements/aiven-customer-notice-aws...
melor··on UpCloud: Better than digital ocean?
We provide UpCloud as one of the cloud options for our SaaS database/metrics/messaging offering at Aiven.io and have been extremely happy with their disk i/o performance.

Here's just a quick "hdparm -t" test I just ran on two random low-end nodes:

upcloud-de-fra: 1028 MB in 3.00 seconds = 342.12 MB/sec

aws-us-west-1: 58 MB in 3.02 seconds = 19.17 MB/sec

I would of course recommend everyone to benchmark their actual workload on each cloud option before making the decision.

melor··on PostgreSQL 9.6 Beta 1 Released
From the Release Notes:

Major enhancements in PostgreSQL 9.6 include:

Parallel sequential scans, joins and aggregates

Elimination of repetitive scanning of old data by autovacuum

Synchronous replication now allows multiple standby servers for increased reliability

Full-text search for phrases

Support for remote joins, sorts, and updates in postgres_fdw

Substantial performance improvements, especially in the area of improving scalability on many-CPU servers

melor··on PGHoard: Tools for making PostgreSQL backups to cloud object storages
A replication slot can be used by defining it in the pghoard.json configuration. However, the slot needs to be created (and removed after no longer needed, important!) manually. We've been planning to add more automatic replication slot management to PGHoard.
melor··on PGHoard: Tools for making PostgreSQL backups to cloud object storages
Both do mostly the same thing with some differences. The biggest difference currently could be that WAL-E uses the PostgreSQL "archive_command" to send incremental backups (WAL files) in complete 16 megabyte chunks, whereas PGHoard uses real-time streaming with "pg_receivexlog", making the data loss window much smaller in case of a disaster.
melor··on PGHoard: Tools for making PostgreSQL backups to cloud object storages
Currently S3 (AWS + compatible), Google Cloud, OpenStack Swift, Azure (experimental), local disk and Ceph (via S3 or Swift) are supported. More can be added quite easily as the object storage logic is behind an extendable interface.

Which vendor neutral protocol are you interested in using?

melor··on PGHoard: Tools for making PostgreSQL backups to cloud object storages
Takes care of realtime WAL streaming, compression, encryption, restoration and backup expiration among other things. Open Source and written in Python.
melor··on Ask HN: Is this a Python bug?
It is a feature http://stackoverflow.com/questions/26595895/return-and-yield...
melor··on Lesser-Known Python Data Analysis Libraries
I use histogram.py from https://github.com/bitly/data_hacks all the time...
melor··on Google.com partially dangerous
Sure is dangerous sometimes for my work productivity...
melor··on Monitoring, metrics collection and visualization using InfluxDB and Grafana
Note that if you are planning to use the latest InfluxDB 0.12 version with Grafana, you must go with the new Grafana 3.0 beta version as the older Grafana versions have issues with the newer InfluxDB versions: https://github.com/grafana/grafana/blob/master/CHANGELOG.md#...