HNHacker News
TopNewBestAskShowJobs

richvdh

4 karma · joined May 10, 2016

submissionscomments
richvdh··on We recovered from nightmare Postgres corruption on the matrix.org homeserver
We certainly haven't ruled out Postgres or kernel bugs here.

> If you ran into a problem like this on ZFS, for example, you'd have very high confidence about whether the disk was at fault

Would we, though? I'll admit to not being that familiar with ZFS's internals, but I'd be a bit surprised if its checksums can detect lost writes. More generally, I'm not entirely sure how practical it would be to add verification at all layers of the stack, as you seem to be suggesting.

We'd certainly be open to considering ZFS in future if it can help track down this sort of problem.

richvdh··on We recovered from nightmare Postgres corruption on the matrix.org homeserver
- We've used pg_repack in the past, though I'm 90% sure we didn't use it in the timeframe in which this corruption must have happened. Anyway, we'll look into this further: thanks for the suggestion.

- Yes, we've done OS upgrades. Back in 2021, our DB servers were on Debian Buster; they are now on Bookworm. We're aware of the problems caused by collation changes, and indeed that was one of the first things we checked; but we're careful to use the C locale for our database, so believe we're safe on that front.

- For the example we gave (index page 192904826, referencing heap page 264925234), the index page LSN is DB4A3/C73ED0C0, and the heap page LSN is DB4FA/4CAAB9D8, so the index page was written shortly before the heap page. The blog post shows the output of SELECT * FROM heap_page_items for the heap page: it looks like a regular empty page to me.