This seems like an instance where BPF could shine. If I had a large magic wand, I’d love to have distributed traces of database query execution with spans or events for significant events like the GIN pending list cleanup described in the article.
We did use DTrace to diagnose what was the root cause of a hanging insert (it was the same pending list cleanup as on this blog post). The whole talk we did was about DTracing PostgreSQL :)
The underlying cause was that `ext4` filesystem was journalled and the `fsync`s were waiting on `jbd2_log_wait_commit`, which offcpu sampling let me pick up.
I don't think I would have been able to trace the kernel call and link it all the way back to the application call without BPF (at least I wouldn't have been able to).
Fun fun :)