Why we spent the last month eliminating PostgreSQL subtransactions
about.gitlab.com
about.gitlab.com
What is about the other end? Why does vacuum need to be a long running transaction and cannot be cut into shorter transactions ?
That this would have been an awesome opportunity for gitlab to show how OSS they are and fund a PostgreSQL developer to allow gitlab to design these boundary pushing designs.
As you can read in the article they also hired Nikolay Samokhvalov (a PostgreSQL contributor) to take a look at the problem. But as you can also read the solution wasn't with PostgreSQL. The solution was modifying their application.
They're also submitting patches to improve subtransaction support in future versions of PostgreSQL (likely v15). This is them doing the right thing for the project.
If this is true, my criticism is even more prescient, as it was merely in the writing of the article where this could have been corrected. It sounds like gitlab already did the thing I recommended, they just didn't highlight it properly.
This isn't directed at you, but HN does not take criticism or alternate viewpoints well. The article is great on many levels. I gave feedback on how they could improve, the real damage is done when someone doesn't care enough to comment.
The conclusion section is lacking in meta analysis and why gitlab. Given the LoE that went into this, gitlab missed an opportunity here. If I were the senior manager on this project, I would want a retrospective on why this was hit in the first place and not as a form of someone to blame, but what about the design, the technology selection, the monitoring, testing, etc. contributed to the outcome we saw.
This was the key takeaway for me.
SubtransControlLock indicates that the query is waiting for PostgreSQL to load subtransaction data from disk into shared memory.
I felt the article fell down for two reasons:
(1) It didn't really articulate the need for transactions in the first place (database integrity). Nor did it discuss the implications on integrity with this change.(2) It didn't articulate the possibilities of other architectures (pushing to a read cache other than PostgresSQL like Cassandra).
I got the feeling they were really pushing PostgresSQL to its limits in their cluster with their load - and it was time to consider another design.
2) They have a very specific, narrow band problem which they solved two ways. What you’re suggesting is several orders of magnitude more work and would likely take years.
They are definitely not pushing the limits of PostgreSQL as a whole, they just hit an edge case with a not heavily used feature. Sub-transactions aren't overly common in most PostgreSQL workloads I have encountered, especially not coupled with long-running transactions.
Great write-up though! I do hope patches to make the sub-transaction cache configurable land, while it's not something I have ran into it would be a great lever to have if I do.
It says they had sub-transactions, not what caused them. Nothing in the article gives concrete details on specific code or logic that was contributing to the issue. The entire "Why would you use a SAVEPOINT?" should have been replaced with "Why we were using SAVEPOINT".
Personally, I was hoping for the post to dig into a specific example and explain why transactions were used in the first place (justification) or why removing transactions or sub-transactions had no negative impact on data integrity.