Hacking on PostgreSQL is hard
rhaas.blogspot.com
rhaas.blogspot.com
https://news.ycombinator.com/item?id=18442941
To quote part of it
> Oracle Database 12.2.
> It is close to 25 million lines of C code.
> What an unimaginable horror! You can't change a single line of code in the product without breaking 1000s of existing tests. Generations of programmers have worked on that code under difficult deadlines and filled the code with all kinds of crap.
> Very complex pieces of logic, memory management, context switching, etc. are all held together with thousands of flags. The whole code is ridden with mysterious macros that one cannot decipher without picking a notebook and expanding relevant pats of the macros by hand. It can take a day to two days to really understand what a macro does.
> Sometimes one needs to understand the values and the effects of 20 different flag to predict how the code would behave in different situations. Sometimes 100s too! I am not exaggerating.
> The only reason why this product is still surviving and still works is due to literally millions of tests!
I think it is one of the largest impact problem in software engineering if it can be improved. Maybe a way to restrict flag interaction and reduce support and test matrix as a result.
One of their strategies was to drop the macro soup and simply program against the "libc we would like to have", and then add compatibility shims to materialise their ideal libc instead of conditional compilation at the point of use.
Personally I consider this a good thing. It's a sign of a really mature codebase where lots of edge cases are known + accounted for.
Even if the underlying code was really well written, simply the number of edge cases hamstring any "quick hacks".
Complex, runs reliably, easy to hack - Pick two
You should be able to work on software because you understand how it works and what the ramifications of a given change are. Tests and code reviews provide redundancy. But here, they aren't providing redundancy, they're bearing the load.
What provides redundancy if tests are missing, broken, or misinterpreted? Have you ever fixed a bug, gone to write a test for it - and found the test already exists but passed spuriously?
It's almost a mark of success of the project. There is obviously a lot of dedication too.
- Postgres documentation is one of the well maintained database documentations. This also means that developers, committers ensure changes to documentations for every relevant patch.
- talk about bugs in postgres compared to MySQl or Oracle or etc databases. Nugs are comparatively lesser or generally rare even if you are supporting postgres services as a vendor with lots of customer. the reason is the efforts involved by a strong team of developers in not accepting anything and everything, there are strict best practices, reviews, discussions, tests, and a lot more that makes it difficult to pass to a release.
- ultimately, more easy is the acceptance of a patch, more the number of bugs.
I love Postgres the way it is today and it still is the dbms of the year and developers most loved database.
I wish we have more Contributors committers, developers and also users and companies supporting Postgres so that the time to push a feature gets more faster and reasonable easier with more support.
Seems like a reason to celebrate the open source model, and specifically here on how to do things better. Not to detract from universal issues for any project on maintainer availability. But, imagine a non oss database vendor with that degree of transparency or velocity, i can’t think of any that are doing anything close unless they got popped on a remote cve, aka prioritized above features or politics on a corporate dev sprint. Aka all software has bugs, it’s about how fast things are fixed, and in the context of oss imho fostering evolution among a diverse set of maintainers and use cases seems to be a better way.
As another example of that, ‘twas a PostgreSQL hacker at MS, that prevented Libxz from going wide because of caring due to perf regression and doing the analysis.
There are some deep lessons about programming in this Factorio Friday Facts:
https://factorio.com/blog/post/fff-366
and I wonder if postgres doesn't look like fig.1 from this blog post, before the refactoring.
However back then MySQL seemed like it went out of it's way to corrupt your data. The only "bonus" is, it did it all silently, so nobody ever noticed until they went looking. With MariaDB(the successor) it's pretty rare that it silently corrupts your data these days.
(Most people wanted speed.)
Trying to find any good information on how to go about this proved super difficult. Well, I wasn't having much luck and just gave up.
Some of what we did might translate well to PostgreSQL, some of it won't, and much of it is probably too expensive and/or too much work. (Then again, it's work that doesn't require an inflight rocket surgeon to accomplish, which means it's doable by a much larger population of developers.)
- we've long had volunteer (and later, employee) "sheriffs" that monitor CI, know how to back things out, and over time get better at recognizing the sorts of problems that come up.
- For slow or expensive tests that don't run on every commit, they'll also take care of "backfilling" test jobs to narrow down which patch or patch stack most likely caused a problem.
- As with most CI systems, there's a staging area that gets a decent level of testing before changes are merged into the main development line.
- Feature gates for larger changes, so things can land in the mainline and be worked on there for a while, with CI regression tests running both with the feature enabled and disabled (as well as feature-specific tests when it's enabled). Good for reducing bit rot.
- Extensive fuzz testing. This would probably need to be specialized to a DB environment, since they're obviously very stateful. Various forms of snapshotting are good. For the browser (and especially the JavaScript engine I work on), it's hard to overstate just how useful this is. I would guess it could work quite well for a DB engine too.
- Lots of resources poured into test machines. With enough machines, good sheriffs, and a rich test suite, test latency doesn't matter all that much. You may not know about the problems for a day or three, but if you can depend on either getting backed out or your feature re-disabled, then you can fire and forget with no guilt. (Ok, the sheriffs will start getting snippy if you bounce a landing too many times, as is their prerogative.)
I'm guessing DB development and testing has tons of idiosyncratic difficulties, but it all sounds so familiar that I think many of the same approaches could work. The inevitable "turning the buildfarm red" should not lead to "spend[ing] the afternoon, or the evening, fixing it..." Complex software is a different beast, and it's unrealistic to expect to be able to break all features down into simple obvious changes. There's just too much going on.
(You still can't handle just anyone committing just anything at any time, though. There will always be a rate of breakage introduction that your system can handle, and it's not hard to go over it.)
Such tests could perhaps be database-agnostic to a degree, verifying that the database behaves according to the SQL-standard?
I was thinking more like lots of concurrent operations, and backups/restores (again concurrent with other DB traffic), and replication, and incremental operations, and failover, and error handling in general. All while varying things that feed into scheduling, etc. Anything nondeterministic is good, though that's not all of it. (Annoyingly, that means that failures will quite often be intermittent, which is a whole can of worms of its own.)
for a lot of stuff you don't need it, sometimes even don't want it
In Rust it's much easier to create robust and performant abstractions that are near impossible to misuse, eliminating tons of potential bugs right away.
In Rust you don't need 10 years of experience and staring with a microscope and at every line to ensure it doesn't introduce issues. In Rust trivial changes and indeed trivial, so you have more time to think about actually difficult parts.
In Rust contributors don't need to learn custom implementation for collections, string and other basic primitives on every project.
And most of all, people actually want to learn and work with Rust, so your contributor pool is expanding, not shrinking.
Timestamps get relativized when we re-up a post [1]. That's an artifact of HN's re-upping system [2].
[1] https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...
[2] https://news.ycombinator.com/item?id=26998308
Sorry for the confusion—I know it's weird but the alternative turns out to be even more confusing and we've never figured out how to square that circle!
It also sounds like internal documentation may be lacking, which isn't surprising.
Obviously easier said than done, but the way he presents things it seems kind of like a push and pray environment.
Any comparison is basic at best.
eg look at the pg docs re: at least 5 different types of indices are implemented https://www.postgresql.org/docs/current/indexes.html
Everything is like that.
(I'm not bashing sqlite, it's great software, but it has a tiny fraction of the product footprint that pg has.)
And that's just core, not even including built-in extensions like bloom.