Looking Back at Postgres
arxiv.org
arxiv.org
My take is Postgres is extremely opinionated, but in such a way that it is beneficial for the long run.
PG adheres to the SQL standard far stricter than MySQL (i.e. select w/o order, ACID). From 2000-2010, it severely slowed the progress of features compared to MySQL (clustering was big topic for a while).
But the problems of hacks, speed and probably a significant one, management/owners, caught up with MySQL.
Feels like tortoise & the hare.
When I think of the term "opinionated" as it applies to software platforms, I think of things that have intended purposes in mind and optimize for the ease of those use cases, potentially to the detriment of people doing other things. It also suggests a "one way to do it" approach.
If that is the definition, I don't think Postgres fits. Postgres tries pretty hard to meet users where they are, and tackle uncommon or emerging use cases. Examples include JSON support, support for writing functions in various languages (javascript, python, perl, R, etc.), and C extensions that can hack the engine in innumerable ways.
Aside from being immediately useful just as a basic relational data store, my impression of PostgreSQL is that the "opinion" is apparent in the way that new things get done. There seems to be a culture of either implementing things well or not at all, and to avoid rushing things just because people want the new feature right now.
Could be I misinterpreted the OP's statement:
> few good people who had the right priorities and made good decisions without drama and without being overly opinionated
I attached the "opinionated" to the people as opposed to the software probably because of the way the sentence was structured.
Perhaps we're all agreeing and saying the same thing.
I really hope we pull off some of these software adages.
What was his first system?
However, if you look at the projects he’s worked on, he has a track record of writing database engines that end up in delivered database systems.
(I think what the paper refers to as "versioned data" is the "time-travel queries" stuff, not multi-version concurrency control which was added after Stonebraker's time.)
My guess is the audience here on HN is less likely to run into these features in their day-to-day lives. By my observations, the HN audience tends to be generalist developer types that will readily just hand off database specialization to the most popular ORM layer for their first class development environment. And I would agree they'll not encounter these things... and if they do, they'll largely be annoyances.
I work in "enterprise" type software and I see and do a lot of specifically database development (for PostgreSQL usually these days). If we consider PostgreSQL's type system as part of the object-relational feature set, I find I make use of that fairly frequently, for example compound types... also understanding that object-relational shouldn't be confused with "object-oriented". In my own work it's primarily the more advanced typing that shows up in both procedural code I write as well as queries in certain cases. Less so, though not completely absent is the inheritance feature; there are uses for it for sure, but a rather narrow set of cases (I mean outside of partitioning, which is the most common use case).
Anyway... I'm not so sure it's really a matter of downplayed as much as it is less generally understood as PostgreSQL has become more popular.
Like any principle or adage, if you are experienced enough to deeply understand the reason for it, you will know when it does or does not apply.
That's not what "an exception that proves the rule" means. It means "you may walk here" implies (proves the rule) there is a rule that you can't walk elsewhere.
In any event, plenty of second systems are fine. It's ridiculously outdated advice. Any time someone replaces their Flask web app with a go or rust version, qed.
In my experience, it means:
"This thing stands out to us as exceptional because it is exceptional; there is a rule which normally applies, but we tend not to notice the rule until an exception to the rule reminds us of the rule by contrast".
I dunno what your experience is like, but I have seen this scenario repeat itself many many times in many places: someone is rewriting something without fully understanding what the old thing did, blindly proclaiming that their rewrite won't have any of the problems of that old clunky thing, introducing new bugs and limitations in the process without adequately solving the problem the old thing was addressing.
All. The. Fucking. Time. Big companies are especially bad at this.
Seems to be a variant of the phenomenon second system syndrome is meant to describe.
That's not Second System Syndrome. That's Joel Spolsky's The Thing You Should Never Do.
Also, I don't think the advice is outdated at all. I've seen it happen multiple times. However, it describes a tendency to be on guard against. It's not some sort of law of the universe.
It's also not about rewriting in Go or Rust (you may be thinking of a completely different piece of advice about not doing ground up rewrites.) The second system effect is about being overconfident from the success of an initial system with limited scope, and being too ambitious in a second system.
Wait... are you suggesting that replacing a flask web app with a go or rust version is in support of or against second system syndrome? That sounds like a textbook example to me - I've seen so many second versions get built mainly for the sake of a new technology that miss a lot of the original requirements that the first version solved just fine.
You wont have that experience (or use it really) if you stop at your first system!
Seriously though, I agree, it's by definition that adages reveal truth. This one just feels so pessimistic and boring.
My version of this would be "the third system is the best one". Make it once, learn the space and the limitations of your solution. Make it second time, learn what isn't useful. Make it just right on the third try!
The righteous "this time we'll do it right" is a powerful drug, and fear of missing something can drive even the most experimented people to stupidity...
We have the replacement system almost ready but won't really be able to swap the thing until the project owner retires (yay politics!).
It's important but not time-critical, so if it dies before he goes we can just eat the downtime and take a couple of weeks to migrate.
https://www2.eecs.berkeley.edu/Pubs/TechRpts/1992/ERL-92-3.p...
The POSTGRES 4.2 tarball includes a bunch of his parallel query work, which you can find wrapped in #ifdef sequent. I assume that this was a precursor to XPRS, and it was ripped out of regular POSTGRES after that. I wonder if anyone knows what happened to the XPRS sources...
While reading, I also ended up highlighting and sharing some of my favorite parts on twitter: https://twitter.com/felixge/status/1272613965219139585
> 1 OPENING
>
> Postgres was Michael Stonebraker’s most ambitious project—his grand effort to build a one-size-fits-all database system.see https://en.wikipedia.org/wiki/PostgreSQL#History and https://en.wikipedia.org/wiki/QUEL_query_languages
But everything changed when someone added SQL. Soon it became the very best SQL standards implementation. That's when I picked it up again, and stayed with it.
I suspect it's actually Microsoft Office with Access and a mature Excel that did it. Everyone was using office at the time (for Word, Excel, Powerpoint and the emerging Outlook), so access was "free"; and many things that dBase&co were used for, were actually becoming easier to do in Excel.
I was surprised that databases are contained in *.dbf files, and no real database-service is present to process database-calls.
A shapefile is actually a collection of files, typically using the same filename with different extensions.
There is a .shp file containing the actual geographical "shapes": points, lines, polygons, etc.
Alongside it is a .dbf (dBASE) file with information about each shape. For example, if you have a list of cities, each city may have a boundary polygon (or multipolygon) in the .shp file, and the city name, population, and other demographics in the .dbf file.
The other required file is a .shx containing indices into the .shp file for faster navigation through it.
There may be other files as well; a common one is the .prj file that describes the projection used in this shapefile.
This collection of files, foo.shp, foo.dbf, foo.shx, etc. is referred to together as a "shapefile".
https://en.wikipedia.org/wiki/Shapefile
Bringing us full circle, GIS programmers and users often import a shapefile into PostGIS for processing. And PostGIS is, of course, a GIS extension to PostgreSQL!
Funny enough, two popular contenders are built on SQLite.
Oh man, you just wet my eyes! What a trip into the past!
I wrote my first paid project in dBase in 1982 when noone knew what "computer mouse" was. I answered the ad in newspaper and initially they wanted me to help them organize text files. I told them about dBase and that was my first project - a simple database to manage their parking lots. It was such early days of programming, we had basically zero agreement! Just half piece of paper that they pay me 10% in advance and short list of features to get remaining 90%. And funny thing is - no copyright agreement! So while I made few months of rent for their project, I ended up paying off my first mortgage by just offering the same piece of software to many other parking corps! It was incredible feeling! Like creating money out of thin air - they sponsor my plane/hotel and I just told them what PC (with DOS) I need to be ready. Then in just few hours of installation and training, I was on my way home with nice check. I frankly don't even remember how much it was back then...
Good times!
I suppose the secret is quickly delivering software that does what someone wants.
I don't want to stir up a religious battle here, but the tooling around MSSQL is just that much better.
And the sql dialect is imo more to my taste.
We’re working on a lot of tooling at https://supabase.io. We’re open source too. If there is anything in particular you’re missing, let me know
How about a UML/ERD diagram tool which can spit out schemas and, given a schema to start with, migrations?
There are a couple dozen of these being offered as SAAS with varying levels of usability and complexity. If you built one of those you'd soon have the attention of everyone who is bitching about paying 20 dollars a month to see their own database in a way that management can visualize for their charts and graphs.
ERD diagrams have existed for decades. Even Oracle gives such a tool away. But for Postgres, the easiest way to create a schema with migrations without a credit card is still to write a Django model by hand, afaik.
Once we've stabalized this we will start on git integrations - link up your repo for to keep everything in code - your types, schema, tables, functions, etc. We are also brainstorming ways that we can offer branching - since we give a full database server, we can possibly "branch" the default "postgres" database into various schemas which you can test with.
> ERD diagrams have existed for decades
If you know of a good open source tool which we should be supporting for this, let me know. Otherwise we don't mind building from scratch!
https://github.com/projectstorm/react-diagrams (diagram library)
https://www.npmjs.com/package/react-database-diagram (an example using the above library)
There is a Node migration library that seems well received but as I hinted above I have always trusted Django's migration system and still use it to bang out schemas if I know the schema will change, so I can't vouch for this lib personally...
woah. This looks powerful.
Thanks for the libraries - i'll research more and loop back here if we end up using them.
The story in all of this is that the only thing keeping Joe Q. Public from using a database instead of a spreadsheet is the very first step: designing a schema. Database design isn't really an impossible task for a non-developer to learn, if it can be visualized. If you start out at a command line and try to explain many-to-many relationships to them in code, you just lost them. People are visual creatures.
Since we're talking about giving people an API that takes care of CRUD and basic view logic at the database itself, and returns the data and the success/fail/delay messaging in a way that is standardized on the UI via javascript objects and async browser functions, the only hurdle that's left is that first one...
...designing a schema.
edit: on your topic of branching into test environments, I am in the process of evaluating low-code platforms for non-profits with limited development money but complex data needs at the moment, so I think we're thinking about the same things.
Everyone and their brother is working on a low-code platform at the moment. A big part of whether those platforms succeed or not is how they handle "last mile" complex logic. By last mile I mean, you can feasibly deliver basic CRUD and data analysis visualizations with drag and drop and query binding in generic javascript UI components. But there's still going to be the 10% complex business logic cases that don't fall within the 90% of an application that the generic components and boilerplate don't cover.
Off the top of my head, you could guide a user through some form logic that picks the first hundred or fifty or whatever rows from their existing tables that match a certain subset of shared keys, and tells them to flag columns which have sensitive data. Once the columns are flagged the tool could jumble the data in those columns and spit out a sandbox without any identifiable sensitive data that they could then safely hand to a contractor or freelancer to work on for custom development.
This is news to me. Any sources for this?
edit: https://medium.com/@calebmer, https://github.com/calebmer, https://calebmer.com/
I misspoke, not a founder of the company but an engineering hire there after he left Facebook, and the creator of PostGraphile. The gist of what I was getting at is there is a lot of activity swirling around putting business logic back on databases, it seems.
We might look back at these days as the high water mark of "peak javascript."
I do feel that Postgres administration mostly caters to command line fans (like myself); the people more used to visual tools like SQL Server Management Studio tend to find it difficult to transition because of this. Even if they might be absolute wizards at writing SQL otherwise.
Part of it is familiarity of course, but there's always something ever so slightly embarrassing about showing people psql when they're used to advanced IDEs.
If it's the latter, it'd be pretty trivial to generate a more mobile-friendly html format. I wonder if arxiv supports something like this.
There's even a (humorous) economics paper on why this is an advantageous method: https://onlinelibrary.wiley.com/doi/abs/10.1111/joes.12318