HNHacker News
TopNewBestAskShowJobs

craigkerstiens

32,067 karma · joined September 20, 2008

I work at Snowflake building best in class Postgres via Crunchy Data, building Snowflake Postgres - a multi-cloud Postgres managed service. Previously I ran product at Azure Postgres, product at Citus data, and a number of product areas at Heroku.

You can find more thoughts from me: https://www.craigkerstiens.com

submissionscomments
craigkerstiens··on An Update on Heroku
It sounds like there were pretty broad layoffs which impacted a lot more than just a focus on enterprise contracts. It wasn't "just" a few enterprise sales people. Engineering may have indeed been the least impacted, but this sounds like biggest round of layoffs to hit Heroku since its inception, not just some right sizing from over hiring.
craigkerstiens··on An Update on Heroku
It sounds like there were pretty broad layoffs which impacted a lot more than just a focus on enterprise contracts. It wasn't "just" a few enterprise sales people. Engineering may have indeed been the least impacted, but this sounds like biggest round of layoffs to hit Heroku since its inception, not just some right sizing from overhiring.
craigkerstiens··on Hacking the Postgres wire protocol
Pgquery was created by the pganalyze team for their own purposes I believe initially for features like index recommendation tooling, but immediately planned as open source. It is indeed a very high quality project with the underlying C implementation having several wrappers that exist for a number of languages[1].

[1] https://github.com/pganalyze/libpg_query/blob/15-latest/READ...

craigkerstiens··on Making Postgres scale
Citus works really well *if* you have your schema well defined and slightly denormalized (meaning you have the shard key materialized on every table), and you ensure you're always joining on that as part of querying. For a lot of existing applications that were not designed with this in mind if can be several months of database and application code changes to get things into shape to work with Citus.

If you're designing from scratch and make it worth with Citus then (specifically for a multi-tenant/SaaS sharded app) it can make scaling seem a bit magical.

craigkerstiens··on Making Postgres scale
Probably as useful is the overview of what pgdog is and the docs. From their docs[1]: "PgDog is a sharder, connection pooler and load balancer for PostgreSQL. Written in Rust, PgDog is fast, reliable and scales databases horizontally without requiring changes to application code."

[1] https://docs.pgdog.dev/

craigkerstiens··on Postgres Is. Introducing PostgreSQL compatibility index
I appreciate the idea of this a lot but am a bit skeptical.

The good thing is this would at least shine a better light on all those that "claim" to be Postgres but really have little to nothing to do with it. Overwhelmingly people are supporting the wire protocol despite being a completely separate database because Postgres is already so universal and it'd be a huge investment to recreate that ecosystem of language drivers and everything else around it.

The reality is even "wire protocol" there are varying levels of support depending on what you're trying to do.

Then when you get down to functionality, it could be "we support the Postgres data types"... well except this one or that one. That's fine and good until a user is surprised 2 years into building an application.

Even the notion of we support all Postgres extensions, well all Postgres extensions don't work together some take hooks and change queries that other extensions want to modify for themselves.

Having worked with Postgres and managed Postgres for a very long time. Postgres is Postgres, there are extensions that modify Postgres, there are forked versions of Postgres, and things that are "Postgres" compatible simply aren't Postgres.

craigkerstiens··on Pg_incremental: Incremental Data Processing in Postgres
Indeed, thought this may have a lot more context with the examples and possibly be relevant.
craigkerstiens··on Amazon Aurora DSQL
It feels very disingenuous to say "Postgres compatible" and have this as a missing feature set. I'm sure they'd quickly argue it's wire compatibility, but even then it's a slippery slope and wire compatible is left open to however the person wants to interpret it.

There is no 'standard' or 'spec' for what makes something Postgres wire compatible.

This feels like a strong overreach on the marketing front to leverage the love people have for Postgres to help boost what they've built. That is not to say there isn't hard and quality engineering in here, but slapping Postgres compatible on it feels lazy at best.

craigkerstiens··on Ask HN: Bluesky Accounts Worth Following for HN Enthusiasts
I've been maintaining a starter pack for Postgres people - https://go.bsky.app/Acp7hmk
craigkerstiens··on Pg_parquet: An extension to connect Postgres and parquet
This alone wouldn't be a full replacement. We do have a full product that does that with customers seeing great performance in production. Crunchy Bridge for Analytics does similar by embedding DuckDB inside Postgres, though for users is largely an implementation detail. We support iceberg as well and have a lot more coming basically to allow for seamless analytics on Postgres building on what Postgres is good at, iceberg for storage, and duckdb for vectorized execution.

That isn't fully open source at this time but has been production grade for some time. This was one piece that makes getting to that easier for folks and felt a good standalone bit to open source and share with the broader community. We can also see where this by itself for certain use cases makes sense, as you sort of point out if you had time series partitioned data, leveraged partman for new partitions and pg_cron which this same set of people authored you could automatically archive old partitions to parquet but still have thing for analysis if needed.

craigkerstiens··on pg_duckdb: Splicing Duck and Elephant DNA
Very much agreed with this general idea, and believe a lot of this was inspired by the team we hired at Crunchy Data to build it as they were socializing it for a while. Looking forward to pg_duckdb advancing in time for now it still seems pretty early and has some maturing to do. As others have said, it needs to be a bit more stable and production grade. But the opportunity is very much there.

We recently submitted our (Crunchy Bridge for Analytics-at most broad level based on same idea) benchmark for clickbench by clickhouse (https://benchmark.clickhouse.com/) which puts us at #6 overall amongst managed service providers and gives a real viable option for Postgres as an analytics database (at least per clickbench). Also of note there are a number of other Postgres variations such as ParadeDB that are definitely not 1000x slower than Clickhouse or DuckDB.

craigkerstiens··on Re: Do people IRL know you have a blog?
I love when friends do this. It's hard to keep up with people and what they're up to. Publishing and letting people subscribe to me is a great way to share things. A few examples of some friends who are doing this:

Justin Searls (fairly known in Ruby and Rails community) mostly quit a lot of various social channels though publishes on some of them one direction. He started a podcast that wasn't meant to be guests of some specific topic, it's just him updating you on things. What he's working on, what he's learning, random stories, etc. - https://justin.searls.co/casts/

Brandur who I've worked with at a couple of places (Heroku previously, and now Crunchy Data) who writes great technical pieces that often end up here also has more of a personal newsletter. While there are technical pieces in there at times he'll also talk about personal experiences my favorite one is some of the unique experiences hiking the Pacific Trail (https://brandur.org/nanoglyphs/039-trails).

craigkerstiens··on Half a century of SQL
This is very much why we built the Postgres playground, which has Postgres embedded in your browser with guided tutorials - https://www.crunchydata.com/developers/tutorials
craigkerstiens··on Show HN: Serverless Postgres
At very first glance this would be much closer to neon, with the separated storage and compute.

Crunchy Postgres for Kubernetes is great if you're running Postgres inside Kubernetes, but is more of standard Postgres than something serverless. Citus also not really serverless at all, Citus is more focused on performance scaling where things are very co-located and you're outgrowing the bounds of a single node.

craigkerstiens··on Crunchy Bridge for Analytics: Your Data Lake in PostgreSQL
Hmm, let me see if we have a link to slides. There was no video recording unfortunately but we can definitely get slides posted.
craigkerstiens··on Crunchy Bridge for Analytics: Your Data Lake in PostgreSQL
It's not datafusion, much more custom with a number of extensions that underly pieces. And we've got a number of other extensions that will be in the works, to the user it's still a seamless experience but we've seen that smaller extensions that know how to work together are easier to maintain. For example we're working on a map type one that knows how to understand the map types within Parquet files within Postgres. In time we may open source some of these pieces, but we don't have a time frame for that and it's a case by case on each of the extensions.
craigkerstiens··on Crunchy Bridge for Analytics: Your Data Lake in PostgreSQL
There is coordination from the Crunchy Bridge control plane to the data plane, that the extension is then aware of.

At this time it's not FOSS, we are going to consider opening some of the building blocks in time, but at the moment they have a pretty tight coupling on both the other extensions and on how Crunchy Bridge operates.

craigkerstiens··on Crunchy Bridge for Analytics: Your Data Lake in PostgreSQL
It's a custom extension and actually a number of custom extensions, with quite a few more planned to further enhance the product experience. All extensions work together as a single unit to compose Crunchy Bridge for Analytics, but under the covers lots of building blocks that work together.

Marco and team worked were the architects behind the Citus extension for Postgres and have quite a bit of experience building advanced Postgres extensions. Marco gave a talk at PGConf EU on all the mistakes you can make when building extensions and best practices to follow–so in short quite a bit gone into the quality of this vs. a quick one off. Even in the standup with the team today it was remarked "we haven't even been able to make it segfault yet, which we could pull off quite quickly and commonly with Citus".

craigkerstiens··on ElephantSQL Is Shutting Down
Fully agree with this sentiment it's very much our focus and goal at Crunchy Data with one big thing I'd add-great support.

I recall seeing them crop up in the early days of building and running Heroku Postgres, they were a very very early managed service provider. To my knowledge they never seemed to grow to massive scale but were a steady business (though I don't know any of the details for sure). That they were still around for over a decade is a testament from a lot of others.

craigkerstiens··on Blazer: Business intelligence made simple
Love the callout to Dataclips. It was easily my favorite least used feature by Heroku customers. Blazer and PgHero both got a bunch of inspiration from some of the early things we built at Heroku and its amazing having Andrew crank out so many high quality projects to make some of the tooling more broadly available.
craigkerstiens··on Hosted Postgres provider ElephantSQL is shutting down
Shameless plug, but we aim to get pretty close to this on Crunchy Bridge (our hobby-0 with 2 vcores starts at $10 a month) - https://www.crunchydata.com/pricing/calculator.
craigkerstiens··on An overview of distributed Postgres architectures
Marco (author) is probably asleep at this point and could give a deeper perspective. He sort of hits on this when talking about disk latency... Depending on your setup and well just from some personal experience I know it's not crazy for Postgres queries to go at 1ms per query. From there you can start to do some math on how many cores, how many queries per second, etc.

Single node Postgres (with a beefy machine) can definitely manage in the 100k transactions per second. When you're pushing the high 100k into millions read replicas is a common approach.

When we're talking transactions, question of is it simply basic queries, bigger aggregations, and is it writes or reads. Writes if you can manage to do any form of multi-line insert or batching with copy you can push basic Postgres really far... From some benchmarks Citus as mentioned can hit millions of records per second safely with those approaches, and even without Citus can get pretty high write throughput.

craigkerstiens··on Fly Postgres, Managed by Supabase
I think a bit of confusion on Fly Postgres vs. the Supabase offering. The earlier was unmanaged on Fly infra.

I'm not sure the full details on supabase as it's more recent.

This is a pretty good breakdown of various database providers and in particular a lot paired with Fly - https://dancroak.com/webstack/

craigkerstiens··on Fly Postgres, Managed by Supabase
I don't think that's correct at all. Heroku Postgres has a central control plane that is monitoring availability and orchestrating things. There are continual health checks that go back to Heroku. In the event of unavailability it sets off a page to the on-call engineer to investigate if systems haven't restored availability.

My understanding of Fly Postgres is they put a lot into the tools to orchestrate, but there is not centralized monitoring and in the event of a failure it is up to you to realize and remediate.

Disclaimer: Was part of the team that built Heroku Postgres, and know the Fly team pretty well but don't personally use Fly Postgres so it's my understanding from the team. We've had a number of customers leverage Crunchy Bridge (build by a lot of the original Heroku Postgres team) use us for the managed Postgres connected to fly.io via Tailscale.

craigkerstiens··on My favorite database shirts
I love the "this guy" piece... Meanwhile he's been the most public person in academia talking about databases over at least the last 5 years, maybe the last 10. He's not only done an awesome job of talking about foundational pieces of databases, but also examining new databases that have come up over the last 10 years or so. He's course is quite open as well, so it's not just the student base that gets to take advantage - https://15445.courses.cs.cmu.edu/fall2023/

The shirts he's generally helped promote and publicize those companies so he has things to hand out to his students, TAs, graduate students.

craigkerstiens··on Postgres: The next generation
This is a common source of confusion for a ton of folks. Anyone can submit a patch, but commit bits are reserved for a much smaller list. The attitude is something like you commit it, you maintain it–so if bugs come in you'll spend your time fixing those for whatever time it takes vs. working on the next shiny feature that you're excited about for the next release.

There was sort of a fuzzy "major" contributors (https://www.postgresql.org/community/contributors/) which were people that contributed major features and then a list of other contributors. Depending on who you talk to this is either dated or a pretty close attempt at reflection of reality but not perfect. In recent years they expanded the contributors to include others that were contributing in non-code ways though it's still a decent place to find people contributing to major feature sets.

Of course this is not to be confused with the core team–which is more like a steering committee. But not so much steering committee of code and feature sets.

craigkerstiens··on Postgres: The next generation
I disagree on this. Yes it's C. But I've heard people comment "I don't like writing C, but I don't mind Postgres C".

The bigger hurdle which Peter mentioned in another thread is simply building up enough expertise with the system and having the right level of domain expertise.

craigkerstiens··on Postgres: The next generation
So many thoughts on this. The community has definitely ebbed and flowed, on this for a while. A few varying pieces of insight with no intention other than to share a bit more on the PG community. And I'm sure some current and former colleagues already in comment threads are going to correct me on nuance of a lot of this.

For several years there were no new committers at all. In recent years the team has tried to be a little more intentional about adding new ones and culling those no longer involved.

About 15 years there was a phase of letting a lot of younger people earn their commit bit. I can recall 3 people by name that all got a commit bit before the age of 25, and they may have actually all been under 22. One of those three shortly after moved on to work outside of the Postgres community, another quietly was busy on other things for over 10 years before coming back, and the third was actively involved going forward. I suspect there was some unease of folks getting a commit bit and then sort of falling off a cliff so it slowed for a few years on adding new folks. Edit - sounds like it was less age driven but maybe still slightly related to some folks falling off that there was a slow down in new committers – tldr - you're not getting a commit bit right out of college for Postgres.

What to me would be interesting but likely hard to gather is what age to people become a committer to Postgres. It wouldn't surprise me if the average age of getting a commit bit is closer to 45 than not. Many folks contributing come to Postgres after other systems work or just don't consider contributing to they're a bit more seasoned because it feels intimidating–I mean patches sent on a mailing list who does that any more? Postgres thats who.

craigkerstiens··on Postgres: The next generation
Just coming to say hi Paul! I recall you being very excited and getting the email after you submitted your first patch and loved it.

I think a ton of the current set of "next generation" came to Postgres a good bit later. Even Tom himself will tell you he did some stuff with images for a few years which undersells himself (tiff, jpg, png - in some form involved in creation of each of those), then found this Postgres thing and started working on it.

craigkerstiens··on When did Postgres become cool?
I definitely recall being on the talk committee one year for DjangoCon, and there was some rough discussion. (Context: The conference was generally a two track conference but keynotes and a few other sessions were single track). One of the single track talks was about Postgres. The discussion was roughly "If we have a Postgres talk we should have another talk like Mongo or MySQL" and the response was roughly "Everyone in Django is using Postgres and if you're not you should be at the talk to learn why you should".

Way more Rails apps used MySQL or other databases, it was largely Heroku winning Rails that led to the strong adoption amongst that community.

Page 1 of 17Next →