973 karma · joined July 16, 2017
admin@viggy28.dev
Short answer: yes, column-level schema changes sync to Iceberg automatically[0].
Logical replication (pgoutput in v1) doesn't actually stream DDL statements. Instead, Postgres emits a fresh Relation message describing the table's current column layout right before the next change to that table. So we diff that against the last layout we knew and infer what changed.
From there we evolve the Iceberg schema in place: flush any buffered rows under the old schema first, then write a new metadata version with the change. What's handled today:
- ADD COLUMN — new field ID allocated; the column's Postgres DEFAULT is carried into Iceberg's initial-default/write-default, so existing rows read back correctly
- DROP COLUMN — removed from the current schema, existing data files untouched
- Type widening — int4→int8, float4→float8 (the changes Iceberg considers compatible)
- REPLICA IDENTITY changes
[0] https://github.com/viggy28/streambed/pull/21Agree, CDC is like Death by a thousand cuts. I believe Debezium has a Java library.
My initial need was Postgres compatibilty. Wanted to give an endpoint that BI and dashboard teams can use to query as if they are querying a Postgres replica. Added more context here https://news.ycombinator.com/item?id=48350820
So the question I started with was: what's the fewest components I could get away with? That led to the architecture here — Streambed connects to Postgres as a logical replication subscriber (same mechanism as a read replica) and streams WAL changes straight into Apache Iceberg on S3, queryable from psql via an embedded DuckDB. There are a lot of edge cases to handle, and it's very much early days.
Welcome any feedback.
How’s it different from Parallel Web Systems?
Hi, we have been working on a personal assistant for knowledge professionals. A lot of challenges around building an advanced RAG, improving reasoning, and custom models
Learn more about the role: https://wellfound.com/jobs/3382533-senior-ai-ml-engineer or email vignesh@buildrappo.com
PS: I am a solo founder, and we are generous with equity
Everyone says early PMF is crucial, but there is no streamlined support there.
Learn more about the role: https://wellfound.com/jobs/3133503-lead-founding-full-stack-...
PS: I am a solo founder, and we allocated generous equity for advisors and early hires.
I don’t remember they have a similar doc for setting up HA.
Hi, I am the founder of Rappo (a network for GTM). We observed that technical startups spend way too much time in the early stages figuring out GTM.
Everyone is like PMF is crucial, but there is little support/help around in getting there.
Learn more about the role: https://wellfound.com/jobs/3133503-lead-founding-full-stack-...
PS: I am a solo founder, and we allocated generous equity for advisors and early hires.
As a database practitioner and a founder of a database startup, I would be curious to see how you approach these challenges and address them. Also, how you make it economically sustainable too.
0: https://github.com/omnigres/omnigres/tree/master/extensions/...
At Omnigres, our north star is to enable developers to laser-focus on business needs instead of fighting technological challenges. We're fighting the complexity and inefficiencies of contemporary stacks by removing them instead of hiding them.
At the core, we are turning Postgres into an Application Runtime. Why? Because we believe that code and data are inseparable in pretty much all of the line-of-business application systems. Turns out, when done this way, applications work a lot faster, require a lot less maintenance, and are simply easier to write.
Our foundation is open source and is available at https://github.com/omnigres/omnigres
We're backed by some great early-stage VCs and looking to onboard people who can move quickly, learn on the go, and maintain the focus on the goals. Another way to look at it is that we want to meet other pragmatic idealists.
Email: founders@omnigres.com
We have been seeing this trend (or pendulum swing) of pushing SQL and simplicity. I am not saying this solution is simple (will leave that for discussion).
Tangential, if this kinda stuff is interesting, folks might be interested in what Omnigres (disclaimer: my employer) is doing with the python integration in Postgres
We're fighting the complexity and inefficiencies of contemporary stacks by removing them instead of hiding them.
At the core, we are turning Postgres into an Application Runtime. Why? Because we believe that code and data are inseparable in pretty much all of the line-of-business application systems. Turns out, when done this way, applications work a lot faster, require a lot less maintenance and are simply easier to write.
Our foundation is open source and is available at https://github.com/omnigres/omnigres
We're backed by some great early-stage VCs and looking to onboard people who can move quickly, learn on the go and maintain the focus on the goals. Another way to look at it: we want to meet other pragmatic idealists.
Email: founders@omnigres.com