HNHacker News
TopNewBestAskShowJobs

hasyimibhar

81 karma · joined November 8, 2019

I'm building Polynya (https://polynya.dev/) and pg2iceberg (https://pg2iceberg.dev/). LinkedIn: https://www.linkedin.com/in/hasyimi-bahrudin/
submissionscomments
hasyimibhar··on Cloudflare K2: serverless event streams
Queues are for actions (do this), K2/event streams are for events (this happened).
hasyimibhar··on Cloudflare K2: serverless event streams
It’s built for Cloudflare ecosystem, IMO using K2 on its own doesn’t make sense. I’ve worked with a startup that uses Cloudflare for everything except for Kafka. If K2 were to exist back then, I’m pretty sure they would have used it.
hasyimibhar··on Show HN: EterDB, a Postgres fork that makes it easy to recover from incidents
The homepage shows an example of an agent accidentally dropping a table.

1. This is such an insane example, I don't get why would you give agents write access to your production db in the first place. Are people really doing this? I don't even have production connection urls on my laptop. Any manual statements executed against the db must be treated as a war-room situation with at least another engineer reviewing your SQL before you execute it.

2. Dropping the table could have easily caused writes to fail. Most likely there is no way to recover these writes (especially if it's from user requests), so it could have lead to loss of data. Reversing the table drop doesn't fix this issue.

hasyimibhar··on It's OK to hardcode feature flags (2025)
Having worked at a large company that misuses feature flag service for everything, here are some examples why you shouldn’t:

- one team uses feature flag for product gating. Feature flag service goes down. Users temporarily got locked out of the features they paid for.

- one team uses feature flag for dynamic pricing by leveraging targeting rules (how hard is it to write a bunch of if else in code?). It’s evaluated against all users, even if they are not active (for analysis reasons). Feature flag service charges by MAU. We have millions of users. Our feature flag service bill is now 6 digits per year.

- one team uses feature flag as literal json store instead of a proper db (god knows why). Someone updated the value but the “schema” is wrong. Shit breaks.

hasyimibhar··on Cloudflare's AI Psychosis
> Examples? Ok -> Let us look at data storage.

> They got D1 (SQLite serverless), Durable Objects with their own SQLite, KV, R2, Queues, and Hyperdrive to speed up external Postgres or MySQL.

Maybe it's just me, but I don't see this as confusing:

- D1 is SQLite on cloud, their go-to relational database

- Durable object is the cousin of durable execution (via actor model instead of saga)

- KV is for caching like Redis (sub-ms read latency)

- R2 is S3, R2 SQL is Iceberg + Athena (for OLAP queries)

- Queue is SQS

- Hyperdrive is bridge to allow CF worker to use native Postgres/MySQL driver (since v8 isolate cannot maintain persistent connection)

I've built some stuff on top of CF, and the data ecosystem is actually useful.

hasyimibhar··on How We Pushed CDC into Postgres
DMS is so unreliable though.
hasyimibhar··on How We Pushed CDC into Postgres
I’m waiting for Cloudflare R2 to eventually support mirroring Postgres into R2 catalog. It seems like a nice fit because they already have R2 SQL.
hasyimibhar··on How We Pushed CDC into Postgres
Yes I’m not suggesting to do this inside Postgres. I’m hoping that a Postgres provider can provide this mirroring capability out of the box, similar to how they provide a connection-pooled endpoint out of the box so I don’t have to self-host pgBouncer. I just want to be able to check a box somewhere and have a table in Postgres automatically mirrored to Iceberg, with guarantee that no data is lost. They can charge more for it, I will gladly pay.
hasyimibhar··on How We Pushed CDC into Postgres
It's interesting to watch how different companies that offer both Postgres and warehousing solution under 1 roof approach the same problem:

- ClickHouse focuses on traditional CDC (ClickPipes) and just make it blazingly fast

- Databricks leans on their unified storage architecture (LTAP) to avoid copying data (though you can argue there is still a copy in the cache)

- Snowflake uses a data mirroring CDC as extension so it runs directly on Postgres

I'm still waiting for a Postgres provider to just let me mirror data directly to Iceberg, so I can plug in my own stateless query engine.

hasyimibhar··on Making Postgres 300x faster for analytics: batching, operator fusion, and SIMD
> Almost always they end up setting up a separate system such as Clickhouse and replicating the data between the two systems. Now they can have one system that's Postgres-compatible, and it's faster than either of the original systems.

I can see the appeal for pgrust for smaller teams who need analytics, and don't want to deal with having to ETL to something like ClickHouse. But beyond a certain complexity, replicating your data to a warehouse or lakehouse _is_ the right approach and more scalable for several reasons:

- Analytics tend to be centralized, i.e. you want data from several Postgres databases spread across multiple teams to be replicated into 1 place, so people can start joining data across the entire business

- Analytics tend to fall under a different team ownership with their own set of non-technical requirements (e.g. data governance)

- Lakehouse architecture (Iceberg + [insert query engine]) is more scalable in terms of cost

- In some cases, you want to be able to swap different query engines depending on the use case, e.g. use PuppyGraph to query your data in Iceberg for fraud analysis

hasyimibhar··on Making Postgres 300x faster for analytics: batching, operator fusion, and SIMD
I'm guessing it's something like MongoDB TTL index[0]. It's useful for huge append-only tables where you want a row to be automatically removed after a period of time. The postgres approach is usually to partition your table by the event/insertion timestamp, and use something like pg_partman[1] to automatically drop entire partitions over time.

[0] https://www.mongodb.com/docs/manual/core/index-ttl/

[1] https://github.com/pgpartman/pg_partman

hasyimibhar··on Show HN: Isopolis – Isometric pixel map of SF
Sorry but that is not pixel art.
hasyimibhar··on The startup's Postgres survival guide
If your domain is analytics-heavy, don't try to optimize your Postgres for analytics. Follow the standard pattern of mirroring your data to a data warehouse and go to town there instead. There will be upfront cost of having to pay for a warehouse and the ETL, but it will be worth it.
hasyimibhar··on Xiaomi-Robotics-1
As a parent with 3 kids and both my wife and I working, folding clothes is the most tedious household chore. It's physically and mentally exhausting (unlike doing the dishes, which is only physically exhausting as I can wash piles of dishes while zoning out). And not to mention take so much time, a week's worth of unfolded pile of clothes for our entire household can take me 1-2 hours to fold uninterrupted (2-3 hours if I do it while watching Netflix or something).
hasyimibhar··on Making 768 servers look like 1
I'm surprised no one has mentioned Multigres yet, they are the competitor of Neki. I'm a big fan of both and has been following them since they were announced last year. I think this blog post is the first time they are talking about the internals of Neki. In contrast, Multigres is being built in public since day 1, you can see their high-level architecture here [1], though I'm still waiting for more details on their sharding model.

[1] https://multigres.com/docs/architecture

hasyimibhar··on Benchmarking coding agents on Databricks' multi-million line codebase
At huge companies, it's hard to prioritize this because it's hard to pin dollar values to removing legacy code, while it's easier to show how building feature X will earn the company $Y amount of revenue. And because of that, there is also no incentive to do it, you don't get promoted by deleting old code, you get promoted by showing how your effort helped contribute to company revenue. At my previous company (100+ engineers, hundreds of microservices), teams that regularly clean legacy codebase tend to be platform teams (cost centers), while teams that struggle to get these prioritized are product vertical teams (revenue centers).
hasyimibhar··on Postgres data stored in Parquet on S3: LTAP architecture explained
How does LTAP architecture deals with major Postgres upgrade? Is it truly zero-downtime for both upstream and downstream?
hasyimibhar··on Postgres data stored in Parquet on S3: LTAP architecture explained
> If you have something like dolt (not affiliated), a version controlled database, you wouldn’t have to slap change dates on anything OR create your historical table. The changes would be implicit in the version history.

Nope, even if I have the ability to see the exact changes of each row, I would still add timestamps everywhere, because timestamp of row change does not equal event timestamp. For example, if I have an order table with status column, and I see a CDC event where status changed from in_progress to completed, I cannot simply assume that the CDC timestamp is the timestamp when order was completed. It's possible that the source database received the event late a few minutes late due to delay upstream, or it's backfilling some missed orders a few days ago. Having a completed_at timestamp (and a bunch of other timestamps for each order lifecycle) would eliminate any ambiguities, and your data analyst will thank you for it.

It's the same thing with row history. You cannot simply assume that your row changes are aligned with the logical history of your entity.

hasyimibhar··on Postgres data stored in Parquet on S3: LTAP architecture explained
It wouldn't be possible to do this with LTAP architecture since (I'm assuming) the individual logical changes are not visible. But honestly I've always seen SCD type 2 table as a workaround due to lack of data modeling experience in the source database. If you design your tables correctly, you shouldn't need SCD type 2 downstream.

For example, if you know your user can change emails, and there might be events from another source that is keyed by user email (e.g. marketing-related events), then naturally you will need some sort of email_history table that has historical mapping of user id to email (you probably need it for audit purposes too). Then in this case there is no need to build SCD type 2 table of user from CDC, it's already there.

hasyimibhar··on Rive, Fast and reliable background jobs in Go
How does it compare to a full-fledged durable execution platform like DBOS[0], which follows the same philosophy? Looks like River does have workflows, but it's locked behind Pro [1].

[0] https://dbos.dev

[1] https://riverqueue.com/docs/pro/workflows

hasyimibhar··on PgDog is funded and coming to a database near you
What about Multigres[0]? It builds on top of Postgres and adds HA (based on Flexible Paxos[1]), sharding, etc. They're still not production-ready, but I'm highly optimistic they will solve a lot of the problems Postgres have.

For example, with Multigres, you should be able to achieve true zero downtime major version upgrade by simply resharding [2]. With vanilla Postgres + pgBouncer, you can only achieve near-zero downtime (few seconds at most), though it's probably good enough for most use cases.

[0] https://multigres.com/

[1] https://fpaxos.github.io/

[2] https://multigres.com/docs#migrate-across-postgres-versions

hasyimibhar··on Migrating from Go to Rust
It is also easier to make your code deterministic with Rust vs with Go, which is incredibly useful if you need to perform deterministic simulation testing + property-based testing. I recently wrote a Postgres-to-Iceberg data mirroring tool [1] in Go, but I ported it to Rust because I wanted the ability perform DST without fighting Go's runtime [2]. But if the domain is not critical that warrants DST, I would still pick Go over Rust any day.

[1] https://github.com/polynya-dev/pg2iceberg

[2] https://www.polarsignals.com/blog/posts/2024/05/28/mostly-ds...

hasyimibhar··on Incident Report: Railway Blocked by Google Cloud (Resolved)
The problem with the us-east-1 outage is that a lot of big companies are there, so even if you try your best not to depend on us-east-1, your third party providers are most likely there. In my previous company, we were completely down during us-east-1 outage because of other dependencies that are beyond our control.
hasyimibhar··on Lakebase architecture delivers faster Postgres writes
In some cases you have no choice but to retain the data, e.g. due to compliance. But the good thing is it doesn't have to be in Postgres. You can periodically offload data to a lakehouse, then delete it from Postgres. If the table is partitioned, delete should be cheap.

I'm guessing with Neon, since their storage is a lakehouse, you get this for free.

hasyimibhar··on Lakebase architecture delivers faster Postgres writes
How does Lakebase compare to OrioleDB[0]?

[0] https://www.orioledb.com/

hasyimibhar··on Show HN: Mljar Studio – local AI data analyst that saves analysis as notebooks
You should check them out, their interface pretty much looks like chat nowadays.
hasyimibhar··on Show HN: Mljar Studio – local AI data analyst that saves analysis as notebooks
How does this compare to open source Deepnote[0]? We use the cloud version (BYOC) at my previous company to replace self-hosted Jupyter notebooks, and it's pretty great.

[0] https://github.com/deepnote/deepnote

hasyimibhar··on Show HN: DAC – open-source dashboard as code tool for agents and humans
I mean that's what the Vega team is doing no? They are building the standard grammar (Vega-Lite), along with an implementation (Vega). And they are already quite established with rich ecosystem, and supports a ton of components[0]. The only thing missing is that it expects a CSV or inline data source. But it's probably not too hard to build an extension that connects to a data warehouse with an SQL query.

[0] https://vega.github.io/vega-lite/examples

hasyimibhar··on Show HN: DAC – open-source dashboard as code tool for agents and humans
Why not use Vega-Lite[0]? It’s my go-to data viz DSL with Claude.

[0] https://vega.github.io/vega-lite/

hasyimibhar··on If I could make my own GitHub
You can skip by running git commit --no-verify. I know this because I also hate pre-commit checks, and I will automatically use it when working with any codebase that has one.
Page 1 of 2Next →