HNHacker News
TopNewBestAskShowJobs

malisper

3,310 karma · joined December 31, 2013

Hi! I'm Michael Malis. I'm the co-creator of pgrust (https://github.com/malisper/pgrust/)

I previously ran Freshpaint (YC S19) for 7 years and before that I led the database team at Heap.

GitHub: https://github.com/malisper/

Blog: http://malisper.me

Email: michaelmalis2@gmail.com

submissionscomments
malisper··on Postgres rewritten in Rust, now passing 100% of the Postgres regression tests
> Threads does not offer any major performance advantage

This is very not true. When it comes to parallel queries, a process model adds a ton of overhead. You can't pass pointers between processes because the address space is different. This adds a ton of overhead in a bunch of different places. For example when doing a parallel hash join, Postgres will have each worker build a local hash table. Then it will take all the tuples out of the local hash table and copy them through shared memory to the leader who will then construct a new hash table. This duplicates a lot of work as you have to hash the tuples multiple times.

A lot of getting to Clickhouse level performance was making better use of parallelism.

malisper··on Postgres rewritten in Rust, now passing 100% of the Postgres regression tests
I spent a couple years managing a Postgres cluster with a petabyte of data. I wrote a couple blog posts from my work then[0][1]. I also wrote dozens of posts on the Postgres internals[2]. I've also given talks on how to generate fractals with SQL[3] and how to write a lisp interpreter in SQL[4].

[0] https://www.heap.io/blog/testing-database-changes-right-way

[1] https://www.heap.io/blog/analyzing-performance-millions-sql-...

[2] https://malisper.me/table-of-contents/

[3] https://www.youtube.com/watch?v=xKoYIvMFnoQ

[4] https://www.youtube.com/watch?v=MPSMH8w7nfw

malisper··on Postgres rewritten in Rust, now passing 100% of the Postgres regression tests
That's something I eventually want to fix. The challenge is the storage format is so integral to Postgres that it's going to be a huge PITA to come up with a novel design.

Right now OrioleDB is in beta. Once that becomes production ready, I'll evaluate incorporating it into pgrust.

For Ben Dicken, he has seen the project: https://x.com/BenjDicken/status/2074512043462603236. We're still working on all the novel features so I don't think it meets his bar quite yet.

malisper··on Postgres rewritten in Rust, now passing 100% of the Postgres regression tests
My approach has changed throughout the course of this project. Throughout most of the project, we were working off of a c2rust translation of Postgres to Rust. That gave us a bunch of Rust code that was unsafe but did pass the Postgres test suite and was fast. c2rust had split Postgres into 1000 different crates. We then went through 1 by 1 and rewrote each crate into idiomatic rust.

This naturally lended itself to a suite of skills to describe how to rewrite a crate from unsafe rust to idiomatic rust. The main three skills I had were 1) a skill for identifying the next crates to port 2) a skill for rewriting a crate and 3) a skill for auditing a crate and making sure there weren't any outstanding issues.

My exact approach for managing subagents changed throughout the project. Initially I was doing parallel coding sessions with Conductor. After dynamic workflows came out, I used that as it was really easy to spin up dozens of parallel subagents and manage it from a single orchestrator. Over time I switched from using dynamic workflows to manually spinning up subagents from a central agent. The issue with dynamic workflows is they waterfall. Each step needs to finish before the next one starts. By manually spinning up subagents, I could have claude start porting a new crate as soon as a prior subagent finished.

malisper··on Postgres rewritten in Rust, now passing 100% of the Postgres regression tests
Can you elaborate on the use case for query multiplexing? Is it so your client would only need to establish one connection with Postgres and then could run as many queries as it wanted?
malisper··on Postgres rewritten in Rust, now passing 100% of the Postgres regression tests
> The rewrite also could have simpler code in some cases

The Rust code is a literal translation of the Postgres code which returns the value at the end instead of an early return.

> I see a lot of MemoryContext

MemoryContext in C is used for multiple reasons: 1) performance 2) keeping track of how much memory has been allocated and where and 3) preventing memory leaks.

Reasons 1 and 2 are still relevant for Rust. The challenge is in C memory contexts are stored in a global variable. Global variables don't work well with the rust borrow checker so I opted for passing memory contexts as function arguments instead.

malisper··on Postgres rewritten in Rust, now passing 100% of the Postgres regression tests
Yep! The new version of pgrust supports batch based execution and a columnar format. I'm curious how you got δx to perform that well? From what I've seen a columnar layout only gets you part of the way and really good parallelism and really fast hash tables seem to make up a significant portion of why Clickhouse is faster.
malisper··on Postgres rewritten in Rust, now passing 100% of the Postgres regression tests
Rust actually made the change pretty simple. The main changes are:

  - Use thread local variables
  - Move everything from shared memory to process memory
  - Use threads instead of processes
I've started to see meaningful benefits by changing the parallel algorithms to use a shared memory space. For example parallel hash joins have to copy tuples through shared memory to pass them between workers. That's just not something I have to do.
malisper··on Postgres rewritten in Rust, now passing 100% of the Postgres regression tests
> why is the parser unsafe at all?!?

The parser was generated by c2rust. The Postgres parser is generated from yacc/bison itself so I didn't bother making it idiomatic.

malisper··on Postgres rewritten in Rust, now passing 100% of the Postgres regression tests
Note that most of the unsafes are confined to the parser which was generated by running c2rust over the Postgres parser. The Postgres parser is itself generated from yacc/bison, so I decided to port it over mechanically rather than idiomatically.

If there's particular unsafes that you think are egregious, let me know.

malisper··on Postgres rewritten in Rust, now passing 100% of the Postgres regression tests
Note that the code I believe you are referring to is from the parser which was generated with c2rust. The Postgres parser is generated from yacc/bison so rather than try to rewrite it idiomatically, I did so mechanically.
malisper··on Postgres rewritten in Rust, now passing 100% of the Postgres regression tests
The 50% is specifically on percona-tpcc[0]. I got there through a mix of batching (postgres processes a row at a time), prefetching, and several handful of other optimizations.

  [0] https://github.com/Percona-Lab/sysbench-tpcc
malisper··on Postgres rewritten in Rust, now passing 100% of the Postgres regression tests
It is theoretically possible to have a Rust port of Postgres support extensions. If you make all the relevant functions and structures ABI compatible with Postgres, extensions should work. The issue is the moment you're dealing with C pointers and C strings, pretty much all the code you have to write is unsafe.
malisper··on Postgres rewritten in Rust, now passing 100% of the Postgres regression tests
The version in the GitHub repo is ~8x slower than Postgres. I have a new unpublished version that is 50% faster than Postgres on transactional workloads and ~300x faster on analytical workloads.
malisper··on Postgres rewritten in Rust, now passing 100% of the Postgres regression tests
It's not used in production. I've been using different benchmarks to compare the performance vs other systems. Namely sysbench-tpcc[0] and clickbench[1]

[0] https://github.com/Percona-Lab/sysbench-tpcc

[1] https://github.com/ClickHouse/ClickBench

malisper··on Postgres rewritten in Rust, now passing 100% of the Postgres regression tests
I'm not sure what you mean? The rust code you're showing mimics the Postgres code: https://github.com/postgres/postgres/blob/2e6578292a9184dcaa...

The boolean being returned is the return value of the function. It's not used to return an error.

malisper··on Postgres rewritten in Rust, now passing 100% of the Postgres regression tests
Hey author here. Wasn't expecting to see this up.

To concisely give an overview of the project, I've been experimenting with using LLMs to build a better version of Postgres. Postgres is 30 years old and we've learned a lot about databases since hten. A lot of the techniques that work for doing a rewrite are also useful for doing a rearchitecture.

I'm now working on a new, not yet published version of pgrust that incorporates a lot of techniques. Currently the new version:

  - Passes 100% of Postgres regression suite
  - Implements a thread per connection model instead of the process per connection model Postgres does
  - Is 50% faster than Postgres on transaction workloads
  - Is ~300x faster than Postgres on analytical workloads. Right now it's 2x slower than Clickhouse on clickbench and I think it's possible to get faster than Clickhouse
If you have any questions, I'm happy to answer them.
malisper··on Rewriting Bun in Rust
If you want to follow along, I've been writing about it on my blog (https://malisper.me/) an you can follow the github repo here: https://github.com/malisper/pgrust
malisper··on Rewriting Bun in Rust
I can confirm a naive rewrite won't make things faster. I've been working on rewriting Postgres in Rust. I rewrote things function by function similar to how Jarred did. Even though the new Rust code mapped closely with the previous C code, it was 8x slower. This was due to myriad of reasons. For example naively converting a C union into a Rust enum can be slower because Rust stores a tag with the enum, while C unions do not.

I've been working on a new rewrite that's focused on beating Postgres on performance. As of this morning I got to 100% of the tests passing and have meaningful performance gains over Postgres.

malisper··on My AI-built PHP engine in Rust passes 17% of PHP-src tests, renders WordPress
> Maybe the takeaway is that 20% is about all the LLM can muster

At this point there's a long list of projects that have used LLMs to rewrite a system in Rust including:

  - Bun (https://github.com/oven-sh/bun/pull/30412)
  - Valkey (https://github.com/ianm199/valdr)
  - Git (https://github.com/gitbutlerapp/grit)
  - Postgres (https://github.com/malisper/pgrust)
With the exception of Bun, these projects were done pre-fable too, so I bet Fable will make these types of rewrites even easier.
malisper··on The C to Rust migration book
> But LLMs are great at this

Not really. The latency for an llm to make a single change is on the order of seconds. If in order to change a single function from unsafe Rust to safe rust requires thousands of changes, it will take hours to refactor a single function.

malisper··on The C to Rust migration book
I made a comment about this yesterday[0], but there's been a massive increase in people migrating from C to Rust due to LLMs.

In contrast to what the C to Rust migration book is recommending (using FFI to integrate Rust with C), I've found it much easier to start from scratch. I recently finished a project where I rewrote Postgres in Rust[1]. For context, Postgres is about one million lines of C code.

On one attempt, I tried using c2rust to convert Postgres into unsafe Rust code. That attempt succeeded in terms of getting working "Rust" code, but any attempt to change any piece to safe rust, would require thousands of changes across codebase. Even though I had working Rust code, I found it infeasible to get to working idiomatic Rust code.

Instead what I found to be more effective was starting a new codebase and rewrite each file from the Postgres codebase into Rust one at a time. This allowed me to guarantee that at all times the new codebase was idiomatic and simultaneously I could make one pass over the Postgres codebase to get working idiomatic Rust.

YMMV, but I found it way easier to generate a whole new codebase from scratch rather than incrementally rewrite an existing codebase.

[0] https://news.ycombinator.com/item?id=48738985#48739882

[1] https://github.com/malisper/pgrust

malisper··on I ported Kubernetes to the browser
A meta-trend I find interesting is there's a lot of projects using AI to rewrite existing systems in new programming languages. Most often in Rust.

  1. Bun rewritten in Rust
  2. Flow rewritten in Rust
  3. The react compiler was rewritten in Rust
  4. Grit is a new implementation of Git in Rust
  5. I've made my own rust rewrite of postgres that passes 100% of the regression and isolation tests[0][1]
I think AI changed the economics of these projects even more than it has the economics for software engineering work in general. Though direct AI code translation is usually slop for me.

One of the many things I did to deal with this was an audit skill that would:

  1. Find a small chunk of code to rewrite
  2. Have a list of things that it was looking for in each piece of code that's being rewritten
  3. Place that next to the code being translated
  4. If that document didn't exist and/or didn't say the code was passing the audit, code wouldn't be merged
  5. As I found problems and anti-patterns I would add those to the skill over time
This by itself still let a lot of slop slip through, but also preemptively caught a ton of issues as part of my overall process.

Complicated old "boring" infra software might actually be the most AI-rewriteable code right now

[0] https://pgrust.com

[1] https://github.com/malisper/pgrust

malisper··on SQLite is all you need for durable workflows
> Isn't concurrency also limited by your machines disk speed for writes, what difference does it make if you write sequentially vs concurrently? Why does concurrency even matter for databases?

For a simplified example, having three processes reading blocks X, Y, Z in parallel is much faster than having a single process read block X, wait for the read to finish, read block Y, wait for the read to finish, read block Z and wait for the read to finish.

malisper··on Bun's experimental Rust rewrite hits 99.8% test compatibility on Linux x64 glibc
There's a few big differences between the Anthropic C compiler and pgrust. The C compiler was built mostly autonomously and as a clean room implementation. OTOH I'm steering codex and using the Postgres source code as a reference. That's leading to the implementation being based more on how pg does things than anything else. If you want to try it out, I compiled it to wasm so you can try it out here[0]. You'll see it's much more faithful to Postgres than a C compiler that doesn't handle type checks.

[0] https://pgrust.com/

malisper··on Bun's experimental Rust rewrite hits 99.8% test compatibility on Linux x64 glibc
I wrote up a bit about my workflow here[0][1]. I'm using conductor.build to manage multiple codex sessions at once. When I hit the rate limit, I'm using codex-auth[2] to switch codex accounts.

[0] https://malisper.me/pgrust-rebuilding-postgres-in-rust-with-... [1] https://malisper.me/pgrust-update-at-67-postgres-compatibili... [2] https://github.com/loongphy/codex-auth

malisper··on Bun's experimental Rust rewrite hits 99.8% test compatibility on Linux x64 glibc
Same but for multi-threaded Postgres[0]. 96% pg regression tests pass after 1 month and 823K LOC. 8 Codex accounts at $200/mo is what i could use up with no Mythos

I've also seen the benefits of Rust for this too. And making the bet that my pg experience will help me make good design choices around many of the things people have been having trouble with in pg for a long time[1]. Excited to see AI make it more possible to improve complex pieces of software than has historically been practical.

[0] https://github.com/malisper/pgrust [1] https://malisper.me/the-four-horsemen-behind-thousands-of-po...

malisper··on Zig → Rust porting guide
> what looks like a massive undertaking for vibe coding

fwiw, I suspect it's less of an undertaking than you may think. I've been playing with AI to rewrite Postgres in Rust[0] over the past couple of weeks and I found the AI to be exceptional at doing rewrites. Having an existing codebase you can reference prevents a lot of the problems you have with vibecoding. You have an existing architecture that works well and have a test suite that you can test against

Over the course of a month I've gone from nothing to passing over 95% of the Postgres test suite. Given Jarred built Bun, I bet he'll be able to go much faster

[0] https://github.com/malisper/pgrust

malisper··on Uber torches 2026 AI budget on Claude Code in four months
I've been working on a project to build a new Postgres based database in Rust[0]. I'm four weeks in and have 93% of the Postgres test suite passing. I've found agents to have worked really well for this as I have an existing codebase that has good architecture that I can point my agents at. It's also easy to debug as I can diff what my agents are doing and what Postgres is doing.

I've had to get multiple codex accounts, but there was a brief period of time where I tried API usage to see how expensive it would be. In about an hour I spent $650 of credits. I had codex estimate how much I would be spending if I was doing pure API usage and it estimated around $10k/week.

For context Postgres is 1M lines of C code. It's looking like pgrust will come out as less lines of code than Postgres and at peak I was adding over 100k lines of code in a day. I would estimate it would take a team of 5 software engineers at least 3 years to get to where I got in a month with a couple Codex subscriptions.

[0] https://github.com/malisper/pgrust

malisper··on Anthropic takes legal action against OpenCode
Since there's a lot of questions about what this means, let me explain.

Anthropic has two different products that are relevant here: the Claude API and Claude Code. The Claude API has usage based pricing. The more you use, the more you pay. With Claude Code, you can get a monthly subscription which gives you a fixed amount of usage. Comparing equivalent token generation between the Claude API and Claude Code, Claude Code with a subscription is much cheaper.

When it comes to third party products such as OpenClaw and OpenCode, Anthropic has made it clear those products should be using the Claude API and not the internal Claude Code APIs. OpenClaw and OpenCode have both been using the internal Claude Code APIs as when a user has a Claude Code subscription, the internal Claude Code API gives you tokens at a much cheaper rate than the Claude API. Presumably Anthropic makes Claude Code cheaper than the Claude API because they are willing to give users a discount for them to use Claude Code vs a competing product such as OpenCode.

It looks like until recently OpenCode tried to get around Anthropic's requirements by offering "plugins" in OpenCode that would allow users to use their Claude Code subscription in OpenCode. This PR mentions as much at[0][1]:

> There are plugins that allow you to use your Claude Pro/Max models with OpenCode. Anthropic explicitly prohibits this.

> Previous versions of OpenCode came bundled with these plugins but that is no longer the case as of 1.3.0

This PR seems to be in response to Anthropic threatening OpenCode with legal action if they keep using the internal Claude Code APIs.

  [0] https://github.com/anomalyco/opencode/pull/18186/changes#diff-b5d5affc6941bf7bb19805cc8f556cd1b9ae73ffd99e520120700536b166f8c0L310
  [1] https://github.com/anomalyco/opencode/pull/18186/changes#diff-b5d5affc6941bf7bb19805cc8f556cd1b9ae73ffd99e520120700536b166f8c0R321
← PreviousPage 3 of 16Next →