HNHacker News
TopNewBestAskShowJobs

malisper

3,231 karma · joined December 31, 2013

Hi! I'm Michael Malis. I'm the co-creator of pgrust (https://github.com/malisper/pgrust/)

I previously ran Freshpaint (YC S19) for 7 years and before that I led the database team at Heap.

GitHub: https://github.com/malisper/

Blog: http://malisper.me

Email: michaelmalis2@gmail.com

submissionscomments
malisper··on RIP, vector database
> Your design choice went from a Postgres design pattern to a Mysql one. The difference is the reindexing cost vs the lookup cost - Postgres optimized for lookup and Mysql does for indexing on writes. Or more accurately, Postgres was better with good schema design using joins & mysql was optimized for a bad design with less normalization where many indexes exist for the same table

You are right that MySQL does better when you have lots of indexes, but I don't think the tradeoff is that the overall Postgres architecture is better with good schema design.

Having secondary indexes point the primary key enables things like undo logging, which obviates the need for vacuums - vacuums being the most painful part of Postgres. On top of that your primary key index will be mostly cached so the cost of the indirection is much smaller than it may first appear

malisper··on It's Time to Investigate the AI Labs
> Why aren't they running agents on isolated computers without Internet access?

They probably will soon, but even air-gapping may not be enough. See stuxnet for instance

malisper··on Training a 4B model to produce 81% faster query plans than Postgres
Even if the query plan was not generated by an llm, you can't verify it will run in an acceptable time. This is one of the biggest unsolved problems in databases
malisper··on Training a 4B model to produce 81% faster query plans than Postgres
> That bug is fixable and verifiable

A bad query plan is not your typical kind of bug. I would definitely not call it fixable. Query planners are inherently dealing with estimations and approximations. If the query planners estimation is off, you're screwed.

Unless you come up with a way to cheaply determine exactly how many rows a query will return, bad query plans will still exist.

malisper··on Training a 4B model to produce 81% faster query plans than Postgres
Funnily enough, you could replace "LLM query planner" with just "query planner" and this comment would still hold true
malisper··on 118M Queries per Second on Neki
> I watched an interview that Casey Muratori did with Tyler Cloutier (SpacetimeDB founder and spokesperson) [1]. One of the points that Tyler is that, given modern CPU architecture with cache lines, a distributed database needs to fan out to at least 50-100 nodes to beat the throughput of a cache-optimized, single node database.

Do you have the timestamp where they are talking about this? The claim doesn't pass the smell test for me. If you're talking about latency, then perhaps. On throughput, I don't understand how a single-node system could deliver higher throughput than a three-node system

malisper··on Claude Fable 5.1
Most of the work I've done with pgrust hasn't had issues with Fable. The only time I've had issues is when building a fuzz tester to find bugs
malisper··on JIT Compiling Code in 5μs
Eval would count as JIT compilation though[0]

[0] https://www.sbcl.org/manual/#compiler-only-implementation

malisper··on JIT Compiling Code in 5μs
Author here. Let me know if you have any questions about the post or about pgrust.
malisper··on JIT Compiling Code in 5μs
> There's no rarity of JITs, it's just that LLVM (and other frameworks) are often used

Except that using LLVM has high latency limitting it's applicability. Postgres just disabled LLVM by default because of this[0].

[0] https://www.postgresql.org/message-id/E1w8GWU-002bSL-31%40ge...

malisper··on There's no reason for software to be slow anymore
> This person doesn't understand how to make efficient code

The author is one of the most knowledgeable people about performance there is

malisper··on A decades-old bug in Knuth's long division (TAOCP Vol II, Algorithm 4.3.1D)
You still get a physical piece of paper that looks like a check; it's just not a valid check.
malisper··on Making Postgres 300x faster for analytics: batching, operator fusion, and SIMD
We've fixed the issue on our development branch. We've only been working with the default collation so when you try a different one it causes an issue
malisper··on Making Postgres 300x faster for analytics: batching, operator fusion, and SIMD
How so? Window functions work exactly the same way. Postgres processes them one row at a time, and you can batch them the same way as you would with sum.

Window functions do make parallel queries more difficult, but that's a different story.

malisper··on Making Postgres 300x faster for analytics: batching, operator fusion, and SIMD
> You’re cosplaying as someone who could build a database.

> In reality you yourself could never build Postgres or Redis or any other database.

> You’re not skilled enough or knowledgeable enough and you’re not willing to put the time in so you’ve simple vibe copied Postgres.

What makes you think these things? I've been writing about the Postgres internals and giving talks about Postgres for a decade and have managed a Postgres cluster as big as 1PB of data

malisper··on Making Postgres 300x faster for analytics: batching, operator fusion, and SIMD
Can you file an issue? We know there are bugs and the work we're doing with formal verification and fuzz testing is to go through all the code and make sure all of it behaves identically to Postgres
malisper··on Making Postgres 300x faster for analytics: batching, operator fusion, and SIMD
We don't expose the priorities right now, but we have the priorities decay over time. That way faster queries get prioritized over long running queries. That should achieve the behavior you're looking for.
malisper··on Making Postgres 300x faster for analytics: batching, operator fusion, and SIMD
Can you file an issue? Our big focus over the next few weeks is to eliminate these issues and that's why we're our formal verification and fuzz testing work
malisper··on Making Postgres 300x faster for analytics: batching, operator fusion, and SIMD
Yes. Email me and Jason at malis@pgrust.com and jason@pgrust.com
malisper··on Making Postgres 300x faster for analytics: batching, operator fusion, and SIMD
I'm certain pgrust can find a long term home somewhere
malisper··on Making Postgres 300x faster for analytics: batching, operator fusion, and SIMD
Thanks!
malisper··on Making Postgres 300x faster for analytics: batching, operator fusion, and SIMD
What do you mean by native TTL? Would that be when rows are automatically deleted if they aren't touched after a certain period of time?
malisper··on Making Postgres 300x faster for analytics: batching, operator fusion, and SIMD
^For context, this is Greg Smith, the author of Postgres 9.0 High Performance[0]. That book was my first introduction to Postgres

[0] https://www.amazon.com/dp/184951030X

malisper··on Making Postgres 300x faster for analytics: batching, operator fusion, and SIMD
> My understanding is that they used c2rust, and then told the LLM/Agents to make the code more idiomatic rust, while also using the PG test suite as a feedback mechanism.

This is correct

> and thus a fork, and should thus have the original license and copyright preserved)

This is not correct. The Postgres license is permissive. We need to include a copy of the license (which we do in the NOTICE file[0]) but we CAN relicense the Postgres code however we want as long as we meet the requirements of the license. pgrust is a derived work of Postgres, but Postgres allows derived works to be under a different license.

[0] https://github.com/malisper/pgrust/blob/main/NOTICE

malisper··on Making Postgres 300x faster for analytics: batching, operator fusion, and SIMD
It absolutely can be embedded. The bigger enabler is replacing the process-per-connection model with a thread-per-connection model. Projects like pglite[0] had to give up concurrency because of it. We also support compiling to wasm so you can embed it in the browser too, which is what powers pgrust.com

[0] https://github.com/electric-sql/pglite

[1] https://pgrust.com/

malisper··on Making Postgres 300x faster for analytics: batching, operator fusion, and SIMD
> I am very disappointed to see the direction: It is moving from a "interesting attempt to recreate system software" to "building flashy but useless demo"

What makes you say this is a useless demo? I can't count the number of people who've struggled to do analytics inside of Postgres. Almost always they end up setting up a separate system such as Clickhouse and replicating the data between the two systems. Now they can have one system that's Postgres-compatible, and it's faster than either of the original systems.

> Everyone who knows a bit about databases knows the difference between execution models and what kind of optimization it brings.

In our last post[0], when we mentioned we were getting close to Clickhouse level performance (now faster than Clickhouse), we were met with disbelief. This post is meant to explain part of how we closed the 300x gap between Postgres and Clickhouse. The execution model being 10x of it.

[0] https://news.ycombinator.com/item?id=48841676

malisper··on Making Postgres 300x faster for analytics: batching, operator fusion, and SIMD
One of the new features we recently built is "test mode". This brings cloning a template db from 100ms down to <10ms making it much better for tests.

If you're interested in trying it out, please reach out to me at malis@pgrust.com

malisper··on Making Postgres 300x faster for analytics: batching, operator fusion, and SIMD
Me and Jason, the two people working on the project
malisper··on Making Postgres 300x faster for analytics: batching, operator fusion, and SIMD
We disabled parallelism in the blog post for demonstration purposes. The 300x slower refers to the clickbench numbers[0] where parallelism is enabled

[0] https://benchmark.clickhouse.com/#system=+liH|pgrs|gQ&type=-...

malisper··on Making Postgres 300x faster for analytics: batching, operator fusion, and SIMD
I would probably dig into the reasons for the differences in the benefit on the test machine and in prod

I had an issue like this for optimizing pgrust. I had an optimization that showed no impact on my test machine (c8g.4xl) and showed a 20% improvement when ran on my mac. It turns out the issue was the instruction cache on the c8g.4xl was being saturated on the test machine but not on my laptop, moving the bottleneck to a different place

If you can consistently reproduce the performance difference, you're already half way there

Page 1 of 16Next →