HNHacker News
TopNewBestAskShowJobs

nikita

1,695 karma · joined October 19, 2010

CEO of Neon Partner at Khosla Ventures Ex CEO of Singlestore I really like database technology
submissionscomments
nikita··on PostgreSQL for Everything
Today CDC, tomorrow an authoritative part of storage.
nikita··on Postgres data stored in Parquet on S3: LTAP architecture explained
If you product is CDC based (peerdb) you don’t want storage to support this :)

This architecture is better for OLTP because all maintenance operations are moved to storage AND it has all other benefits such as LTAP that emerge from having a scalable storage.

nikita··on Postgres data stored in Parquet on S3: LTAP architecture explained
Conversion is async. The whole point is to never deal with CDC which is error prone and taxing Postgres with occupying a replication slot and burning memory and cpu in the OLTP system.
nikita··on Postgres data stored in Parquet on S3: LTAP architecture explained
Recent data plus working set is always in Postgres page format.

Historical data when pushed to s3 is in parquet. This happens async - not on the transaction hot path.

So older data below certain LSN is on s3 in parquet available to all analytics processing. Hot data is on page servers in page format for OLTP.

You can be smart in querying both representations for real time analytical queries

nikita··on Postgres data stored in Parquet on S3: LTAP architecture explained
Safekeepers keep a window of WAL in Postgres WAL format and doesn’t have an external API.

It streams WAL to pageservers and Postgres read replicas

nikita··on Lakebase architecture delivers faster Postgres writes
* Rate limiting on proxy in front of compute fleet

* Large tenants are broken up into shards, reducing hotspots

* Each shard is throttled to a fixed req/s rate

* We do not run pageservers at their redline in terms of CPU load, so there is some slack to take up bursts

* Capacity quotas which selectively throttle write traffic to the largest databases if they are competing with others for disk space, until the larger database is migrated away.

nikita··on Lakebase architecture delivers faster Postgres writes
Generally the more throughput the system supports the better. In this case we were hitting limits (btw each operation is many queries of different sizes) and the customer observed higher latencies which is typical if the system can't sustain the throughput required.

After this change latencies are back to normal and throughput increased.

nikita··on Lakebase architecture delivers faster Postgres writes
Lakebase is referring to the fact that in addition to disaggregated storage s3 is authoritative storage for older data.

Since data is on s3 (or lake) you can perform direct to s3 type operations like data loading, reading this data by engines that are not Postgres and more

nikita··on Lakebase architecture delivers faster Postgres writes
Lakebase is OLTP.
nikita··on Lakebase architecture delivers faster Postgres writes
We provide you fully managed Postgres. Lots of our customers use it for lots of small instances of Postgres since using Lakebase is so lightweight.

Small and large instances benefit from this performance optimization.

nikita··on Lakebase architecture delivers faster Postgres writes
This applies to our storage implementation. In Lakebase architecture storage serves pages and it doesn't always have the most recent version of the page and therefore it reconstructs it on demand.

In the past we relied on Postgres compute to periodically send a full page so reconstructive a page was always a bounded process. Once we turned it off (and got all those perf gains) we got another problem: unbounded page reconstruction which we had to solve separately.

nikita··on Lakebase architecture delivers faster Postgres writes
This specific perf improvement is orthogonal to HA.

However generally disaggregating storage makes HA simpler and allows for things like zero downtime patching: https://www.databricks.com/blog/zero-downtime-patching-lakeb...

Read replicas can be "shallow". You don't need to replicate all the data to create a replica. This allows to create them very very quickly (sub second).

All the extension still work. We don't support Citus today, but mostly because customers are not asking for it rather due to technical limitations. We support lots of extensions: https://docs.databricks.com/aws/en/oltp/projects/extensions

nikita··on Lakebase architecture delivers faster Postgres writes
I'm a VP on Databricks and former CEO of Neon. Happy to answer performance related or any other questions here.
nikita··on A recap on May/June stability at Neon
replit
nikita··on Bridged Indexes in OrioleDB: architecture, internals and everyday use?
(Neon CEO)

Not really. OrioleDB solve the vacuum problem with the introduction of the undo log. Neon gives you scale out storage which is in a way orthogonal to OrielDB. With some work you can run OrioleDB AND neon storage and get benefits of both.

nikita··on Fauna Service Winding Down
One of the services that can replace the fauna service is DocumentDB Postgres plugin (+proxy that is not open sourced yet, but will be shortly). It's available on Azure, but I can also see other Postgres Providers will start picking this up.

https://github.com/microsoft/documentdb

nikita··on Postgres Just Cracked the Top Fastest Databases for Analytics
This is an exciting project. Few highlights: - Query processor is DuckDB - as long as it translates PG type system to DuckDB typesystem well - it will be very fast. - Data is stored on S3 in Parquet with Delta or Iceberg metadata. This is really cool. You don't need to push analytical data through WAL - only metadata goes into WAL. This mean fast loading at least in theory, and compatibility with all the Delta/Iceberg ecosystem. - Once they build real-time ingest, you can just push timeseries into this system and you don't need a second system like Clickhouse
nikita··on Microsoft open sources PostgreSQL extensions to muscle in on NoSQL
The word is it's a serious effort on the part of Microsoft. It's missing a MongoDB wire protocol which they plan to opensource as well. In the meantime it's possible to use FerretDB for that.

I think the biggest use case is big data and dev platform that need application compatibility and wrapping Atlas is less attractive for some reason.

nikita··on PostgreSQL is the Database Management System of the Year 2024
Oh, that's coming very very soon! London and San Paolo. Def in Q1.
nikita··on PostgreSQL is the Database Management System of the Year 2024
(neon ceo). What region are you in?
nikita··on QwQ: Alibaba's O1-like reasoning LLM
I remember having a hard time solving this in my first year of college. This is really impressive.

So I have this problem here: given that for positive integers m and n, 0 ≤ x[m + n] ≤ x[m] + x[n], I need to prove that the limit of x[n]/n exists. Hmm, okay, let's see how to approach this.

…

Alternative approach: maybe I can show that a_n is convergent by showing that it's a Cauchy sequence.

But that might be complicated.

Alternatively, perhaps I can use the subadditivity to bound a_n.

Wait, let's consider the following:

Since x is subadditive, x[m + n] ≤ x[m] + x[n].

Let me set m = n, then x[2n] ≤ x[n] + x[n] = 2x[n].

Therefore, x[2n]/(2n) ≤ x[n]/n.

So, a_{2n} ≤ a_n.

That means that a_{2n} ≤ a_n.

Similarly, x[3n] ≤ x[2n] + x[n] ≤ 2x[n] + x[n] = 3x[n].

So, x[3n]/(3n) ≤ x[n]/n.

Similarly, x[4n]/(4n) ≤ x[2n]/(2n) ≤ x[n]/n.

It seems like a_{kn} ≤ a_n for any integer k ≥ 1.

This suggests that as n increases, a_n tends to decrease or stay the same.

But wait, in the example where x[n] = n, a_n = 1 for all n, so it's constant.

In the example where x[n] = sqrt(n), a_n = 1/sqrt(n), which decreases to 0.

So, in some cases, a_n is decreasing, and in some cases, it's constant.

But in general, a_n is bounded below by 0 and is decreasing along powers of 2, for example.

Wait, from x[2n]/(2n) ≤ x[n]/n, and x[4n]/(4n) ≤ x[2n]/(2n), and so on, so along the sequence n, 2n, 4n, 8n, etc., a_n is decreasing.

Maybe I can use this to show that a_n has a limit.

Let me consider n = 2^k, for integer k.

Let b_k = x[2^k]/2^k.

Then, from x[2^{k+1}] ≤ x[2^k] + x[2^k] = 2x[2^k], so b_{k+1} = x[2^{k+1}]/2^{k+1} ≤ (2x[2^k])/2^{k+1} = x[2^k]/2^k = b_k.

So, b_{k+1} ≤ b_k, meaning that b_k is a decreasing sequence.

Since b_k ≥ 0, it must converge to some limit, say L.

Now, I need to show that x[n]/n approaches L as n approaches infinity.

Given that b_k is decreasing and converges to L, and for general n, a_n is bounded below by L, and perhaps approaches L.

Wait, more carefully:

For any n, there exists k such that 2^{k-1} < n ≤ 2^k.

Then, x[n]/n ≤ x[2^k]/2^{k-1} = 2(x[2^k]/2^k) = 2b_k.

Since b_k approaches L, for large k, b_k is close to L.

But I need to show that x[n]/n is close to L.

Wait, perhaps I can bound x[n]/n between b_k and something else.

Given that 2^{k-1} < n ≤ 2^k, then x[n] ≤ x[2^{k-1}] + x[n - 2^{k-1}] ≤ x[2^{k-1}] + x[2^{k-1}] = 2x[2^{k-1}].

Therefore, x[n]/n ≤ (2x[2^{k-1}])/n ≤ (2x[2^{k-1}])/2^{k-1} = 2(x[2^{k-1}]/2^{k-1}) = 2b_{k-1}.

But b_{k-1} approaches L as k increases.

Similarly, since n > 2^{k-1}, x[n]/n > x[2^k]/2^k = b_k.

But b_k ≤ b_{k-1}, and both approach L.

Therefore, x[n]/n is squeezed between b_k and 2b_{k-1}, both of which approach L.

Therefore, x[n]/n approaches L as n approaches infinity.

This suggests that the limit exists and is equal to L.

Therefore, the limit of x[n]/n exists.

Alternatively, if I recall correctly, this is exactly what Fekete's lemma states.

So, in conclusion, the limit of x[n]/n exists.

*Final Answer*

\boxed{\lim_{n \to \infty} \frac{x[n]}{n} \text{ exists}}

nikita··on Neon's Free plan now includes 10 projects
Neon CEO...

This is possible from the cost perspective due to the architecture of Neon and the ability scale to 0. Happy to answer questions!

nikita··on PostgreSQL 17
(neon ceo). We have lots of examples of this. Here is one with Retool. https://neon.tech/blog/how-retool-uses-retool-and-the-neon-a...
nikita··on PostgreSQL 17
A number of features stood out to me in this release:

1. Chipping away more at vacuum. Fundamentally Postgres doesn't have undo log and therefore has to have vacuum. It's a trade-off of fast recovery vs well.. having to vacuum. The unfortunate part about vacuum is that it adds load to the system exactly when the system needs all the resources. I hope one day people stop knowing that vacuum exists, we are one step closer, but not there.

2. Performance gets better and not worse. Mark Callaghan blogs about MySQL and Postgres performance changes over time and MySQL keep regressing performance while Postgres keeps improving.

https://x.com/MarkCallaghanDB https://smalldatum.blogspot.com/

3. JSON. Postgres keep improving QOL for the interop with JS and TS.

4. Logical replication is becoming a super robust way of moving data in and out. This is very useful when you move data from one instance to another especially if version numbers don't match. Recently we have been using it to move at the speed of 1Gb/s

5. Optimizer. The better the optimizer the less you think about the optimizer. According to the research community SQL Server has the best optimizer. It's very encouraging that every release PG Optimizer gets better.

nikita··on Going open-source as a VC-Backed company
At neon we only worry about hyperscalers particularly Amazon. But they already have Aurora so we just open source everything under Apache 2.0

Being open is extremely important to us to build trust and we had this since day 1. VCs are fine with it because monetization is all cloud

nikita··on Top features in Postgres 17 (plus our contributions)
CEO of Neon here. Happy to answer question about Postgres development.
nikita··on Understanding Neon's Autoscaling Algorithm
Neon CEO here.

Happy to answer any questions

nikita··on pg_duckdb: Splicing Duck and Elephant DNA
Are you guys planning to opensource your work at crunchy?
nikita··on pg_duckdb: Splicing Duck and Elephant DNA
Very soon, it needs to get a bit more stable.
nikita··on pg_duckdb: Splicing Duck and Elephant DNA
Postgres IS missing an analytics engine. benchmark.clickhouse.com puts it at the bottom of the list and ~1000x slower than @duckdb and @ClickHouseDB.

Here are the scenarios and how to address them

1. Query Parquet and Iceberg from Postgres. When Parquet files are stored in S3 Postgres should be able to run analytical queries on them.

2. Postgres should allow creation of columnstore tables inside Postgres storage subsystem. Analytical queries on top of these table should be FAST. Top 10 on Clickbench fast. This allows to run analytics without S3 and have super low latencies for analytics.

3. Postgres should allow creation of secondary columnstore indexes to speed up analytical queries in mixed workloads. This is super useful for Oracle migrations since Oracle had this feature for a while.

So How do we get there? 10 years ago it would be a MASSIVE project, but today we have @duckdb - super fast analytical engine with an open license. The work is still not trivial, but it is much much simpler.

First you need to integrate an analytical query processor into Postgres and today @duckdblabs announced github.com/duckdb/pg_duck…. Yay and congrats!

This plugin runs duckdb alongside with Postgres and integrated Postgres syntax with the @duckdb query processor (QP)

With that it now can trivially query external files from S3. This addresses scenario 1.

With that it now can trivially query external files from S3. This addresses scenario 1.

Building columnar table requires either implementing columnar storage from scratch or integrating duckdb storage into the Postgres subsystem. You can of course let duckdb create duckdb files on local disk, but then all the Postgres machinery: replication, backup, recovery won't work

Duckdb tables have to mapped into 8kb Postgres pages pushed through the Postgres WAL for replication, recovery and transactionality. This will give us scenario 2

Scenario 3 is even more work. You need secondary index maintenance and it will require hybrid query execution. We will need to modify Postgres executor so that it can mix and match regular Postgres query operators and "vectorized" query operators from duckdb. Or built vectorized operators into Postgres

Scenarios 2 and 3 will take some time, but I'm excited for this roadmap: this will unlock a huge world for millions of Postgres users and simplify the lives of many developers dealing with moving data between transactional and analytical systems.

Page 1 of 13Next →