Select ’Hello, World’: Serverless Postgres Built for the Cloud
neon.tech
neon.tech
I'm curious about the business side of things, to give some context, some open source models that startups use today are:
1. The "open core" model, Gitlab being a good example. They try to split features that are open or closed/enterprise depending on the buyer.
2. The AGPL model, Mongodb used to do this, today a popular example is Grafana and their collection of products.
3. The Apache + cloud backend model, the core being standalone working with Apache license while building a value added managed service. I think this is what Synadia is doing with NATS.
4. The "source available" model, not really open source, but worth mentioning as it's very popular recently. Examples Mongodb, Elastic, Cochroachdb and TimescaleDB. This is often combined with open source such that some parts are open source, others source available.
With this as a reference Nikita, how would you explain how Neon thinks in regards to licensing and eventually building a healthy business? It's obvious a managed database service is the money maker, but how do you think around compeditors taking the project and building managed services without or with minimal code contributions? I'm sure you guys have thought a lot about this, would be interesting to hear some thoughts and reasoning for or against different options.
(Note: This is not meant to be an extensive explanation of these business models just a high level overview. If I have miscategorized some company above feel free to correct me in a comment.)
Please do apply! we are hiring around the globe!
Do you expect candidates to have background in database development?
Also, any advice for someone looking to transition their career from backend to database development?
Neon is open source, and in Rust, I will start with that and would try to contribute.
Also, do you ask leet code style questions in the 1:1 rounds? More details would be appreciated
One thing I haven't seen with "serverless" databases is an easy way to dictate where data is stored. Mongo has a pretty clever mechanism in their global clusters to specify which regions/nodes/tags a document is stored in, and then you simply specify you want to connect to nearest cluster. Assuming your compute is only dealing with documents in the same region as the incoming request, this ends up working really well where you have a single, multi-region db, but in practice reads/writes go to the nearest node if the data model is planned accordingly.
A real world example of how I am using this in Mongo today: I run automated trading software that is deployed to many AWS regions, in order to trade as close to the respective exchange as possible. I tag each order, trade, position, etc. with the exchange region that it belongs to and I get really fast reads and writes because those documents are going to the closest node in same AWS region. The big win here is this is a single cluster, so my admin dashboard can still easily just connect to one cluster and query across all of these regions without changing any application code. Of course these admin/analytics queries are slower but absolutely worth the trade off.
We are working on deploying the proxy globally and routing read traffic to the nearest region.
We also have some multi-master designs in collaboration with Dan Abadi. But this will take a second to build.
db-uri = "postgres://authenticator:mysecretpassword@project_name.cloud.neon.tech:5433/postgres" db-schema = "api" db-anon-role = "web_anon"
I always liked this model since it gives the best isolation and not having to code with multi-tenant in mind really simplifies development and security. But maintaing an infra for this is hard and costly.
[1] https://docs.microsoft.com/en-us/azure/architecture/isv/appl...
This makes it super easy.
How does the memory cache work? The architecture looks great, but so much of PG efficiency is cache hit rate. Do the pageservers have a distribution plan? Plans to add something like DynamoDB Accelerator?
I have a question about performance though. Let's say I have a table with millions of rows. Would it automatically scale to resources to get me sub second queries on it?
Ultimately that's what I desire in such a solution. Something that will let me throw whatever at it and will handle the scaling to keep whatever queries I throw at it fast.
Everybody today knows that if you're gonna change your database [schema], you need a system to migrate the DDL changes intelligently so you don't break something, and then run that in lower environments to test, then promote it up to higher envs and run it, then deploy your apps that use the changes. But of course it may be impossible to revert those changes after that point, requiring an entire database snapshot restore. So not only are there serious operational concerns, making any operations around this time-consuming and frought with peril, but you need to set up a migration solution (language-specific or framework-specific or agnostic) and make sure you architect your application to only make changes in a specific way. (all of this, by the way, is only necessary because the database is one big mutable state machine)
...whereas if it worked more like version control, you could make any change, commit it, and get a commit ID. If that causes problems, if you could just `cloudpg revert $change_id`, then there would be no need to carefully architect the app, changes wouldn't be fraught with peril, and we could be more agile with database-driven design. The database would obviously need to be intelligent enough to figure out how to revert any change, which is why this has to be a database-specific feature and not just a git revert.
Use Case 2. Merging Changes
This sort of follows on the above (making changes more agile). If you have branches, and 4 different devs are working on 4 different database changes, how do you merge and deploy those all safely, and handle reversions safely? Well if all database changes had versions, and we could diff the changes between versions, then we could treat the database like a Git repo and merge/rebase all the changes to the database together at the same time as the code. Again, no need to go back and refactor migration scripts or the app design, because the database is essentially just version-controlled code.
The same things would apply to upgrading/downgrading database versions, bringing up or restoring new servers, possibly even making replication easier, maybe other things we haven't thought of yet.
There would probably need to be a fallback mechanism if, for example, a new column was created and new data was entered into it, and then the revert removes the column. Probably it could keep pointers to such things ("there is a database D with a table T with a column C and rows [a,b,c,d]"), so that if the change is re-reverted later, an extra merge instruction could pop the reference back into place like nothing happened. Somebody with an actual CS background must have better ideas than me :)
* Make a schema change
* push into prod
* accumulate some new data
* revert just the schema change
In our discussions we call this separating schema and data and allow you to have different schemas on the same data.
It's tricky to do in Neon due to the fact that storage knows nothing about schemas. It stores page with no idea what's on them.
But we have some ideas how to do this with logical replication where we will run a transform on top of logical replication stream to keep two branches in sync. Not this year though.
Doing so for a database seems less desirable from an availability perspective, especially with high-throughput databases.
But, because SQL has conflict resolution by cancellation of one of the two conflicting modifications, I don't think that it is reasonable (or even possible) to merge 2 divergent databases in a single way that always conforms to the needs of the developer and/or application.
Seems like there's no upper limit for scaling up reads, just wondering how this architecture affects write throughput. Would love to hear more!
I think we can do a lot of good things here over time and have plenty ideas. But for now it’s a single writer system. Good news is that there is so much open source tech around Postgres that it might not be a gargantuan task in the future
I have a hunch that Pageservers contain the disk pages, so storing the B-tree (or may be LSM) pages and compute traverses those pages to find the relevant page/pages. I am curious about how does it fit in together and fetches from disk/Pageservers work
I'm sure Neon's design can handle the hobbyist wordpress DB or personal project DB for <=$10/month since it can scale down to 0, and it likely would not be terrible for a cold-start on a small personal website to be 1-3 seconds for the first page load -- sure, bad for SEO, but I think a LOT of people want a DB that is managed and simple for <$20/month -- and again, I'm very hopeful Neon will have a nice pricing number here when they figure out their model, since a very tiny shared tenant should be very inexpensive for their model.
Multi-cloud and open source.
1) roll back long transactions and enforce upscale
2) wait for a better moment to upscale (potentially forever)
3) try to do a live migration of running Postgres to another node (like VM live migrations, or CRIU-like process migration) and preserve long-running transaction
So far, we plan to start with some combination of 1+2 -- should be fine for web/OLTP kind of load. But ultimately, we want to arrive at 3), but that approach has way more technical risks.
Still, there are no technical limitations. Our test suite already uses Neon-specific SQL functions from a C extension (https://github.com/neondatabase/postgres/tree/7faa67c3ca53fc...). At the very least, providing a lot of popular extensions out-of-the-box is on our roadmap once we figure out the security, no special repacking needed. As compute nodes should already be pretty isolated from each other, I don't think allowing arbitrary code will require a redesign.
CockroachDB supports the PostgreSQL wire protocol and the majority of PostgreSQL syntax. Not 100% compatible but most stuff works the same.
https://www.cockroachlabs.com/docs/stable/postgresql-compati...
Regardless, what are some of the trade offs this implementation makes? (aka the cons)
- Buffer cache evictions of freshly dirtied pages are bad. This is because we request pages with a hint on the latest change was evicted, and with newer and newer changes being evicted from buffers you might start to get limited by the write-through latency to Pageserver instead of only Safekeeper.
- Commit latency can be not great due to cross-AZ communication -- with 3 safekeepers in as many AZs, the second slowest response is the limiting factor
- Write amplification in the whole system is quite high. Plain PG does ~ 2x (WAL + Page), while we have many times that. Admittedly, these writes are spread around several systems, where any one system only really needs to write 1x WAL volume for the data that it is responsible for, but Pageserver is currently configured for something like 4x write amplification due to 4 stages of LSM-tree compaction.
So in layman terms, cache purging is slow (unreliable?) and writing is slower. (which is to be expected)
Maybe another way of begging the question is, what apps would be suitable to build with Neon and what would not?
Anything with a write-based working set that fits in the buffers of the primary instance;
Databases that are unused most of the time;
Databases with a lot of read replicas (potentially with Neon only providing the read replicas, not the write node/hot standby);
Apps that want to run analytics on the data, but don't want to transfer O(datasize) data to a different DB every time, and also don't want to deal with the problems of long-running transactions.
Not very suitable:
OLTP with writeset that doesn't fit in the caches;
Databases that need <1ms commit latency.
Seems like from a cost standpoint Neon does a good job of avoiding EBS entirely and thus this particular problem would be sidestepped. Is that right?
Looking at the design they are going for, I expect they can scale horizontally really well. I'm really looking forward to this as a low-cost hobbyist DB
Next, we have not yet optimized for replication outside Neon Cloud, nor do we have sideloading of extensions, so to use Debezium you'd have to self-host Neon for now.
You can already build and run all base components on your local machine, see instructions here: https://github.com/neondatabase/neon/tree/d11c9f9fcb950ac263... . You can run the tests instead if you want more insight into how a particular piece works. You probably want to attach your S3-compatible storage to Pageserver and Safekeepers; some Ansible scripts with command-line flags are in the repository.
To run Neon components on multiple machines, you should be able to create the `.neon` data directory via `neon_local init` and then share the generated configuration files across machines and tweak network settings. You can refer to our documentation to understand the terminology and the intended hosting configuration: https://neon.tech/docs/storage-engine/architecture-overview/
However, there are still two missing bits: the self-hosting documentation and the Neon Control Plane (web UI + K8S-based compute nodes orchestrator). So you don't get automatic scale-to-zero at the moment out-of-the-box, although all hooks and the PostgreSQL proxy we're using at pg.neon.tech are there.
We consider open-sourcing the Control Plane, so stay tuned. As for documentation and support, Nikita has already answered.
We are not planning to support on prem deployments commercially. We are working with partners like Percona to do it eventually, but those conversations are too early to commit to anything.
How does using local SSD instead of default EBS help achieve higher throughput though? I get it's cheaper and lower latency but I fail to see how the rebalance magic solves the throughput tradeoff for all unless they throw more ram at it.
The rule of thumb of software cloud margins for infrastructure is 60-70%. So we need to be there over time at least. So sorry for non answer, I hope we will sort it out soon.