Quicker serverless Postgres connections
neon.tech
neon.tech
So we introduced so many optimizations just because we can't persist state. I can't help but think this is being penny wise pound foolish; the problems this article is solving wouldn't have been problems in the first place if you choose a boring architecture.
Serverless postgres-as-a-service isn't a stretch at all.
This thread does remind me of a previous thread, I believe it was something like optimizing a basic string operation (concatenation?) by going down to assembly and SIMD. Few people have a need to know the details of that particular implementation, and mere mortals cannot comprehend the wizardry going on there either, but that's just fine because if it gets incorporated into standard libraries or the underlying implementations that said standard libraries use, everybody benefits without ever having to understand the inner workings.
Heroku and similar things are the sweet spot right now for me. I don't have to deal with much, and I get exactly the amount of flexibility that I need. Sure it'd be a little nicer not to deal with Express on my own in a NodeJS environment, but right now it's easier to deal with than it is to avoid. Some day I'll happily go serverless.
There's always instances with state somewhere; it's just a question of how smart your serverless infrastructure is about pooling the instances.
So you have:
a) No enforceable isolation, reliance entirely on good behaviors and module systems
b) No independent scaling
c) No isolation of security components so instead of a minting service and an edge service you have one service that does both
d) Single language lock-in
I could go on and on an don
... of course, the best-paying jobs are at the places that really do need that scale, so of course we all want an excuse to play with those tools, even when it's not the right business choice. And it's not as if software folks are the only ones with a principal-agent problem, in business. Oh well.
I dislike using boring as a measure of sufficiency. It makes me wonder if those drawn to the clever solutions carry a "I know better than you" perspective, the kind of dogmatic optimism that tanked SVB.
With most services, until you go over a few million users concurrent users, you really don't need to horizontally scale.
There is ScyllaDB for a much better reimplementation of Cassandra/Dynamo, but wide-column databases are still best for niche scenarios, especially as RDBMS are rapidly evolving into natively distributed architectures.
Everything from StackOverflow to ad serving systems to high-frequency trading exchanges are run this way.
Quora, Pinterest, Twitter, etc are all just big app instances talking to DB instances, with separate systems for background processing, queues, and caching. Are you suggesting that they would scale better with serverless functions instead?
Here's a list of architectures: http://highscalability.com/all-time-favorites/
I'm curious if Cloudflare, or other serverless platforms offer the ability to generate short lived JWTs assigned to each worker. Combined with some configuration (pub key exchange, claim setup, etc) a platform like Neon could use these JWTs to establish identity. Sorta SSO for workers without the CPU overhead.
Seems like a safer approach than basic auth with fixed pws.
let value = await env.MY_KV.get(keyName);
No setup needed, at least in code. You create the binding either through the configuration UI or API.So far we've mostly used this technique to connect Workers to other Cloudflare-provided services like Workers KV, but I'm super-interested in the idea of third-party bindings. Hopefully, you'd be able to configure them through an OAuth-like flow, where the Cloudflare dashboard redirects you to the third-party service, that service prompts you for permission, then redirects back to Cloudflare, and you never have to copy/paste a single secret.
Of course, under the hood this would all be backed up by strong authentication, but it's high time we stop making application developers waste time thinking about this stuff.
(I'm the tech lead for Workers.)
https://developers.cloudflare.com/workers/runtime-apis/mtls/
I can't cite examples while I'm being ushered by a distracted Uber driver but this explanation goes beyond the documents.
Also, please extend Cloud Workers to be written as NanoVMs.
We implement the Postgres frontend with a forked version of PgBouncer, and we changed the authentication method such that when the user authenticates, we issue them a JWT which we store as a session variable. That session variable has the same security properties as a cookie in a web browser (the user can change/manipulate it, but if it's signed by us we can trust its claims).
That's the simple explanation that skips over the multi-tenant part. I don't want to derail from the thread - Neon is very cool, and we are actually experimenting with it right now, for storing the Seafowl [1] catalog [2] when deploying to "scale to zero" services like Google Cloud Run or AWS Lambda, which don't have persistent storage.
It should be a lot faster than a purposefully slow kdf AFAIK.
I reckon the current hype-cycle will be just about at its peak, when that happens.
(FWIW I like serverless—but then, I liked CGI, so of course I do)
The article ends with the statement that you are pretty much done here for now. Would optimizing your TLS termination not maybe offer some more ways to speed this up? Or is that also already fully optimized?
I did not realize before that your approach with Websockets actually meant that there was no application/client side pooling of connections. What made you choose this approach over an HTTP API (as for example PlanetScale did) anyway?
No, we don't do early termination yet, but it makes sense to try it out too. Here we mostly concentrated on how far we can get in terms of reducing number of round-trips.
> I did not realize before that your approach with Websockets actually meant that there was no application/client side pooling of connections. What made you choose this approach over an HTTP API (as for example PlanetScale did) anyway?
To keep compatibility with current code using postgres.js.
That makes a lot of sense - not needing an additional driver/client package is indeed a good point. Any plans to add a HTTP based API though anyway?
I have a number of workloads I'd migrate to you over the next few weeks if you're ready for it. What is the state of Neon?
Is it ready for small-volume but business critical loads?
I'm struggling to find support guarantees and SLAs on your website, do you have them yet?
What is your off-site backup story? Can I export my backups to something like R2, S3, or B2?
Many of my customers are running in Cloudflare Workers. Their volume is low, but it's business critical (DB downtime means business downtime - business will fail if we lose all the data).
> Is it ready for small-volume but business critical loads? Depends on business criticality. But generally yes. We are debating internally when we are going to announce it.
> I'm struggling to find support guarantees and SLAs on your website, do you have them yet? Good feedback. It's in the works.
> What is your off-site backup story? Can I export my backups to something like R2, S3, or B2? Since our storage is integrated with s3 and supports branching we can treat branches as backups. > Many of my customers are running in Cloudflare Workers. Their volume is low, but it's business critical (DB downtime means business downtime - business will fail if we lose all the data). We haven't lost data yet. There are a lot of redundancy in the system. Safekeepers, S3, branches (backups).
Neon[2] is Postgres that has a surgically enhanced data layer that reads from a distributed set of nodes and object storage, but is still a single Postgres instance that's started and stopped on-demand and not automatically sharded. Because it's real Postgres, all functionality works including extensions without dealing with issues from sharded datasets, but also without the horizontally scalability and instant start (so far). However because the data layer is improved, the actual compute node is much more efficient in both startup and processing so it works well in many scenarios.
CockroachDB[3] is a proprietary distributed database built to Postgres wire/data protocol with natively separated storage and compute. Their serverless plan also starts and stops instances as traffic comes in but with fast startup because of their specific architecture. Because it's not real Postgres though, there's a good bit of missing functionality.
There's also TiDB[4] from PingCap which is similar to a MySQL version of CockroachDB, and Yugabyte[5] which is somewhat like Neon but with both distributed compute and data layers also using real Postgres components - however neither has a serverless offering.
1. https://planetscale.com/ 2. https://neon.tech/ 3. https://www.cockroachlabs.com/ 4. https://www.pingcap.com/tidb/ 5. https://www.yugabyte.com/
This is a good analysis of network time and optimization. I'd love to see a followup exploring the impact on the server side; Considering how Postgres was historically not very efficient at setting up new connections but has been steadily getting better with each major release. It would also be great to see what portion of the request is taken up by the setup versus the query itself, to have some idea of cost.
In that case, do you manage the possibility of multiple clients opening a websocket connection to the same to “the same” server (which may actually be a different instance of the same proxy)
From there we are thinking to expose <project_id>.read.neon.tech and <project_id>.write.neon.tech and for read queries route traffic to the local replica.
It's not set in stone yet, but that's the gist of it.
- Data safety and privacy. As I'm UK based - can my data be restricted to the UK only (no magic round the world trips inside your backend)? do you comply with the GDPR?
- Can I set up postgres (WAL) replication to a non-neon database?
The idealist in me would rather this kind of thing become a standardized extension of address + transport protocols. But every person I talk to would rather stack 12 protocols on top of each other than work on improving the existing status quo. It feels like a truism of human societies; everyone wants improvement, but nobody wants fundamental reform.
Ed: oh, i see - the usecase is going from a serveless app to pg? I suppose any vm-like serveless node could be made to speak witeguard, too...
The problem that S3 doesn't love small files. So we organize pages in LSM trees on the pageservers and offload layers to S3.
I hadn't heard it before and am not finding references googling.
I'm curious what problems you have with small files on S3; or if someone wants to feel free to point me to a link to discussion of this apparently known fact!
Lots of stuff on the web. Quick googling found: https://www.upsolver.com/blog/small-file-problem-s3
Is the recommendation to have a per-dev branch off the main DB?
Can't stop being mad at the misuse of "serverless". Even worse than using crypto for cryptocurrencies. Words have meaning...
A true serverless solution, is for me something that takes care of this automatically by autoscaling / sharding. I should just have to point my database connection to an URL and it would not require any thoughts on my side. If I suddenly get a peek during Black Friday it should just handle it transparently.