Omnigres: Postgres as a Platform
github.com
github.com
Just a quick note to the readers: we're still in the early days, in a pre-release mode. A lot of things work, but not everything is there yet and there are some bugs (we know some for sure!). A lot to come as we progress: first-class Python support, CLI, schema & package management, etc.
Our initial users are using it with our support and our timeline is sometimes advised by their milestones.
That said, happy to see you try it out and please join our Discord (https://discord.omnigr.es/) if you want to chat.
So i Tried omnigress to build a simple api that does few db calls a month back to mock real world scenario. And damn, the numbers were quite thought provoking. Comparing those numbers with a similar api built using fastapi and asyncpg, there was if i remember correctly at least 4x to 5x throughput increase.
Even though the project is in quite nascent stage, the idea is promising. This tackles the concept or should i rephrase, pain of having multiple platform to run a service and the needed expertise that is necessary.
If the scaling part and dev convenience can be fully sorted, and also the security aspect like many have highlighted, this could very well have the potential to change the game.
I wonder to what extent this improvement comes from not using Python, as opposed to using this alternative solution? For example after Python to Go migrations it’s common to see this kind of speed up I believe.
An important thing here is that we see Omnigres as a polyglot runtime with a database inside (Postgres) and we want people to use languages they prefer.
https://yrashk.com/blog/2023/02/16/what-happens-if-you-put-h...
In the instance of short-living queries, we are able to get this by effectively removing the communication latency.
Many thought RDBMS is dead (at least 10 years ago), well here we are. Dynamic languages were all the craze and now pretty much all those languages added type safety (Typescript, recent Python versions)
Looks like what goes around comes around.
Folks like Martin Casado noticing the trend [1] and explained what was the challenge [2]
[1] https://twitter.com/martin_casado/status/1705969880225513815
[2] https://twitter.com/martin_casado/status/1706045633571041731
Why would someone use this? Because they only want to deal with "one thing" in production. Dealing with multiple stacks is annoying. People show stuff like kubernetes as a way to handle multiple stacks. That stuff is also complicated.
There is still a huge opportunity for a methodology (accessible to non-operators) to get systems in production easily, where we wouldn't feel the need to force it all into these kinds of models.
Having said all of that, this is pretty fun looking!
However I doubt shoving everything into the DB will get us there - besides some hardcore PG enthusiasts like OP & co.
Sqlite together with the application layer of your choice seems to make much more sense to me, but hey we don't have to all like the same stuff (better we don't).
Also without reading too much in detail but I wonder how testing/monitoring etc would be supported
Nomad is a simpler way than K8S to run not just containers, but plain old binaries. Can’t beat K8S network effect. Technical merit or economical efficiency are secondary factors in tech trends.
- In my experience the performance of the stateful DB server has been the biggest bottleneck when scaling - it's much easier to scale the stateless application servers which sit between end user requests and your DB in a traditional architecture. So usually I'm wanting to move as much work as possible away from the DB in order to squeeze the most out of it before needing to shard the DB or move to a different solution, rather than moving more responsibilities into the DB.
- It's frankly pretty scary to load a C extension into postgres which is opening ports and parsing requests etc - bugs in it could crash the server or open security holes, and if you were able to exploit a vulnerability you'd be able to grab any of the data in the DB and easily exfiltrate it. This would be less of an issue if using this for an internal service which isn't directly exposed to the internet, but it still could make it easier for an attacker to escalate their access. (This isn't a 'it should be rust' comment really, even if this was in rust it would still be pretty worrying).
- Even if think you only need simple CRUD actions, over time you tend to need more and more logic around those actions. Authentication, verification, triggering processes in other systems, maybe you make schema changes and need to adapt requests from old clients, etc. It's really nice to have a more heavyweight application server where you can implement that logic - I'm pretty skeptical that row level permissions, triggers etc will be able to cleanly handle all those as you add new requirements over time. This applies also to other tools for more directly exposing your DB ( e.g. PostgREST ). IMO starting off using a tool like this is really just setting you up to have to do a pretty painful rebuild later on.
Am I missing something here, maybe I have misunderstood the intended use case?
- about scaling: you have to get very far before saturating a single postgres server. A lot of applications certainly do get to that point, but most don't. And once you get there, scaling postgres is definitely more work than scaling a stateless service, but it also gives you a lot more in terms of performance, reliability and further scalability.
- about C being scary: as a rust afficionado, I am not going to contradict you. But postgres itself is already C, and @yrashk is not just any C developer. Notably, he contributes to postgres itself.
- about managing complexity: postgREST, Omnigres, hasura, SQLPage and other tools that simplify building directly on top of the database never require exclusive access to the database. You can always put some of the complexity outside if you need to, when you need to.
Thank you for saying this. Mostly when postgres as a platform is brought up, the horizontally scaling ppl will often mention the parent thread. I have being developing since Apple computer had floppy disk, and there is rarely many situations where I need to saturating a single postgres server.
And even if one did get to the situation where that happens, with introduction of hydra, or other postgers columnar db, we can just put that in. Most user will never get to a point where they need to saturating a single postgres server. And also keep in mind when processing large row data, writting stuff in middleware is just not as efficent or fast as in postgres when it has native access to data and data manipulation.
On scaling, yeah a single postgres server can handle a lot. For us we were well past the million user mark before running into serious issues. However, a lot of how we were able to keep postgres working for us as we grew was by shifting work from postgres to our stateless services like I alluded to before - e.g. making our SQL queries as simple to execute as possible even if it means more work for the client to piece the parts back together.
If everything had been running inside the database we wouldn't have had that option and we'd probably have hit scaling limits much earlier - I guess we could have split off the traffic to the highest traffic endpoints and have those handled by a separate service calling the PG db, but then you get into issues with keeping the authentication etc consistent.
Re security - yep, PG is already using C to parse untrusted inputs from the network, which is also scary, but it's (hopefully) well reviewed and mature code - and even so, I wouldn't want to expose PG's usual wire protocol port to the internet, so it's hard to imagine exposing HTTP from postgres to the wild west.
Ultimately it probably is just a question of the sort of project it's being used for - if it's for something that's not going to get need to get to larger scales, handle a lot of complexity over time, or pass security reviews and your main goal is simplicity, then maybe an approach like this is a good option. I've just found that things tend to start off looking small and simple and then turn out to be anything but, so I'd rather run `rails new` and point it at a standard PG server - which would be just as simple and productive when you are starting out, and can keep scaling as your customer base and team size grows up to the size of Shopify, Github, or Kami (shameless plug).
The project is interesting, and thought provoking, because it goes against the often recommended "good practice" of separating storage and compute. Doing the exact opposite has a lot to offer in terms of performance, simplicity, and speed of development.
I am currently also working on a database-first web application framework [1], with different goals and use cases, and I bet we'll see more of these in the future.
This gets me thinking... has developer conventional wisdom ever recommended binding things together? Or does it only ever recommend more separation, more abstraction?
It works and can scale far enough that it allowed entire corporate websites to be run from an oracle database.
I read the README, but it focuses on the practical aspects of getting the solution to work, not its philosophical justification.
edit: When I say "disaster" for the MS SQL Server attempt, I don't believe anyone ever experienced a breach, etc. But, as I recall, there were no large customers (MS caters to the larger corporate crowd) with IT deparments willing to risk exposing their DB servers to the "DMZ zone" of public Internet access (or only 1 layer past is, such as behind a proxy). So from that standpoint the disaster was providing a solution nobody (with money and corporate experience) wanted. I guess for hobbies or Silicon Valley MVPs, the concept might have some legitimate appeal.
I've always felt that this approach is the right way to build applications. Application servers seemed like an unnecessary layer to me, essentially serving as intermediaries that merely passed data from the database to the browser. In the past, they played a more critical role in generating HTML, but nowadays, application servers are primarily used for handling APIs. Consequently, they often lack meaningful tasks to justify their existence.
Having your code closely integrated with the data also has the benefit of improving performance.
Would not recommend this path, but that is probably because I understand deployments and platforms. if they can fix the ergonomics, modernize the process, I could see it working out.
I suppose replacing your OS with postgres could have its advantages.
The intention we have is to take what's right with the colocation approach and take the learnings of what was wrong with those before (such as the above or Illustra).
Essentially, we're aiming create a modern polyglot runtime that has an embedded database with all its functionality (Postgres), with a contemporary, lightweight D3X.
As you rightly point out, the idea behind Omnigres is to make a lot of these concepts available broadly.
Is this something that omnigres addresses?
Having http server tightly coupled with the database server makes for a responsive combo.
We are still blocking workers, but prototyping improvements that will alleviate this.
I think you have a good project on hands, happy to see you could establish it as a tech startup. Over the years my mind keeps wandering into this place where an apps are much tigther integrated with the database. Where a full class of engineering tasks (and bugs caused by them) is eliminated completely. I'll be checking your progress.
This way we don't need to handle the difficult parts of JWT (forced expiration, etc.) and the mental model becomes rather simple.
Also, this [1] seems intriguing. How do containers connect to the db? What would the performance differences to the "internal" approach? Is this feature more like Lambda or for long running processes? Or something else? In any case very interesting.
Thank you!
[1] https://github.com/omnigres/omnigres/tree/master/extensions/...
We are adding first-class Python support right now. It's already possible to extrsct stored functions from decorated functions and their type hunts and we're working on providing standard Python APIs like DBAPI, WSGI support, etc. We have a branch on which we ran Flask applications inside Postgres. As it matures, it will be merged and documented.
As for the containers, they receive the database credentials over env variables currently. The performance characteristics of such applications aren't as good currently. The intended use case for this is third-party apps and legacy pieces of own applications.
Please keep in mind that this extension hasn't received much updates in the past couple of months but there will be upcoming changes to simplify ot a lot and provide more functionality. We also have a future experimental goal of going all the way through a runtime like crun to remove moving pieces.
Perhaps not same level as containers, but WASM runtime could be a powerful addition here. Or a container running said WASM code. I’m thinking more about untrusted client code for ad hoc data analysis and such.
One additional question: is Postgres foreign data wrappers going to be supported?
A good ressource to start exploring would be this site with compare features between vendors: https://modern-sql.com/
The interface we have with databases could be so much better.
I agree that something like INSERT INTO t VALUES a=1, b=2, c=3
would be much more readable.
I remember doing an experimental improvement allowing to do FROM where SELECT what syntax.
I agree there's value there. Omnigres is able to intercept query expressions to do the augmentation.
What’s the migration story? Migrations are what always trip me up, especially when there’s dependent logic.
However, we're not quite satisfied with this and working on a more sophisticated system that would allow us to derive incremental changes where possible, lint schema, load application code with the right dependencies on types and other functions where necessary.
Unless you mean edge as opposed to mainstream?
We believe that a practical edge [backend] requires the presence of data next to the code, which is precisely what Omnigres promotes.
We'll see what the future will bring, as the real edge of computing is in user's hands.
There's a keen interest to get there for Omnigres applications, but the shape of that is still somewhat vague.