Show HN: EdgeDB 1.0
edgedb.com
edgedb.com
I do have a couple minor quibbles:
- Serializable transactions are expensive, and that deserves to be an explicit caveat. Not everyone knows this, and it’s an important thing to put up front.
- Some of the language in this post are in CAP theorem territory but neglect to directly address that. I’d like to see how client usage compares with direct Postgres usage (idiomatic for each insofar as such a beast exists) in a Call Me Maybe. I know that’s a lot to ask in a 1.0 announcement four years in the making, but I hope it’s a priority to get this in front of Aphyr.
Edit: oh and I definitely look forward to this being further distinguished from an ORM, because even though I can see the blue and black dress my mind keeps switching it back to gold and white.
Thank you!
> Serializable transactions are expensive, and that deserves to be an explicit caveat. Not everyone knows this, and it’s an important thing to put up front.
We've not seen a major difference in our benchmarks (though maybe our benchmarks are wrong :-)). EdgeDB tends to produce very short transactions, so that helps. EdgeDB also knows if your statements are read-only or not, so we have the ability to steer these into a read-only transaction, though this isn't implemented yet.
- The overhead discussed in the docs, which is ~negligible for lots of use case and a perfectly reasonable tradeoff for those.
- The overhead of retries, which with appropriate defensiveness can effectively become an indefinite lock in, but undetected by, the client. When automated by an abstraction layer, this can become pathological pretty easily depending on usage patterns.
The most realistic alternatives are to provide a lower level abstraction (eg “I don’t want your guarantees, I want your errors”), or to provide other isolation options (eg “I don’t want your guarantees, I want my errors”). But there may well be opportunities here because EdgeDB knows as much as it does about the schema, and positions itself as a SQL replacement rather than a companion so it can potentially optimize for at least some of those cases at query time.
That sounds complex enough to boggle my mind, but if y’all are up to it I’ll be excited to see how it goes!
> We've not seen a major difference in our benchmarks (though maybe our benchmarks are wrong :-)).
You’ll likely not see anything noteworthy without specifically creating concurrency contention which specifically causes the kinds of pathological retry scenarios I mentioned. I’d be shocked if there isn’t at least a good starting point in either the Postgres test suite or Call Me Maybe. (Seriously though, I want to read Aphyr’s take on this project.)
Here's what I think makes EdgeDB special: it's a DB that replaces the tediousness of ORMs with a better core that can be cross-language / cross-platform. I've implemented tons of APIs, first REST, then GraphQL all of them on top of ORMs (Django, Peewee, SQLAlchemy, Mongoose and more). When prototyping was great, but scaling them became quite challenging, specially if you want to have a performant outcome when retrieving data.
EdgeQL is an incredible useful abstraction that will prove itself in a few years. Long life EdgeQL. Keep up the good work!
Does it compare to RethinkDB in this regard? The "fluent" native query language, ReQL, was one of its best parts.
Been keeping a close eye on Edge, had even considered it as a primary database, and probably will in the future!!
As much as I adore the ergonomics improvements I really am more interested in the performance, replication, scalability story, with the likes of cockroach db reaching maturity in 2022.
But as a postgres replacement in general, I would highly consider using edge.
it's as close as you can get to having actual magic sprinkles that make your code go faster.
My example/requirement: I have a user wanting to find best-matching blog posts. Every post is tagged with a given category. There could be 100+ categories in the blog system and a blog post could be tagged with any number of these system categories. A user wants to see all posts tagged with "angular", "nestjs", "cypress" and "nx". The resulting list should return and be sorted by the best matches, to those of least relevance. So, posts that include all four tags should be up top and as the user browses down the results, there are posts with less matching tags.
What I've seen with SQL looks expensive, especially if you search with more and more tags. I may just not know what to search for though, re. SQL. Is there a query against a graph database that could accomplish this?
I just happen to have a very similar requirement to yours and was also wondering.
> Every post is tagged with a given category. There could be 100+ categories in the blog system and a blog post could be tagged with any number of these system categories.
My point is, I don’t see SQL query as expensive for this kind of use case. There are easy and native ways to do it.
In case you would like a top notch performance, Redis might be a way to do it. Even a reverse-index would achieve great performance.
+---------+-------+---------+-------+-----------+
| post_id | title | content | other | fields... |
+--------+------+
| tag_id | name |
+------------+---------+--------+
| tagging_id | post_id | tag_id |
But it seems like the core of the request is still something like: SELECT post_id, count(1) AS count
FROM taggings
WHERE tag_id IN (3, 8, 255)
GROUP BY post_id
ORDER BY count DESC
(off the top of my head; I haven't checked this for any kind of correctness)And I don't see why that query suffers as you add tags...?
------------
EDIT responding to below [HN believes I am a problem user who should only be allowed to make so many comments per day]:
< that is pretty much what I meant by “I see just one table” as you don’t need any joins
Well, assuming you're doing this because a user is interacting with your site via some kind of web interface, you can set the interface up to deliver you tag_id values directly, but you'll still need to do a join with the posts table so you can present a list of posts back to the user instead of a list of internal post_id values.
So I guess
SELECT t.post_id, count(1) AS count, p.title, p.url
FROM taggings t JOIN posts p ON t.post_id = p.post_id
... with tag_names := {"angular", "nestjs", "cypress", "nx"},
select BlogPost {
title,
tag_names := .tags.name,
match_count := count((select .tags filter .name in tag_names))
}
order by .match_count desc;
Which would give you a result like this: [
{
title: 'All the frameworks!',
tag_names: ['angular', 'nestjs', 'cypress', 'nx'],
match_count: 4,
},
{
title: 'Nest + Cypress',
tag_names: ['nestjs', 'cypress'],
match_count: 2,
},
{
title: 'NX is cool',
tag_names: ['nx'],
match_count: 1,
},
];> EdgeDB does not treat Postgres as a simple standard SQL store. The opposite is true. To realize the full potential of the graph-relational model and EdgeQL efficiently, we must squeeze every last bit of functionality out of PostgreSQL's implementation of SQL and its schema.
I don't see how this and what you're saying can both be true at the same time. Is EdgeDB tightly coupled to the implementation of PostgreSQL, or isn't it? Is there really a chance that EdgeDB could support other databases, or not really? I don't think there's anything wrong with the answers being "yes" and "no", respectively; that's actually what I'd expect. It would be more unusual to try to do this in an implementation-agnostic way.
So to only get blog posts with matching tags we would need to add a filter „match_count > 0“, right?
Update: I am very excited about EdgeDB :)
with tag_names := {"angular", "nestjs", "cypress", "nx"},
select BlogPost {
title,
tag_names := .tags.name,
match_count := count((select .tags filter .name in tag_names))
}
filter .match_count > 0
order by .match_count desc;SQL definitely must be replaced, but the silly pseudo English syntax is one of the things we want to get rid of, not retain.
Found this: https://github.com/edgedb/edgedb-rust
With SQL, I have a mental model of how things work under the hood. For instance, I think of each table as being stored separately on disk, containing "rows". And the rows are really just equally-sized data blocks that are laid out back to back. B+ trees, with leaf nodes that point to (or just are) the rows, are used for indexes.
When I'm designing SQL schemas, I use this mental model to make guesses about performance. And when my queries are slow, I look at the execution plan.
My question is, how can I develop a similar intuition about EdgeDB? Under the hood, how are types and links stored in Postgres? And if I'm having performance issues, can I see an execution plan?
At the physical schema level [0] or at the conceptual schema level [1]?
This answer from edgedb CTO might clear the latter up; https://news.ycombinator.com/item?id=30291538
As for the former, I guess it is the same as however Postgres (pg) chooses to represent the edge-db tables. EdgeDB (graph on pg) sounds like Timescale (timeseries on pg [2]).
[0] https://en.wikipedia.org/wiki/Physical_schema
[1] https://en.wikipedia.org/wiki/Conceptual_schema
[2] https://blog.timescale.com/blog/timescaledb-vs-influxdb-for-...
* Every edgedb type has a postgres table
* "single" properties and links are stored as columns in that table (links as the uuid of the target)
* "multi" properties/links are stored as a link table
So it's basically just translated to a relational database in normal form
So... just an ORM ;)
Btw, if you folks have time, then EdgeDB should consider penning posts like the ones timescale has been doing for 3 years or so, in its march to industry leadership.
I think what I'm trying to understand is this: if I use EdgeDB in production, how often will I end up dropping down to the SQL level to debug things? If I'm trying to debug a slow query, can I do it at the EdgeDB level? Or will I have to open a PostgreSQL terminal, see how things are laid out there, run EXPLAINs, check the slow query log, and so on?
When I use ORMs, the answer to this is "pretty often". The ORM makes my application code cleaner, but I still need to have a complete understanding of the underlying SQL representation in order to ensure good performance and debug errors. I'm curious how that compares to using EdgeDB.
We're working on proper query profiling now, but in the meantime the "slow query" problem you see with ORMs happens quite rarely with EdgeQL. For starters a lot of ORM performance issues are causes by the fact that they secretly do a bunch of roundtrips under the hood. EdgeQL queries compile to a single SQL query always. Also, since we target Postgres exclusively, we can produce queries that take full advantage of its (rather preposterous) power and performance. We extensively test the performance of things like extremely deep/wide fetching, lots of nested & complex filter expressions, computed properties, subqueries, polymorphics, etc. Obviously nothing is 100% but we're pretty confident in saying that EdgeQL performance is good.
"What is a graph-relational database? EdgeDB is built on an extension of the relational data model that we call the graph-relational model. This model completely eliminates the object-relational impedance mismatch while retaining the solid basis of and performance of the classic relational model. This makes EdgeDB an ideal database for application development."
In a classic relational model everything is a tuple containing scalar values. Graph-relational extends the relational data model in three ways:
- every relation always has a global immutable key independent of data (explicit autoincrement keys aren't needed)
- this enables us to add a "reference type", which is essentially a pointer to some other record (i.e. a foreign key)
- attributes can be set-valued, so you can have nested collections in queries and in your data model.
This is what lets us do `Movie.actors.name` instead of a bunch of `JOINs`, because `actors` is declared as a set-valued reference type in the `Movie` relation.
FWIW, I do think there's a space in the market for a thin wrapper over Postgres (or MySQL) which would automate certain optimisations such as whether to index a particular table. It always struck me as perverse that that optimisation was delegated to the developer, when it's no more subjective or application-specific than a thousand other automated optimisations the engine makes. I'd be really interested if your project covered that.
There are countless permutations of the choices that database designers face, so it's a shame there aren't mature products for more of them. I hope this particular permutation turns out to be a good one for lots of people :)
And you - or hypothetically the end user - could change the backend, e.g. to Cockroach for better horizontal scalability, while trusting that EdgeDB will only rely on Postgres's public API at least in meeting its own public API/contract?
[0] It's hard to make that analogy with Postgres b/c it only has one storage engine, but of course the separation still exists.
From the docs:
Performs a recursive search on a collection, with options for restricting the search by recursion depth and query filter.
https://docs.mongodb.com/manual/reference/operator/aggregati...
It was considered bad practice and eventually got deprecated. Since 99% of the time the collection link in a given attribute is fixed and known in advance so it’s just duplicate information.
EdgeDB has:
- Full schema model with indexing, constraints, defaults, computed properties, stored procedures
- A query language that replaces SQL. If there's something you can do in SQL that isn't possible in EdgeQL, it's a bug.
- The query language is backed by a full type system, grammar, set of functions and operators, etc.
- A set of drivers for different languages that implement our binary protocol.
By any definition, EdgeDB is a database. It's a new abstraction built on a lower-level abstraction: Postgres's query engine. Both abstractions indubitably fit any reasonable definition of "database".
Basically: just because there's a declarative object-oriented schema doesn't mean this "is just an ORM" (unless your definition is quite pedantic).
If I build a DB Schema in EdgeDB, can I interact with the underlying Postgres instance using regular SQL?
Edgedb is more akin to hasura than to a database in traditional sense.
We realize that the database must be extensible and flexible, so non-EdgeQL UDF will become a reality (and if things work out the way we hope they will, they'll be amazing and far beyond what you can do with plpgsql).
Is a graph-relational database something completely disjointed from a graph database? Or do they share some performance improvements to some use cases? Also does EdgeDB keep the advantages of a true graph database even being based on Postgres?
This is unsurprising, because we just invented the term :-)
> Is a graph-relational database something completely disjointed from a graph database?
Graph-relational is still relational, i.e. it's a relational model with extensions that make modeling and querying graph-like data easier. And in apps everything is graph-like (hence GraphQL etc). An important point is that graph-relational, like relational is storage-agnostic, i.e. it makes no assumptions on how data is actually arranged on disk.
Pure graph databases, on the other hand, encode the assumption that data is actually _physically_ organized as a graph into their model and query languages.
I guess the word "graph" is simply too overloaded in computing.
Long before Arango, Orient etc.
The company is absolutely nowhere near to you guys in terms of marketing, but their thing works with more than 40 enterprise clients so far.
So you definitely did not ‘just’ invent the concept. A lot of companies approach the problem one way or another…
What would be the benefit / disadvantage in each case?
I have a couple of questions around this.
Firstly, what happens to the performance when I have a sizeable resultset of set-valued data?
I've seen similar ideas implemented in the past that look fine for the Movies and Actors or Books and Authors examples but fall apart badly when you query a number of fields (20+) that have sets within them, which can happen on say, a sizeable reference database of marketing information.
Another question: How deep in the graph can I go, and how much circular reference protection is there? E.g. if I query Movie.actors.movies.actors?
I'm interested in graph databases and data modeling and while it offers some convenience I'm always skeptical but hopeful (mostly from having lost a lot of hours) that these problems have been solved sufficiently to keep performance good in practical use cases.
> How deep in the graph can I go,
As much as you want, though the path must be explicit, EdgeQL currently doesn't have any way to say "traverse link foo recursively".
How deeply is EdgeDB integrated into Posgresql? Any chance it could be used to query other databases eventually?
Very interested in what data engineers think about this project!
I am not a developer, but the founders (including Yury Selivanov, Python Core Developer, see also https://github.com/MagicStack) and the fact that these people have been investing in the project for four years already, make me think that EdgeDB can be an important project for the database world!
Questions:
1. What is the story for replication currently? Can I use EdgeDB with Posgres replication tools like Stolon or Patroni, running EdgeDB against the proxy they expose?
Or does EdgeDB plan/need to have its own replication?
Googling this, I found this previous HN post (https://news.ycombinator.com/item?id=19640689) saying:
> Tooling for that will be coming in the next few alpha releases.
2. "A builtin migration system that can reason and diff schemas automatically or interactively"
How do you deal with the fact that Posgres does not offer transactional DDL (e.g. ALTER TABLE)?
In our Posgres, we had to use advisory locks around migrations to avoid concurrent schema changes invoked by concurrently starting servers which run migration.
movie_reviews
.filter(_.actor.name.lowercase() == "Zendaya")
.groupBy(_.title, _.credit_order, avg(_.ratings))
.sortBy(_.credit_order)
.take(5)
vs. select
Movie {
title,
rating := math::mean(.ratings.score)
actors: {
name
} order by @credits_order
limit 5,
}
filter
"Zendaya" in .actors.nameThat said, the functional variant is a likely way to represent EdgeQL in programming languages.
Also, it looks like your version also does implicit joins (like EdgeDB), but I'm not sure how they would work in that style.
[1] Or maybe only movies where Zendaya is in the first 5 credited actors; I'm not sure.
With EdgeQL, we conclusively solved a lot of these fundamental issues. Now, we're able to use EdgeQL as the foundation of our query builders, which is a huge advantage. We built the first version of the TypeScript QB in ~4 months) and it immediately leapfrogs all the major ORMs in power & expressiveness. But that's only possible because the hard work of designing EdgeQL was already done.
We'll be working on communicating how our query builder works and how it compares to ORMs in some of upcoming posts, stay tuned.
Live launch stream: https://www.youtube.com/watch?v=WRZ3o-NsU_4
How do you folks solve the large amount of joins that are the result of graph queries? Any worst-case-optimal multi-way-join secret sauce :D?
Also, with DBs like Datomic competing in the same area, do you have an immutability/versioning story?
We use array_agg, so we often avoid data serialization altogether. Our binary protocol just lets the data messages pass through with the original binary encoding. And because we fully control the schema, we can make all sorts of interesting optimizations, like implementing high-perf data codecs on the client side to unpack data fast.
Hasura has the frontend safe API and strong authz going for it. Is that something you might also do, or are you focused on serving the backend? End-user row and column level authz gives me a lot of peace of mind when writing bigger queries.
See also this reply by Colin: https://news.ycombinator.com/item?id=30293544
\q or \quit
If you get stuck, run \help for a list of commands. $ python
>>> while True:
... pass
...
^CTraceback (most recent call last):
File "<stdin>", line 1, in <module>
KeyboardInterruptCan someone use Postgres extensions with EdgeDB, like TimescaleDB, PostGIS, Zombodb and Postgres_fdw?
How are sum types (called enums in Rust) modelled in EdgeDB? (like Rust's Result). Do I need to define it with inheritance, where each variant inherits from it? What about adding specific syntax for sum types?
edit: also, I see there's a WIP custom #[derive] for Rust in https://github.com/edgedb/edgedb-rust - can it store Rust types on EdgeDB? Something like http://diesel.rs/ or even https://lib.rs/crates/turbosql
Sum types can be modeled with inheritance. You can create an abstract base type and use that as target for the links. You can then derive types from it and write polymorphic EdgeQL queries to select/match data.
Rust client is still work on progress, not really open to tinkering unless you want to experiment.
Do you offer fully managed service to host this or do I need to spin up some compute instances of my own ?
Can't wait for edgeDB cloud!
Technically speaking though, using GraphQL as a data model is quite limiting: just basic scalar types and a handful of custom types implemented by DGraph (Polygon, etc). By contrast EdgeDB schemas support a much wider range of granular types, constraints, computed properties, link properties, etc.
Perhaps more importantly EdgeQL is a full query language with composable syntax, a standard library of functions and operators, for loops, primitive literal manipulations (e.g. string indexing and slicing), the ability to cast values to different types, subqueries, etc. Basically it's a complete query language contrast GraphQL is closer to an ORM in that it really only supports CRUD.
I think DQL is marginally more powerful though admittedly I'm not very familiar with it: https://dgraph.io/docs/get-started/
This seems great. I have a highly interrelated data model (think parsed natural language text with annotations), and writing SQL to keep all of that data aligned and synced is a pain with an ORM.
Am I right in imagining this works similarly to how Apple's CoreData does? It lets you build objects linked to each other but handles all of the joining and syncing of the data for you to keep the object model in mind.
In the same vein will you be creating a Swift client?
We position EdgeDB as a database server, not a library, partly because you can interact with it from different programming languages. But we design our client library with focus on API composability, check out our edgedb-js library for example: https://www.edgedb.com/docs/clients/01_js/index
As for the Swift, it would be great to have it one day. I have a counter question: how many of you use Swift to write server-side logic?
And EdgeDB is very needed because it's just an Apple library currently.
Conceptually, my impression is that EdgeDb is relational tables (highly-typed) queried, combined, and modeled as nodes in a graph/tree. Is that conceptually correct?
I don't, but I would like to if I could deploy Swift code to Azure. I love Swift as a language.
I've dabbled with CoreData myself and it's a phenomenal API. Apple comes up with a lot of great stuff. In fact, given how much server-side Swift there is these days, we'll have to look into building a Swift client library at some point.
I hope EdgeDB is immediately recognized for the value it will bring to developer productivity everywhere!
Now we just need an SQLite version of EdgeQL (perhaps EQLite) for local data and it can completely replace SQL in my life.
Why no C API? Pretty much every language has a C FFI, PostgreSQL's libpq is in C too, and then it's easy to make bindings from other languages. But with a few individual implementations in other languages it's not as easy to use from more languages. Always makes me to wonder about these decisions when I see projects heading that way.
Unrelated to the subject, but the text on the website is #B3B3B3 on #FFF, which is hard to read as is, and violates WCAG recommendations. When overwriting it with a custom CSS, it's visible that headings have weird margins, covering parts of text before them.
While looking for more, I noticed that the website suggests to `curl | sh` things in the introduction, which is quite awkward too (and a subject of light flamewars), adding another barrier.
The brief project description made me to wonder how well it abstracts out PostgreSQL (and how to work with the databases it creates via PostgreSQL itself, if it's possible at all, how to debug it when things will go wrong), but after brief skimming there's no PostgreSQL bits in sight: just a special shell, a dedicated language, its own drivers. Which is a bit scary.
Neither have I found a description of its graph-relational model, how it's built on top of PostgreSQL, how one can be sure that it'll work more or less smoothly, how things like profiling are done (or is it designed to never need explicit profiling/optimization?), is there more to it than PostgREST-like interface abstracting out the SQL bits for DDL too.
Looks like an interesting project overall though.
Sorry, but dissing proven technology like this just makes me vomit. Show me _how_ you can beat SQL in performance and features on the frontpage, or I'm just happy to go along with what I already have.
Scroll down. They do exactly that.
Here are are some links:
Pointed critique at SQL: [1]
Benchmarks: [2] and [3]
We'll be adding a dedicated benchmarks page to our website soon.
[1] https://www.edgedb.com/blog/we-can-do-better-than-sql
That classic ORMs are kinda slow is probably not a surprise to anyone, the others which either get to compile the full query or have a hand in controlling the schema are more interesting.
They also seem more similar to your product, running as a server and managing the schema, so most worth comparing.
Edit: The fact that "Raw SQL" ends up being a suboptimal query because of the node driver limitations, which then gives you "Way faster than even raw SQL!!!" graphs also leaves a weird taste. I guess if you are comparing programming language level solutions fair enough.
"Faster than SQL" is, of course, relative and depends on "what SQL"? EdgeQL compiles into a single query that uses PostgreSQL-specific features. This is a guarantee. No matter how large or complex your query is, if it compiles, it compiles into a single SQL query. Manually written or ORM-generated SQL tends to be "multi-query" due to the whole "standard SQL composes badly" story. And this matters, because if the roundtrip network latency between the client and the server is 10ms, EdgeQL will get you a response in ~10ms, whereas a multi-query approach will in (~10ms X <number-of-queries>) even if every individual query is super-quick to compute.
But just for clarity in the discussion thread here, Hasura also compiles to a single query when only Postgres is being hit and I'd expect performance to be quite similar...
Ofcourse, if the GraphQL query requires federation across multiple Postgres databases or multiple databases or databases + other GraphQL / REST APIs, then Hasura breaks them up into multiple queries with a worst-case performance of a data-loader type set up.
> Scroll down. They do exactly that.
Where?
How's that dissing anything? They are saying they have a good abstraction that will make you more productive. While it might not be, this is literally what makes software the amazing tool it is, building layers of abstraction that make you more productive.
You seem to have gotten it all wrong, how could building something on top of SQL be dissing SQL? It's as much of a diss as C is a diss to assembly and so on.
> Show me _how_ you can beat SQL in performance
It is literally a postgres instance underneath, if edgedb is malleable enough, you don't need to beat anything, you'll have the query you want (sans edge cases and some minor overheads).
> I'm just happy to go along with what I already have
Go ahead bud, I'm pretty sure it won't be legally enforced to use any time soon so you're good to go.
The fact that this low effort comments get upvoted in HN drives me mad, just the contrarian attitude will get you points no matter the depth or effort. You don't seem to have even read the landing page or tried it, maybe you did but your comment doesn't reflect that, this is the product of 4 years of effort by some devs that are trying to innovate, dissing it without even fully seem to understanding makes me want to vomit.
It's an amazingly lazy and inconsiderate way to respond to people who've put a bunch of effort into building something.
> It is literally a postgres instance underneath,
Exactly. So. Why? Someone created a database abstraction? God damn, that has almost never happened before! ;)
> We shall do better than SQL The EdgeQL language looks cool, and I'm sure querying via a graph structure makes certain problems easier in some use cases. However as much as people have complained about SQL, it's just so ubiquitous there needs to be a very good reason to switch away from it. Not having to write joins isn't really a good enough reason, in my opinion.
> The true source of truth I'm not sure why this means EdgeDB is better. Tons of applications use a traditional or cloud SQL database as the source of truth right now. This section seems to imply with microservices you no longer have a single source of truth. But if they're trying to say a microservice system should instead us a single common database that breaks separation of concerns and moves us into an annoying situation where you have a bunch of services communicating via a shared database.
> Not just a database server It sounds like they have a solid client, which is awesome.
> Cloud-ready database APIs > The vast scale of modern application deployments requires that inelastic computing resources are managed very carefully. Until cloud-native databases reach complete functional and performance parity with traditional databases, we will have to contend with the fact that the database is a scarce resource.
This used to be true, but is definitely no longer true. Cloud-native databases are everywhere and incredibly common. See any major cloud, https://www.cockroachlabs.com/, or any of the tons of other database solutions.
It's great to see a new database coming out - innovation in the space is super important. However this announcement reads like marketing speak, and is light on the details. When I see a new product I want to hear things like: - about how it scales - what the architecture is - why is it stable and trust-worth enough to put my data on - is it multi-node? How did they make it serializable? - how fast is it? Performance is super important.
Based on their website it seems like a thin skin over postgresql. If that's the case I'll just use postgresql. If it's a clustered new and advanced database, then I'll be wary about trusting it for anything real.
> The EdgeQL language looks cool, and I'm sure querying via a graph structure makes certain problems easier in some use cases. However as much as people have complained about SQL, it's just so ubiquitous there needs to be a very good reason to switch away from it. Not having to write joins isn't really a good enough reason, in my opinion.
Oh, it goes much deeper than not writing joins. There's no single ORM out there that can implement a TypeScript query builder like ours, see the example in [1]. This is only possible because of EdgeQL composability, but that composability required us to rethink the entire relational foundation.
> > The true source of truth
> I'm not sure why this means EdgeDB is better. <..>
This section implies that EdgeDB's schema allows to specify a lot of meta / dynamically computed information in it. And soon your access control policies. Take a look at the work-in-progress RFC [2] [3] to see how this is more powerful, then say, Postgres' row level security.
> > Not just a database server
> It sounds like they have a solid client, which is awesome.
Also lightweight connections to the DB so that you can have thousands of concurrent ones without load balancers, built-in schema migrations engine, and many other things. In fact we have so much that it's challenging what to even highlight in a blog post like the 1.0 announcement.
> Cloud-ready database APIs
> This used to be true, but is definitely no longer true. Cloud-native databases are everywhere and incredibly common. See any major cloud, https://www.cockroachlabs.com/, or any of the tons of other database solutions.
Not to pick on CockroachDB (they have an amazing product and company, we love them), but you should benchmark local install of Postgres and Cockroach to see yourself that scalability still has a significant cost in performance.
[1] https://www.edgedb.com/blog/edgedb-1-0#not-just-a-database-s...
I don't understand this claim. Can or does? This all just compiles down to SQL right?
To me the problems with existing DBMS are still the same (as 15 years ago) complexities in setup/clustering/backups/rollbacks/schema updates, even setting up db clients are PITAs in many environments.
SQL is not really a feature to focus on (imho), it is simple enough even for non-tech people). We tried to get out of SQL long time ago anyways (ORM).
Anyways wishing luck, the team seems awesome!
I think EdgeQL (and the data model) has potential for being a lot more than accessing the database. For example, you could have a distributed system use this data model across various typically replicated data stores - include all caches and "materialized views" of various kinds. Extend EdgeQL to define various caches and computed values from the core schema. Then you could use it to represent querying the same data but with different freshness. Now this becomes a compelling idea. Essentially it gives you a well defined way to declare and manage various computed values, laggy caches without having to manually implement all that coordination.
1. define a view representing the object fields you want to cache 2. define eviction and properties of the cache 3. use queries to query either the cache or the db
Specifically, do all of the above using edgeql and not have to write the cache consistency or serialization/deserialization logic.
SELECT
title,
ARRAY_SLICE(ARRAY_AGG(movie_actors.name WITHIN GROUP (order by movie_actors.credits_order asc)),0,5)
avg(movie_reviews.score)
FROM movie
JOIN movie_actors on (movie.id = movie_actors.movie_id)
JOIN person on (movie_Actors.person_id = person.id)
JOIN movie_reviews on (movie.id = movie_reviews.id)
WHERE person.name like '%Zendaya%'
group by titleI immediately assumed it was some kind of distributed db running at the edge but it seems this is not the case.
When do you drop down to SQL / where does the boundary end for edgeQL?
(And, is it likely to be possible to build it against standard interfaces like JDBC - or is it too different?)
> InvalidReferenceError: object type 'default::User' has no link or property 'Name' > Hint: did you mean 'name'?
Which is the expected response (same as the edgedb cli)
1. How does EdgeDB compare to Supabase? Thinking both of realtime functionality and row-level security. 2. If I was to use EdgeDB instead of Django, how would I go about it? In other words, how can I set up a batteries-included, server-side web app?
[1] https://github.com/edgedb/rfcs/blob/865bc48f4050ced99447bd77...
On a tangent note, I find it is honestly annoying to have to maintain a separate account/authorization/vpc connection just because I want to use a database. I would much rather startups work with cloud providers to offer it natively or we invent a better way to interoperate with clouds. The explosion of small "cloud" that offer one service each isn't pleasant to work it especially when they don't have a terraform provider.
Re second point you can deploy edgedb to your cloud of choice. As for working with cloud providers to make integration better for the user, it's a bit too early for us to comment.
Cloud of choice is great, also need region and AZ of choice too for any serious production database. What is killing performance is round trip time to the DB. The language seems to help on the number of round-trips so that's good.
Yep, you got it! :)
1. If other clouds want to offer a hosted EdgeDB offering, they can do so. We'll have a tight integration with the `edgedb` CLI and we're confident we can beat them on the developer experience which is our #1 priority.
2. We'll likely use GitHub for auth though this isn't set in stone yet. We're still planning out the workflows surrounding EdgeDB Cloud.
3. As for the explosion of cloud silos—that's a very real phenomenon with a boring reason: hosting is one of the few ways to make money building OSS software. There are preposterously valuable OSS frameworks + libraries that haven't made anyone a cent. We implemented "one-click" deploy buttons for the major clouds to the extent it was possible to do, but for fully open-source companies like EdgeDB there's aren't many routes to sustainability outside of paid hosting.
Yeah thats ok for startups I guess.
Yeah sad state of affair really, though fully solvable by tech IMO if we had some protocol for integrating services in clouds. In any case, please don't underestimate the terraform integration. I really could not care less about the one-click deploy. What I care about is maintenance and integration with my existing infrastructure as code.
https://www.edgedb.com/blog/building-a-production-database-i...
Thank you for sharing.
Congrats!
One use case I have (solved by some current graph DBs) is performing a complex full text search (boolean operators, stopword removal, fuzzy matching etc.), which alters search ranking based on graph properties (e.g. a node with more edges might get a reduction in score, the root of a tree might get a boost, etc.).
I'm still going through the tutorial, so haven't got to grips with what the DB is capable of yet.
We have a software that's a GraphQL interface to a database that's populated with a project (let's call it the indexer project) we do not control. It would be great if we could check the database schema for problems it might have to be used as an EdgeDB database. Then we would migrate our application to use EdgeDB, while the indexer keeps loading our database through direct interaction with PostgreSql.
I wish GraphQL were more like this and considered built-in features like where clauses and cursors, instead of having to add those over the top with loose conventions.
I don't see how can they retain the performance of the relational model when using graphs on top of PG unless they are using PG as storage k/v layer only.
But EdgeDB isn't ideal for storing loosely typed graph data, neo4j is built for that.
Can use these techniques on top of typical, existing relation model tables as well. Consider a join table, or even a polymorphic join table, that gives you the edges. But you can use unions and stuff to join up data across various tables. Can get really creative with it.
This is very true and might happen in a far future.
a graph database is not useful because of its query language, it's useful because of its performance characteristics when bringing linked data — it could never be performant on a Postgres backend.
Eg. vs Cypher or the likely-Cypher-compatible forthcoming GQL standard? https://www.gqlstandards.org/
Standardization projects typically take awhile, specifically for something complicated as a query language spec.
1. SQL/PGQ (ISO/IEC JTC1 9075 part 16) -- This adds language to create property graph views on top of existing SQL tables and write property graph queries in a GRAPH_TABLE function in an SQL FROM statement.
2. GQL (ISO/IEC JTC1 39075 Database Language GQL) -- This is a full declarative property graph database language to create, maintain, and query graphs. This includes support for both descriptive and prescriptive schemas.
The Graph Pattern Matching language is identical between the two standards. For more details about GPM, see https://arxiv.org/abs/2112.06217
The ISO process has a defines series of milestones. I will spare you the details at the moment.
SQL/PGQ will start a Draft International Standard (DIS) ballot in July 2022 and so will be a published standard next year - 2023.
GQL will finish a Committee Draft (CD) ballot this month (February 2022) and should be ready for a DIS ballot in early 2023. However because GPM is shared between SQL/PGQ and GQL, the GQL Graph Pattern Matching will be stable when SQL/PGQ goes to DIS ballot.
For a little more detail on the GQL standards process and content, take a look at the talks from the LDBC TUC meeting August 2021: https://ldbcouncil.org/event/fourteenth-tuc-meeting/
One quick question - I see that you have a Rust client, however it's marked WIP, how usable is it in its current state? Any idea on a timeline of it becoming an official binding?
Very interested in this! I've been wanting to build a generic HATEOAS server and this feels like this might be the right database!
EdgeDB is designed to do its job validating and compiling your schema and queries and then get out of the way. In other words, once a query was first parsed and compiled, the cost of the next trip via EdgeDB would be similar to that of pgbouncer, i.e. we'll simply send the compiled SQL to Postgres and proxy the results back to the client. This is why our data protocol uses Postgres framing and encoding.
impoart * as edgedb from "edgedb";
I started a container and connected it, but instantly got an email from heroku telling me i had 19500 of my 10000 rows, and 215 tables. Not really any hobby-project viable priced hosting options in their documentation.
Now I'm not against OSS movement and in fact I'm working at a company doing exactly this, but releasing your bread-and-butter as an open source project leaves you vulnerable to peer plagiarism.
Instead, you should just release the core version so others will build an open ecosystem around your mainline product. This will secure your market fundamentals by making sure no one could overshadow you especially in a fiercely competitve market of database.
So please stop open sourcing too much, I don't want to see the same ill fate again and again for great products like RethinkDB (and its downfall) that could change the world
To be clear EdgeDB is the core version, albeit a large core. We have many ideas about value-adds. But the goal for open sourcing it is indeed to foster an open ecosystem.
I have a question regarding the migrations, in the website you say it's safe to run them with automated flow, we know it uses Postgres under the hood and we know that sometimes migrations on large datasets can cause downtime in the database clusters. How do you handle them? I guess since you don't have a backing "users" table for a "User" model adding / removing fields is not actually happening the same way it happens in a normal relational database, thus it's not a blocking or resource consuming operation for you?
How it is a new database? Or an advanced orm?
- Full schema model with indexing, constraints, defaults, computed properties, stored procedures
- A query language that replaces SQL. If there's something you can do in SQL that isn't possible in EdgeQL, it's a bug. Most ORMS provide some sort of language-specific API for writing queries and generating SQL under the hood—that's dramatically different than providing a new query language.
- A full type system, grammar, set of functions and operators, etc. A set theoretic basis for all expressions in EdgeQL. https://www.edgedb.com/docs/edgeql/sets
- A set of drivers for different languages that implement our binary protocol.
EdgeDB is a new abstraction built on a lower-level abstraction: Postgres. Both indubitably fit any reasonable definition of "database".
ORMs work fine for relational data, until there's a lot of edge data on the joins. I looked at the docs on mobile and couldn't find an answer, how does EdgeDB handle data on joins? E.g. GraphQL "connection" types with edges.
(Tao uses mysql and stores graph data as pretty much key/value pairs, and then a layer on top to query it)
> would I have to care about the Postgres schema, or is that abstracted away from me?
EdgeDB takes care of everything for you. You wouldn't know it's Postgres underneath unless we told you.
I haven't worked at Facebook so I'm not super familiar with TAO, only heard about it from my friends. But sometimes I pitch EdgeDB as a database with which you don't need to build your own TAO at your company. we give you much more capabilities than a typical database.
Q: What are your plans for sharding / scale-out?
At Notion, we run ~480 logical schemas spread across ~32 Postgres databases [1]. Our data model [2] (sorry for the blog post spam) has a recursive / graph-like structures, and we make use of columns or jsonb attributes like `{ table: Table, id: UUID }` which sounds like your polymorphic links feature. It seems like EdgeDB's model lines up well with how we already use our database.
I did a quick cmd-f here, on the linked announcement, and in your docs looking for "shard" and "scale" but didn't find any relevant results. Postgres needs a Vitess!
Q: Do you have plans to support EdgeQL embedding or SQLite?
I am always looking for a way to compile "better than SQL" languages down to SQL. I like Datalog in this area because it's a composable way to define relationships/facts/derivations but no one is putting serious business effort to this idea (honorable mention to logica [3]). EdgeQL also fits the bill -- queries aren't logical, but they are composable -- plus looks easier to teach than most Datalog variants.
I took a peek in the repo and saw that a few EdgeQL components are written in Rust [4]; are you considering porting more logic to Rust? That would make EdgeQL much more embeddable - it could run in WASM or linked into an Android/iOS app. My pie in the sky dream is to use a single composable query/logic language to define all my relations and queries, and then compile that stuff so it works the same on both the client, server DB, and data streams (for incremental materialized views, ideally on both client & server).
If EdgeDB had a Lite version that ran on SQLite, we'd be 66% of the way there.
(To get the materialized view bits on the server, a mad scientist might already be able to point EdgeDB at Materialize [5])
[1]: https://www.notion.so/blog/sharding-postgres-at-notion
[2]: https://www.notion.so/blog/data-model-behind-notion
[3]: https://opensource.googleblog.com/2021/04/logica-organizing-...
[4]: eg https://github.com/edgedb/edgedb/tree/master/edb/edgeql-pars...
Thanks!
> Q: What are your plans for sharding / scale-out?
Sharding is planned, though there is no set design yet, this area is in early research phase currently. Thanks for sharing your experience by the way! Learnings from the field definitely help. A traditional read replica scale-out is already supported and we are building integrations with Postgres orchestrators (for failover, replica discovery etc). Oh, and automatically routing read-only queries to read replicas (with some controls for lag) is something that we plan as well.
> Q: Do you have plans to support EdgeQL embedding or SQLite?
Possibly. Depends on the application and performance expectations :-) PostgreSQL is really special in its ability to deal with complex queries. We already have a toy EdgeQL interpreter in the codebase [1], which is mostly used to quickly prototype syntax and validate semantics. It would be great to scale it up to something that can work with persistent stores (even if dumb and slow).
> are you considering porting more logic to Rust?
Yes, that the long term plan.
[1] https://github.com/edgedb/edgedb/blob/master/edb/tools/toy_e...
> And tracing what is happening from beginning of query to result?
EdgeDB exposes a Prometheus endpoint with a bunch of metrics. We don't have a public API for tracing individual queries yet.
Edit: Wow, after starting to read the book you guys have done a fantastic job of making it a joy to learn. Learning a database technology by following dracula? amazing.
PLEASE PLEASE PLEASE giving this first class support in rust. The select statments reminds me of structs and you even have enums.
select Movie { title, rating := math::mean(.ratings.score) actors: { name } order by @credits_order limit 5, } filter "Zendaya" in .actors.name
Do I need to write some metadata layer so this higher level query would know what's going on? Similar to LookerMl?
I'm particularly excited about the HTTP and the GraphQL interface. This could potentially mean no need for a backend. One thing I am curious about it the authorization part of it - how can I limit the result set, similar to Row Level Security in Postgres?
You might need to worry about this in two senses:
1. Graph constructs aren't first-class citizens in the backend. Various potential performance improvements will necessarily be missed.
2. PostgreSQL is ok/good for transactional work, not good for analytical work, like columnar DBMSes.
Looking forward to giving it another try soon!
To quickly learn EdgeQL I recommend our online interactive in-browser tutorial: [1]
We also have a book, it's called Easy EdgeDB, check it out here: [2]
[1] https://www.edgedb.com/blog/building-a-production-database-i...
And edit: why is there no Java/Kotlin client?
If we're going to replace SQL the silly 1960s pseudo-English syntax is one of the things we want to get rid of, not retain.
Are you going to support some kind of subscriptions for realtime?
We perform quite favorably in benchmarks, see our old blog post with some: https://www.edgedb.com/blog/edgedb-1-0-alpha-2#results
> Is 1.0 your MVP?
EdgeDB is ready for production and is light years ahead of its first technical preview MVP release published a few years ago.
should be retrieve
Edit: oh, a relevant reply https://news.ycombinator.com/item?id=30293064
OrientDB positions itself as a multi-model NoSQL database. Where's EdgeDB positions itself as a relational database and a successor of SQL.
OrientDB?
How you guys compare yourself to dgraph database?
If you use a graph database to run graph algorithms on your data, then NO. Although we'll be working on adding support for recursive queries to EdgeQL in the near future.
Is there a roadmap I can subscribe to somewhere?
Looking forward to trying this out!
Would a graphql API be part of your roadmap?