HNHacker News
TopNewBestAskShowJobs

RedCrowbar

438 karma · joined September 20, 2012

Working on hard problems @ Vercel.

Previously co-founder and CTO of Gel Data [1].

[1] https://geldata.com

[ my public key: https://keybase.io/elprans; my proof: https://keybase.io/elprans/sigs/p0bjVrvsPB2J-0pE_ZgpeT-gli6jsZ45wHgeBaMKNNY ]

submissionscomments
RedCrowbar··on EdgeDB 2.0
Demangling the schema into readable SQL schema is quite trivial already, as it's a fairly normal looking schema, except we use schema object ids in place of names (to simplify renaming). Turning EdgeQL into "readable" SQL is probably also possible, depending on your definition of "readable" and the complexity of a query, though the current codegen does not prioritize SQL readability at all.
RedCrowbar··on EdgeDB 2.0
Yes, you can easily put OQL in the list of prior art EdgeQL was inspired by.
RedCrowbar··on Show HN: Prisma Python – A fully typed ORM for Python
That's correct. Prisma is an ORM with all the pros and cons of being one. EdgeDB, on the other hand, is a brand new graph-relational database server, built on PostgreSQL.

ORMs have the ability to work with multiple RDBMS implementations, but for that they trade away expressive power and efficiency. Prisma is fairly slow, especially when your query is fetching multiple relationships, because it does multiple DB roundtrips to fetch parts of the result and reconstitutes it on the client. ORM APIs are generally very limiting, as there is no general way of doing server-side computation (e.g. do a comparison on a substring of a string property or simply do arithmetic).

EdgeDB does not have these problems, because its query language, EdgeQL, is designed to be efficiently embeddable into a programming language without any loss of expressive power or performance. At this moment we have a fully-featured TypeScript/JavaScript builder [1], with Python and Go coming soon.

[1] https://www.edgedb.com/docs/clients/01_js/index#the-query-bu...

(Full disclosure: I work on EdgeDB)

RedCrowbar··on Show HN: EdgeDB 1.0
Yes, the plan is to allow sending read-only queries to read replicas automatically. This isn’t implemented yet, but all the requisite pieces are there.
RedCrowbar··on Show HN: EdgeDB 1.0
There is no way to call UDF SQL (or PL/pgSQL) functions from EdgeDB, because there is no way to _define_ or manage them. The only place that is allowed to do this is the standard library (and, in the near future, extensions).

We realize that the database must be extensible and flexible, so non-EdgeQL UDF will become a reality (and if things work out the way we hope they will, they'll be amazing and far beyond what you can do with plpgsql).

RedCrowbar··on Show HN: EdgeDB 1.0
That’s exactly right.
RedCrowbar··on Show HN: EdgeDB 1.0
We compile EdgeQL queries into SQL currently, because it makes the architecture simpler and less us run on unmodifed Postgres, but conceptually nothing stops us from targeting the query planner directly via an extension or an alternative frontend that consumes EdgeQL IR direclty.
RedCrowbar··on Show HN: EdgeDB 1.0
> Congrats on the milestone!

Thanks!

> Q: What are your plans for sharding / scale-out?

Sharding is planned, though there is no set design yet, this area is in early research phase currently. Thanks for sharing your experience by the way! Learnings from the field definitely help. A traditional read replica scale-out is already supported and we are building integrations with Postgres orchestrators (for failover, replica discovery etc). Oh, and automatically routing read-only queries to read replicas (with some controls for lag) is something that we plan as well.

> Q: Do you have plans to support EdgeQL embedding or SQLite?

Possibly. Depends on the application and performance expectations :-) PostgreSQL is really special in its ability to deal with complex queries. We already have a toy EdgeQL interpreter in the codebase [1], which is mostly used to quickly prototype syntax and validate semantics. It would be great to scale it up to something that can work with persistent stores (even if dumb and slow).

> are you considering porting more logic to Rust?

Yes, that the long term plan.

[1] https://github.com/edgedb/edgedb/blob/master/edb/tools/toy_e...

RedCrowbar··on Show HN: EdgeDB 1.0
It's usable and functional, because we use it in our CLI. It's WIP, because we haven't yet committed to an API, especially in async. Rust is a bit hard in that department :-)
RedCrowbar··on Show HN: EdgeDB 1.0
It's something we want to do at some point, but unlike Hasura, which operates on GraphQL which is conceptually much simpler and limited, EdgeQL would be much harder to fit onto a "pipelined polling" model that Hasura utilizes to implement subs.
RedCrowbar··on Show HN: EdgeDB 1.0
EdgeDB is graph-relational, not a pure graph database, and so the performance characteristics of traversing links are that of a relational JOIN. Which, of course, depends wholly on the size of each relation being joined. So, if you want to select the list of actors for _every_ movie in your database and there are lots of movies, it'd be a pretty expensive operation. If, on the other hand, you want to select some relationships on a handful of objects (or even just one), then it doesn't really matter that much how deep your link traversal is, because all of the steps would be fast index scans.

> How deep in the graph can I go,

As much as you want, though the path must be explicit, EdgeQL currently doesn't have any way to say "traverse link foo recursively".

RedCrowbar··on Show HN: EdgeDB 1.0
The retry logic in clients is fully configurable, you can disable retries and get your TransactionSerializationError if you want that.
RedCrowbar··on Show HN: EdgeDB 1.0
This is not something we plan to do in the near future, but it’s also not outside the realm of possibility. We picked Postgres because of its power, quality and unparalleled extensibility, but we are also very careful to not leak any implementation details into our interfaces.
RedCrowbar··on Show HN: EdgeDB 1.0
> I’ll add to the positivity

Thank you!

> Serializable transactions are expensive, and that deserves to be an explicit caveat. Not everyone knows this, and it’s an important thing to put up front.

We've not seen a major difference in our benchmarks (though maybe our benchmarks are wrong :-)). EdgeDB tends to produce very short transactions, so that helps. EdgeDB also knows if your statements are read-only or not, so we have the ability to steer these into a read-only transaction, though this isn't implemented yet.

RedCrowbar··on Show HN: EdgeDB 1.0
The problem sounds like something that could be solved with a GIST index. EdgeDB doesn't yet have a way to specify the index type, though, mostly because we aren't sure what would be the best way to do it without things becoming too Postgres-specific in schemas.
RedCrowbar··on Show HN: EdgeDB 1.0
We need to figure out the formal API and packaging format for extensions. We're working on it.
RedCrowbar··on Show HN: EdgeDB 1.0
It's a long story :-)

https://www.edgedb.com/blog/building-a-production-database-i...

RedCrowbar··on Show HN: EdgeDB 1.0
Security policy will be part of the next release. See draft RFC [1], although note it's likely not going to be the final syntax.

[1] https://github.com/edgedb/rfcs/blob/865bc48f4050ced99447bd77...

RedCrowbar··on Show HN: EdgeDB 1.0
We've got some benchmarks in an earlier blog post [1].

EdgeDB is designed to do its job validating and compiling your schema and queries and then get out of the way. In other words, once a query was first parsed and compiled, the cost of the next trip via EdgeDB would be similar to that of pgbouncer, i.e. we'll simply send the compiled SQL to Postgres and proxy the results back to the client. This is why our data protocol uses Postgres framing and encoding.

[1] https://www.edgedb.com/blog/edgedb-1-0-alpha-1

RedCrowbar··on Show HN: EdgeDB 1.0
We'll post some benchmarks against Hasura et al soon.

"Faster than SQL" is, of course, relative and depends on "what SQL"? EdgeQL compiles into a single query that uses PostgreSQL-specific features. This is a guarantee. No matter how large or complex your query is, if it compiles, it compiles into a single SQL query. Manually written or ORM-generated SQL tends to be "multi-query" due to the whole "standard SQL composes badly" story. And this matters, because if the roundtrip network latency between the client and the server is 10ms, EdgeQL will get you a response in ~10ms, whereas a multi-query approach will in (~10ms X <number-of-queries>) even if every individual query is super-quick to compute.

RedCrowbar··on Show HN: EdgeDB 1.0
We will run and support EdgeDB for you. Here's an expanded answer: https://github.com/edgedb/edgedb/discussions/3377
RedCrowbar··on Show HN: EdgeDB 1.0
Unfortunately, JSON aggregation destroys type information, so you can't reason about that your SQL query actually returns anymore.
RedCrowbar··on Show HN: EdgeDB 1.0
Also, this query is wrong because we want movies where Zendaya played, but _also_ other actors in order, possibly without Zendaya, so you really need to do the actors join twice.
RedCrowbar··on Show HN: EdgeDB 1.0
EdgeDB currently is. Graph-relational and EdgeQL are not.
RedCrowbar··on Show HN: EdgeDB 1.0
EdgeQL is designed to replace SQL, not graph query languages. Think of it as SQL getting a proper type system and GraphQL capabilities of reaching into deep relationships in an ergonomic way.
RedCrowbar··on Show HN: EdgeDB 1.0
It runs as a separate (stateless) process between the client and the PostgreSQL server. There was a talk about the details of the architecture on the live stream today: https://youtu.be/WRZ3o-NsU_4?t=5294
RedCrowbar··on Show HN: EdgeDB 1.0
There are similarities with TAO and it's not a coincidence. Facebook engineers recognized that a better data abstraction and API was needed for productivity. EdgeDB follows the same logic.

> would I have to care about the Postgres schema, or is that abstracted away from me?

EdgeDB takes care of everything for you. You wouldn't know it's Postgres underneath unless we told you.

RedCrowbar··on Show HN: EdgeDB 1.0
> It's the first time I hear about graph-relational DBs.

This is unsurprising, because we just invented the term :-)

> Is a graph-relational database something completely disjointed from a graph database?

Graph-relational is still relational, i.e. it's a relational model with extensions that make modeling and querying graph-like data easier. And in apps everything is graph-like (hence GraphQL etc). An important point is that graph-relational, like relational is storage-agnostic, i.e. it makes no assumptions on how data is actually arranged on disk.

Pure graph databases, on the other hand, encode the assumption that data is actually _physically_ organized as a graph into their model and query languages.

I guess the word "graph" is simply too overloaded in computing.

RedCrowbar··on Show HN: EdgeDB 1.0
Thank you!
RedCrowbar··on Show HN: EdgeDB 1.0
Because that only gives you actor names, not records, and also because arrays aren't a universal SQL feature.
← PreviousPage 2 of 6Next →