This scaling could work a variety of ways. Records might be horizontally sharded across a bunch of database instances. Caching is layered above the database to protect against hot database queries, etc. The generation logic now needs to understand how to partition a single GraphQL query across the shards and caching layers.
It's definitely not impossible, and in some circumstances it may be easy. But it couples your GraphQL implementation very tightly to your current database system in a way that may make it difficult to rearchitect your database system at the next order of magnitude scaling problem - "We really need to redo how we think sharding and caching, and our GraphQL's direct interaction with the data layer is going to make this a challenge." Maybe it's easier if the company is new and is GraphQL-only. An older company that has a mix of a few data access technologies might wish the GraphQL left the data layer concerns to the data layer.
Have to admit I haven't had this issue yet and I think it's a nice problem to have at that time as means too many paying customers :)
I know it's counter intuitive, but we think of this as an opportunity to solve those problems at the Hasura layer. Sharding and caching are hard but also not impossible to make declarative if you know enough about how the data is modelled, how data ownership works (authorization) and how the cache is managed (caching hot queries, caching complex queries but infrequent, invalidated on time or data).
I'm not trying to say the problem isn't hard, I'm trying to ask what it would take to solve the problem. :) We're also actively working on solving them. Imagine if those challenges got solved for a sufficient number of use cases!
Incidentally, databases are getting better at preventing users outgrowing them too while preserving the SQL API as much as possible, or providing a transform. Citus / YugaByte / Planetscale / Cockroach / Spanner to name a few.
Have you guys thought about looking at some other tools that might be symbiotic? Hasura's interface would be an incredible companion to something like Dremio (also open source), since it is a fully-SQL engine across a number of heterogeneous systems, with probably the most advanced caching I have ever seen in a system like it (Apache Arrow in-memory and columnar storage). If you can make your query interface adapted to tools like Dremio, you shouldn't have any problem exploiting caching in those kinds of platforms - might be even more efficient than building into Hasura for some use cases.
Edit: I think there is a newer type of interface architecture forming here, and I think GraphQL represents one of the pieces of the puzzle, would be really interested to see more detail on Hasura's take on that.
You can do this in PostgreSQL, it supports sharding natively through its FDW ("foreign data wrapper") feature.
Sometimes you are creating an application and building its underlying datastore at the same time. Hasura is great for this.
But sometimes you already have a disparate ecosystem of services and datastores and you want to provide a GraphQL API for that. This use case is better-served by other approaches.
Great tool. Personally I don't trust GraphQL completely.
I had to come up with a deliberately contrived UI to find a realistic query that was more than 4 relationships deep. Most were around 3 (though double-nested connections were fairly common within this). The more common problem was around overall data size, where the bottleneck was JSON serialisation overhead and network transfer size.
Query whitelisting is the winning approach for 1st party APIs. If you’re opening things up to third parties, e.g. the Github API, complexity analysis becomes the more valuable solution.
Scheduled for the 1.3 release this coming week!