The GitHub GraphQL API
githubengineering.com
githubengineering.com
In hindsight, sending a query, written in a query language from client to server seems obvious. So obvious, that I think I've seen it before...
select login, bio, location, isBountyHunter
from viewer
where user = ?
It's ironic to me that it took Facebook reinventing SQL (or a graph database equivalent thereof [2][3][4]) and Github embracing it to legitimize this practice, since if you were doing this before, you were judged in the eyes of your peers and clients for not being "RESTful" (the fake-REST kind [5]), as if everyone was just itching to PUT and DELETE blobs of JSON of your poorly mapped resources to quasi-hardcoded, templated [6][7] URLs.What's old is new again, but this time I'll take it.
[1] http://graphql.org/blog/production-ready/ [2] https://neo4j.com/developer/cypher-query-language/#_about_cy... [3] http://tinkerpop.apache.org/ [4] https://www.w3.org/TR/sparql11-query/ [5] https://news.ycombinator.com/item?id=12479370#12480408 [6] http://swagger.io/specification/#pathTemplating [7] http://raml.org/developers/raml-200-tutorial#parameters
We coordinated with Facebook to do that. ;)
If you only give your users access to views instead of physical tables you get the same result as GraphQL, just with different syntax.
> over the Internet
What does this really mean? How does this substantially differ from the now-facetious term 'web scale' [1] ?
Modern RDBMS'es offer encrypted communications between client and server [2][3][4], but nothing is stopping a deployer from putting a proxy between the database and the Internet-resident end-user to translate between authentication mechanisms. In fact, this happens in every single API today, where the Basic or OAuth or Cookie-Auth request comes in the HTTP body, and gets made into a database query in a query language the DB understands -- SQL or something else.
GraphQL needs to you bring your own authn/z [5][6][7] much the same way that your run-of-the-mill HTTP API needs to you to bring your own authn/z, so I can't accept 'secure' as an innovation over SQL. "Scalable?" How?
I genuinely think this tech is neat but let's consider what it gives you. GraphQL packages together the query language known as GraphQL, the IDL/typespec language known as GraphQL, the server spec known as GraphQL which a GraphQL-compliant server must implement [8], inside which you have to hook up your actual datastore to the GraphQL server [9][10][11][12]. It's a useful tool to build a client-server system where you're going to make a backend query in the end, but I wonder if its full potential will be realized when there are GraphQL-capable datastores to skip that extra data munging layer you have to implement yourself; but if SQL were passed on the TLS-secured wire instead of this bespoke format that no DB yet understands, would we collectively freak out? What's the difference?
[1] https://www.quora.com/What-is-an-explanation-of-the-punch-li... [2] https://technet.microsoft.com/en-us/library/ms191192.aspx [3] https://docs.oracle.com/cd/E11882_01/network.112/e40393/asoj... [4] https://www.postgresql.org/docs/9.5/static/ssl-tcp.html [5] https://medium.com/the-graphqlhub/graphql-and-authentication... [6] https://github.com/mostr/graphql-auth [7] http://stackoverflow.com/questions/34952792/how-do-i-structu... [8] https://facebook.github.io/graphql/#sec-Execution [9] https://www.reindex.io/blog/building-a-graphql-server-with-n... [10] https://www.compose.com/articles/using-graphql-with-mongodb/ [11] https://medium.com/apollo-stack/tutorial-building-a-graphql-... [12] http://stackoverflow.com/questions/35940528/how-to-connect-g...
Aggregating various data points manually in a predefined way may be less flexible than GraphQL, but it's much easier to optimize when things get complex imo
Similarly, I'd expect Facebook to be thinking about Internet attacks due to real-world experience with hostile users all the time, versus a database company where the typical use case is logins from authorized employees and DBA's. They may try to get security right but it's not going to be top of mind and they have other concerns.
SELECT uid, name, pic_square
FROM user
WHERE uid = me() OR
uid IN (SELECT uid2 FROM friend WHERE uid1 = me())https://github.com/rmosolgo/graphql-ruby
https://github.com/shopify/graphql-batch
https://github.com/github/graphql-client
We <3 your work and are thrilled to have built this with you!
Please make sure to give us feedback during this alpha stage! https://platform.github.community/
And it's so similar to ORM's issues all the industry experienced past 20 years. But perhaps more dangerous due to the public nature of many APIs.
If the client requests exactly what it needs, that shouldn't be more stressful on the server-side than spamming REST requests for all the same resources. Plus, it's easier to optimize when you know what the client wants. If there's something expensive, you could, for example, cache/index something extra. If the client were doing it themselves with a series of REST calls, you wouldn't be able to understand the real use-case. Even if you did know what aggregation they really needed, you wouldn't be able to fix the problem without updates to both the service and the clients.
Either way, it's easier to set sane limits than craft un-DOSable APIs. There is always a cost to satisfying queries. If you're trying to run a free service, it's a much bigger concern. If you're paying the bill, you're incentivized to investigate expensive/slow calls.
The nice thing of REST calls in the current form is that they are that - just calls. With proper monitoring you could just see which ones do you get more or less and with these or those parameters. They can be optimized as best possible, but separately. You are right, it needs more analytics to figure out a series of calls (based on some token?) and maybe bundle them up, introducing a new endpoint (thus not breaking old clients).
But yet again, that is that one "query". With GraphQL it could be anything, and that's what bugs me. I find it challenging, in a good way.
Another thing what I'm also not sure about are the queries themselves, or rather, the number of different ways you can write a query. Multiple users can request the same data, or almost same, with queries written in different ways. Backend developers should then guarantee that those queries will be executed in a similar way, with predictable performance. I guess in a similar way SQL query optimization does. I had the "joy" of working with a database that had hugely different performance just with trivial changes in the query (it was not relational, actually it is discontinued now, thankfully). It was a huge PITA. I wouldn't like to serve an API like that.
GraphQL essentially moves a lot of complexity that you usually have on the client side (determining what things you need, making several subqueries, combining results) to the server side. This is great for performance or bandwidth constrained clients, but it might impact the required server performance in a negative way.
I experimented a little bit with GraphQL as a query language for services on performance constrained embedded systems instead of more RPC based approaches but have gone back to the latter one as it's far more predictable and I often didn't need the flexibility of arbitrary queries.
I would get more specific but it depends on which programming language you want to use. Check out the code examples & links to libraries in different languages on http://graphql.org/code/
This was pretty much all the documentation we had, and it's more a design analysis of edge-vs-node authorization: https://medium.com/apollo-stack/auth-in-graphql-part-2-c6441...
Edit: Our eventual solution looked a lot like
class SomeTypeOfResolver {
@allowIfAny(rule1, rule2, rule3)
someProperty;
@allowIfAll(rule4, rule5)
otherProperty = defineRetrieverFunction();
}Does anyone know any best practices if you want to adopt this is in an existing application using a relation database (i.e. PostgreSQL). I don't know how to implement this without causing N+1 queries. (or Worse).
For example:
{ Post {
title,
content,
Author {
name,
avatar,
},
Comments(first:10) {
..
}
}
}A naive implementation would cause a lot of query, for each "edge" a query.
Nobody has static files sitting on a server anymore except for static auxiliary assets (CSS, scripts, fonts) or if you're actually running a website with static content, which is exceedingly rare. Everyone has some kind of request router that parses the URL and the body and figures out what to do next, makes a query to a backing database, then assembles and massages the response to make it look like the mediatype the client expects.
And with clients basically creating their own queries, I imagine the performance implications will be less predictable than with a more rigid REST API.
To get better caching, one could:
- Canonicalize the query to identify queries that look different but are actually equivalent; use the hash of the canonicalized query as a cache key. You can do this on the edge cache, or on the backend.
- Cache more full-bodied resources, where additional fields are present, and perform the filter at evaluation time.
While you can indeed perform larger, more complex requests, GraphQL by nature forces queries to explicitly ask for everything you want to get back. As a result, we're not wasting any capacity giving you back a bunch of data for an entire object that you don't need like we would in a normal REST API request.
The thing that I'm most excited about with all this is the fact that we're building new GitHub features internally on GraphQL as well. This means that unlike a traditional REST API, there will no longer be any lag time between features in GitHub and the GitHub API.
API is a first-class product now. API consumers get features as soon as everyone else!
Please make sure to give us feedback during this alpha stage! https://platform.github.community/
In Postgres, straightforward approach to query such data is based on JOINs and it's absolutely inefficient. This can be dramatically optimized with recursive CTEs, arrays and loose index scan approach, but GraphQL by default it will do straightforward JOINs, right?
I hope GraphQL has (or will have soon) ways to overwrite/redefine queries, but again, it leads us to the same problems "patch driven development" that everyone hated in ORMs during decades. That's why I'm saying that GraphQL is "a new ORM", but it's more dangerous due to it's openness and proximity to web users, that's why it can bring even more dev and devops pain to the world that ORMs did during decades.
Just as a relational database itself can sometimes generate better query plans from its query analyzer if you feed it what you are really after in one big, slow query that narrows to very specific rows rather than lots of small queries that return lots of rows quickly. Amortized against the database's time (CPU, memory) and bandwidth that slightly slower query is still sometimes a big win for overall performance.
(Given that most GraphQL services are typically backed by relational databases, it should probably not be a surprise the savings sometimes get passed right along.)
The main riddle to me is: in case of RDBMS in backend, how can we guess in advance which indexes are needed and how can we forbid/limit all "heavy" queries?
Also, where appropriate, you can still gain a lot with caching.
Simple example: cache the query response containing all fields of a table, and extract the fields required for the response.
Even if I would use query caching, I cannot imagine how it would help me to deal with really long queries (say, lacking proper indexes). Caching doesn't help when your query runs for 30 seconds – nobody will wait (and produce cache for others) so long.
Experimenting with this we often saw 50-70% reduction in the payload being sent to the clients in some requests. If I only need the first, last and avatar from the User object there's no need for my response payload to suffer because other requests need 30 fields from the same object.
Implementing this without causing a lot of N+1 queries is the tricky part and that's where we're currently investing most of our time.
Awesome to see Github adopting this and releasing it to the public API.
That's the dream. We'll see how reality plays out.
For reference, we actually launched with some deprecated fields (see "databaseId" on the "Issue" type -- database IDs will be phased out for global relay IDs eventually) if you want to see what they look like.
Granted, our clients are all Facebook engineers, so we have some pull in helping the migration away from deprecated fields, and GitHub will have to find the right process which works for a broader set of API consumers, but not only is this theory a good one, it's considered GraphQL best practice.
I've sparingly in my free time been working on a project that does exactly this with a REST API[1]. It's in an entirely unfinished state, but the linked documentation is a decent example of the types of queries possible.
- You can pretend that each GraphQL query is a JSON object (which, it actually is)
- You can pretend that each GraphQL schema that you declare is actually a JSON-Schema document, which some people use to specify in a machine-readable way your API's inputs and outputs will look
- You can pretend that each GraphQL resolver, which the piece of code you have to write (on the server) to actually dig up the result of a query, is a function that parses your incoming JSON, validates it against your JSON-Schema, and then reaches out to your datastore to produce a result. You'd then have to construct another JSON document which matches your response schema, stuff the data in it, and return that to the user. Except that in GraphQL, you only have to supply the resolver, the rest is handled by the framework.
You can of course do this by hand and many people do (most obviously when you see APIs that include arguments like "operator=eq" or "limit=100" or "page=25"), but GraphQL gives you the tools to do this with less effort, and end up with a cleaner API by passing everything in the query body. And the GraphQL server saves you from having to manually build up the JSON text of every single response.
Reading through your docs, I've seen this style of API in enterprise settings where there was a backing relational database and the designers were basically trying to expose the underlying database through HTTP. It can get the job done, but GraphQL gives you nicer abstractions, a cleaner way of passing parameters, and conveniences like a real type system (known both on the client and server side) and you only have to supply your resolver function implementations.
Still, pretty much everything else you mentioned seems doable with REST. I'm using Marshmallow schemas in Python which seem to act in a similar way on a field by field basis as a GraphQL resolver does. I'm not sure what exactly I'd be getting outside of a slightly nicer/cleaner abstraction by moving to GraphQL, but maybe that's enough?
Public and private This application will be able to read and write all public and private repository data. This includes the following:
Code Issues Pull requests Wikis Settings Webhooks and services Deploy keys
Why'O whyai?
As a personal opinion, I also feel OData exposes an API that is too tightly coupled to the persistence layer. GraphQL objects and properties are all backed by arbitrary "resolver" functions, which means you can stitch together multiple/legacy backends to generate your response.