The dark side of GraphQL: performance
twitter.com
twitter.com
graphql-js is the reference implementation of GraphQL, so it's not any random library.
* Though Ent’s creators emphatically claim it is not an ORM, that phrase does give you a rough but directionally-accurate idea of what it is.
There is nothing in the posts that identifies NodeJS as the culprit, and based on the info I'd be very surprised if it was. It seems most likely that the type validation is what is taking so much time. But then again, strong types are one of the main benefits of GraphQL. If anything, I've found Node to be one of the easiest and most "natural" server languages for GraphQL, and I have implemented GraphQL servers in Node, Java and Python.
But if you're fetching a lot more data than you need for a typical UI, you might run into bottlenecks.
I'm still surprised it had that much overhead though, as I generally think of node as being relatively competitive with Java outside of heavy computation.
One issue I see a lot of (which happens to be part of the problem in this case) is people overfetching data on their list UI (e.g. search results) so that the detail page data is already in the client side store. But this is usually a premature optimisation and ends up causing more problems than it solves.
If you click on that you can get performance flame graphs / tables, memory profiling, and there's also a REPL for the process, as well as a list of loaded source files so you can set breakpoints through there if you like, as well as modify the files on the fly if you need something more for debugging.
It can be very useful, and works pretty much the same as the normal web devtools.
For example reference implementations of JSRs have very very rarely been usable/used in production.
I always expect the reference library to be optimized for correctness, formal proofing, brevity, simplicity and clarity, not (compilation nor execution) speed.
I don't know why anyone downvotes as they do, but the previous post is an irrelevant argument about semantics, so in my opinion it deserves to be downvoted.
Actually, now that I think of it, it's a little worse than that. The OP is being criticized for not understanding how to debug or improve the performance of their dependency while actively engaged in figuring out how to debug and improve the performance of their dependency. (People respond with questions, OP provides substantive answers, there's a back-and-forth and OP forms an idea it's related to a deeply nested schema, and so on.)
Nothing I said is a critique of Ben Awad, the Twitter OP. I assume we've all used dependencies that we don't completely understand, no?
To rephrase my point: The dark side of semi-opaque dependencies like ORMs, application frameworks, etc. is that when the magic doesn't work, one may not be able to easily determine where to start in order to address issues.
EDIT: Removed unnecessary paragaph to simplify response.
Maybe the optimizer picks a poor plan and you can't figure out how to make it work better. Maybe the schema has redundancy you can't change or the indexes aren't suitable for that query. Maybe it's auto parameterizing constants and the query with the problem has a parameter causing different behavior than the original constant used in optimizing the query, or maybe your query with 1000 elements in an in list worked great in memsql or whatever but is slow unexpectedly in the database you ported your app to. There are downsides to everything.
I also don't think it is hyperbolic say he doesn't have basic debugging skills. What qualifies as basic debugging skills? Like he isn't capable of using a debugger and introspecting code? He can't use a print statement and look at code? Debugging an E2E bottleneck is not trivial.
me { company { listings { projects { id name } } } }
This will initialize: a User, a Company, Listings and Projects of all listings.I can also write this in SQL using a couple of joins and return an array. The memory consumption is trivial in comparison to the original request.
GQL is an idea not an implementation. I don't believe there's anything preventing actual software from optimising this case. Or am I missing something here? The query defines what you're asking for so extra data does not necessarily need to be fetched.
but I do think it's related to my nested object https://twitter.com/benawad/status/1212407236284338176
In Elixir with Absinthe we can resolve to the specific fields we need and we don't load the entire records then slim down.
One annoying thing Absinthe did by default I noticed was fetching entire object from DB, even though the GraphQL query only returns it's ID. For example query below would fetch each person from a DB, even though we already had a list of person IDs on the friends level:
{
friends {
person {
id
}
}
}Which then makes me wonder why this is the "dark side" of GraphQL? Isn't this just not optimizing a query somehow or using a cache effectively? Is it really the nature of GraphQL that's causing this to be slow or just programmer error [1]?
I've used GraphQL in production services as an alternative to a rest endpoint (which I didn't care for) and I don't think sheer the nature of the validations ever caused that much slow down or rather, more plainly put that GraphQLs design would not necessarily cause such poor performance on such a small set of data.
/shrug I dunno, if this were me, I'd just assume I had written a bad query or validation somewhere. And to be fair to the author, they only made a post on twitter to reflect on their problem, not to say GQL has a dark side (at least from what I read in the thread).
[1] We all have made programmer errors, are likely making some "now," and will for sure make more in the future. No reason to feel bad about it, we're all human :) Mistakes are just a part of life no matter how "good" at things we are.
'Slow response times for large documents'
https://github.com/graphql/graphql-js/issues/723#issuecommen...
It seems that the graphql library is performing a lot of validation, and that's slowing things down. I expect validation to be a pure-compute task, and this is Javascript, so I suspect this is really a "working with large amounts of data in Javascript is slow" issue - but that's just at a glance.
I've ran into similar issues with FastAPI and DRF when dealing with really large payloads.
We had something like this in our backend, but this long times is usually meant that something wastes event loop and just blocks everything from execution.
It could be anything for example it could be async hooks that makes ~1000 times slower if you are using a lot of promises (since resolving fields often are just promises) since overhead is per promise. In general in latest nodejs you can do huge amount of promises and they have little to no overhead, but, again - something wrong with nodejs setup, some library populate event loop or something deeper in nodejs internals. It is not an issue with gql itself since if you have gql performance issues that means that your server is super slow in processing like anything. Our team was shocked by performance and it turns out that NodeJS is super fast and it is some libraries (like sequelize) that kills the performance, but gql is not one of them.
Like react, it eschews performance for the sake of enterprise level scaling. This shouldn’t come as a surprise to anyone, being both of these came from one of the largest dev organizations in the world.
That's strange, because I thought the main selling point was to consume only the data you need. The client specifies exactly which fields it wants. Then it doesn't over-fetch. To make things higher performance.
It's true that the client doesn't over-fetch (and also doesn't need multiple round-trips), but at least when I tried the gql-js library it required the server to over-fetch: it would ask for individual records, and then do the field plucking/record joining itself; there was no way to intercept the query along the way to find out which fields it needed so you could only fetch those.
I get the impression that the server libraries were designed to work with a document store or "fat" REST API that is only capable of taking a single ID and returning the entire record. In this situation it makes sense to have a separate middleware server to keep the big fetches and round-trips inside the datacenter and only give the client exactly what it needs, and needing a little more server power isn't a big deal. But, if you need to do something more sophisticated (even something as simple as only fetching certain fields from the datastore), they were no help whatsoever; when I was looking into it there wasn't even a way to parse the query into an AST and do the rest of the query planning yourself.
I think there might be a disconnect or misunderstanding in the developer about this. GraphQL is sort of like the Flux pattern for MVVM architecture. It isn't so much a thing as an idea.
Your mileage may vary, but on the project at work where we tried to utilize GQL, it became apparant that it sometimes is incredibly complex to map between a GQL query and an efficient db-query.
We're slowly migrating to REST for preformance's sake. The front-end devs might complain since they can't just write a query to get exactly what they want, but I'm quite sure our customers won't complain about 50KB being served in 30ms instead of 10KB being served in 400ms.
But as you start to re-use REST APIs across multiple pages/apps, you often end up:
- over-fetching: you'll receive some data that you don't actually need, but that it's required by another page
- over-querying: the data you are over-fetching might have required extra queries (or even worse calls to an external system)
- cascading requests: if you are working with nested data you might have to call the server multiple times, often in a sequential manner
Also in my experience the performance of REST APIs tends to get worse over time, because as many developers work over the same APIs it's almost unavoidable to keep these APIs lean. Just before Christmas we had to spend some time figuring out why a fairly simple GraphQL request was taking almost 2 seconds. Turns out a developer had accidentally introduced an n+1 in the underlying REST API. The n+1 was on a field/relationship that we didn't need/use, but with REST you don't usually get to pick what you want to load so...
Solving these problems with REST is possible but not trivial, while GraphQL mostly solves them out of the box. So while REST could be faster than GraphQL, in my experience it usually ends up being slower.
The main selling points for me are:
- speed of development: once you have a complete graph, you can add new pages/features at a fraction of the time, and often without touching the backend at all
- type safety: you can generate typescript/flow types for your queries, giving you type safety from db to client (assuming your backend has types)
- query co-location: you can have the query (or a fragment) inside or next the component that uses it. Need a new field in a specific component? Just update the fragment, any page that includes it will get it by automatically
and last but not least developer experience. Having worked for a few years using graphql + apollo + react + typescript (on both a personal and a reasonably famous large website) I can honestly say that it feels like living in the future. As a mostly backend developer, I have never enjoyed working in the frontend so much.
I'd love an explanation for why you'd think this to be generally true, and for any category.