Show HN: GraphCDN – GraphQL CDN with edge caching and analytics
graphcdn.io
graphcdn.io
When I built my last startup, Spectrum, we spent months building custom caching for our GraphQL API from scratch. It never worked well enough to alleviate our scaling troubles as we could only cache data for unauthenticated users since we had no invalidation. (since we open sourced it all before GitHub acquired us, you can even read through my terrible code[0])
When Tim told me he had built a prototype of a CDN specifically for caching GraphQL query results with proper invalidation, my first thought was: "Finally!" Not only had he made it possible to cache POST requests (which GraphQL requests usually are), he had made it possible to purge cached query results per specific GraphQL object. For example, when a user edits their name the API can call a purgeUser(id: $ID) mutation and any cached query result that contains that user's data is invalidated.
GraphCDN is based on Fastly Compute@Edge under the hood, which is really the main reason we were able to spin this up so quickly. Huge shoutout to the folks building that!
We'll be around all day to answer any questions you have about GraphCDN — ask us anything!
[0]: https://github.com/withspectrum/spectrum/blob/alpha/api/apol...
I'm just wondering why you decided on Fastly over Cloudflare Workers or AWS Lambda Edge for this?
Fastly has much faster cache purging - it can purge any content globally in about 150ms. Purging is one of THE crucial parts which make GraphCDN work. With such a fast purging, we can deliver Read-after-write consistency . In our tests with Cloudflare that took a few seconds.
Additionally, in our tests, Fastly was in general just a bit faster.
While we're also JavaScript and V8 fans, our edge layer on Fastly is written in Rust - which is a pleasure to use, especially for a product as ours. Fastly runs it as WASM on the edge - they can load the worker in about 40 microseconds.
Fastly Compute@Edge is still in limited availability, but if you get a chance, I highly recommend checking it out!
Based on the comment on max below that's not read-after-write consistency but it just becomes eventually consistent but ideally within ~150ms. What would be the consequences of bypassing the cache for read after write, does that break your model in any way? (I guess at the very least you will lose your statistics feature)
You can bypass the cache if you want to, but it's not necessary. As it would just be like a Cache MISS, that would totally work.
We have a limit - you can only "tag" up to 1k items within one query. If you have more items in there, we can't purge all of them.
Due to the nature of the broadcast protocol of Fastlys purging implementation, I assume that it scales quite well - they have really big customers using this for a few years already.
We're using Fastly ourselves and I was not aware up until today (and in fact I'm taking your word for it, since I can't find it in the docs) that fastly provides consistency guarantees for purges.
Hmm, how can that be? Light in a straight-line fiber optic cable would take 200ms to traverse a greater circle around the world. Add some bends, relays, routers, not to mention servers, and it'll only get longer...
Are you sure the content is really being purged globally in that time, and not just from your local point of presence?
(Disclosure: I'm an engineer on Cloudflare Workers. I have no idea how Fastly does purges. Just trying to understand what you mean here...)
Looking at the 150ms now, which is about 2.3 times the 67ms, this kinda seems reasonable.
And yes, this is global purging - I just reconfirmed it with someone from Fastly.
Travelling half of the distance is not good enough, because if you haven't received _confirmation_ of the purge, then you can't trust that it really happened. There could be network errors, etc.
Sorry but it's not physically possible to do a global purge in 150ms.
What was it about this problem that made you need Compute@Edge rather than the standard Fastly/Varnish VCL-backed caching? If you tried that route, how far were you able to get with Varnish before needing the new Compute@Edge product?
We also support a feature we call "scopes" - where we both support cookies and headers in the "Vary" header so you can cache user-specific data only for the user with a specific "Authorization" header for example.
In order to support scopes with cookies, we needed Compute@Edge and could not make that work with VCL on its own. But believe me - we tried. Our first version was actually mostly built on top of VCL.
Could you not abandon GraphQL? Returning non-customized responses, while taking more bandwidth, is much more cache friendly for the client, proxies and particularly servers - where you can do simple yet powerful things (e.g. version-based invalidation and pre-generated payloads).
Like, I went to https://spectrum.chat/explore and clicked on "Tech. It uses GraphQL to load the communities.
Why not just hit /v1/communities?category=tech
which would loosely translate into one of:
-- if you want to serialize the object on each get
select * from communities where tags @> array['tech']
-- if you can look up the communities in a cache by id
select id from communities where tags @> array['tech']
-- if you have a high read / write and can just pre-serialize the payload and then glue together the response
select summary from communities where tags @> array['tech']
If I understand correctly, you're suggesting to cache on a data-layer level (below the app in the stack) instead of what we do - above the app in the stack - just in front of the client.
That is also a valid approach. It'll need more custom code in your application and has the disadvantage, that your origin will still be hit on every request, while we fully cache on the edge - cached queries won't hit your server anymore and will take load off origin and have the minimum latency possible.
> making sure, that when content changed, to invalidate it properly on the edge is not trivial
The trivial way to do this is simply not to invalidate. Include a version in the cache key. When content changes, the version increases, and clients get new versions. (I think this is sometimes called lazy cache invalidation). It works well with LRU caches. It's still a hit for the top-level keys, but you avoid all the heavy data load and rendering.
Granted, that doesn't solve every case, but I'm not sure it's fair to say that it'll take more custom code when the initial approach took "months building custom caching" that never really worked. Also, I think you'll find your origin under less load due to having fewer variants.
[0] https://blog.codecentric.de/en/2020/05/how-to-secure-a-graph...
I'd love to see a write up/blog post for working specifically with Hasura and their auth system. Would probably be helpful for seo purposes as well.
Are they supported?
Do they count as a single or multiple requests for pricing purposes?
Do you/will you handle caching individual request layers? For example if I need part of my query to always be fresh, but some expensive part can be cached for hours, is this possible? And somewhat related, what about keying on operation variables?
We don't currently cache partial queries / request layers. For now, I would recommend splitting the query into two requests — one that loads the uncached data and one for the cached data.
We're definitely thinking about this exact use case though, stay tuned!
Partial query caching appears to be the unicorn we all would like to chase, capture, and ultimately study.
GraphQL is powerful, But with great power, comes great responsibility, and it is too easy to dig your own grave if blindly jumping in with it.
I wonder if we can learn a thing or two, or just flat out steal, some of the concepts used in flagship relational databases. For example, how Postgres has a query optimizer tucked away secretly under the hood that attempts to alter queries to be more efficient.
Does Compute@Edge have this limitation?
It's a pleasure to share this announcement with you! In all the GraphQL projects I worked on, it was always a pain to get the caching and security right.
Instead of you all spending time on building your own caching and security solutions, you can check out GraphCDN! It has powerful caching with invalidation in 150ms all around the planet. Powerful analytics showing you on a Query-level, how fast your queries are.
We're super grateful to be able to announce this today - ask us anything!
How do you handle smart invalidation? Or more specifically, how do I trust that you're handling smart invalidation _correctly_. Looking at the site, it indicates that calling a mutation like `editUser(id: 5)` would presumably invalidate the User type record with ID 5.
But how do you actually do this reliably? There's nothing in the spec that would indicate the argument ID maps to a record of a certain type with an id field of the same value. Max indicated that you make assumptions based on the return type of the mutation, e.g. editUser has a return type of User, therefore you can infer the relationship. This might be _generally_ true, but it's not 100% reliably true. Additionally, my mutations _never_ just return the naked entity type like this, there's always a wrapping payload type (philosophically, the mutation payload should contain points to all the parts of the graph that _may_ have changed as a result of the operation). Editing a User doesn't _just_ edit a User, the effects on the graph can propagate far and wide. Another point here is that my mutations are rarely just CRUD operations, but more CQRS in nature, they're built to support a specific system capability rather than allowing generic write operations.
The problem of smart invalidation seems to have the exact same shape as smart store invalidation/updates after a mutation in the client. Even after all these years, Relay only does fairly superficial automatic updates, you nearly always have to use custom updaters (or client-side directives in some cases) to get the client-side store back in sync after anything but the most trivial mutations.
We do support wrapping payload types of any kind because we invalidate _all_ objects returned from a mutation. For example, if you run a mutation like editUser { user { id posts { id } } } we will invalidate any cached query result that contains that user and any cached query result that contains any of those posts!
You are right that smart invalidation can never be 100% reliable, which is why we have the Purging API to manually purge records you know changed from your backend. I think most customers are going to use the manual Purging API, however we also have a bunch of customers with use cases for whom the smart invalidation suffices.
As far as I have understood.
You can check out some more examples here: https://graphcdn.io/docs/cache-purging
This has a couple of limitations that you'd also expect from a CDN cache for REST requests. However, I believe the interesting part about GraphCDN is that it can do more to look at the exact queries and mutations that are run to invalidate queries more precisely.
So, it's likely worth saying that it's not that CDN caching GraphQL is hard, but getting invalidation and a high cache hit rate (just as with REST APIs) is hard.
I don't think a normal HTTP api provides anything to help with cache invalidation after a write operation, does it?
Etags can be used to check and invalidate content.
However, that is missing the point. What we need to talk about first is, how you want to invalidate your cache. Do you want to set a TTL of 60 seconds? That might work for certain apps - both in REST and GraphQL, but many apps can't afford stale content for such an amount of time.
You'll need cache invalidation when content changes. That on its own is a hard problem, no matter if REST, GraphQL or any other protocol. And it is one of the main reasons we built GraphCDN: Making it easy to purge the cache, when relevant content changed. How? We give you a purging api (also GraphQL) and additionally GraphQL has the concepts of mutations. Once you run a mutation through GraphCDN, it'll detect the relevant entities involved and purge the cache accordingly.
So - yes, in GraphQL caching on the surface might seem harder - but we're not just solving the "I can't cache POST requests" problem, but rather give you powerful cache purging - which is only possible due to the well-defined structure of a typed GraphQL Schema.
Because of that, we're actually thinking of providing REST "connectors" one day - turning REST into GraphQL, so you can have one unified interface that is easy to cache and invalidate.
Or you can use Vercel's stale-while-revalidate which will update the cache periodically while (temporarily) serving stale responses.
Most of our customers even use both things together to reduce the likelyhood for stale content.
The frontend teams may not want direct access to a backend teams database. But they do want the backend to be flexible, and graphql allows for that.
It solves some problems on the querying side (that not everyone has). OTOH implementing it server-side is a major pain unless you rely on third party stuff like Hasura or GraphCDN.
Personally I'll keep using REST as the default for the foreseeable future, and only use GraphQL when the problems it solves are more painful than the problems it introduces.
I'm not even sure "major pain" suffices to describe it. It's such a mine-field to implement that if it's tractable and non-insane to do so for one's project, then the surface area of one's API must have been so tiny that using GraphQL was entirely unwarranted in the first place.
[EDIT] I take that back: it can be fairly easy if your dataset is 100% public, read-only, and you don't care whatsoever about performance or limiting abuse.
There are a lot of plays in this space that try to move the database, or serverless-functions closer to the end-users, but in all likeliness if you're already building a single-page-app the static content is already on a CDN and close to your customers, so this gives you a very easy way to increase performance dramatically without having to modify your infrastructure.
Awesome!
So we're happy about this level of abstraction, because any app can use it _today_!
Some questions:
How it compares with Apollo Cloud on feature set terms?
My graphql server load is like 20 request/s average. At first the pricing looks a little bit intimidating for me, but running the numbers it looks like $500/m, is that right? Hopefully it will offset some of my origin servers costs.
What count as a request? Just request coming from the “outside” or also calls to purge for example?
I’ll be trying GraphCDN soon, maybe even today.
Good luck
It was cheaper to build our own solution with plugins than it was to use their solution.
Compared to Apollo Cloud: We're mostly focused on the caching part right now and have a different architecture where we are in your stack. Apollo runs a sidecar next to your application. We are a proxy in front of your API.
When it comes to the analytics part - which Apollo rather calls metrics, I think Apollo gives you field-level information, while we for now just have query-level information. However, we are fully server agnostic - you don't need to use Apollo Server. Any GraphQL API works. You just need to switch the URL in your clients. We even have customers just using the analytics part for now and disabling the caching in the beginning.
For the pricing: That is correct - you'd have about 50mio requests a month, so $500. However, the pricing there is not set in stone and we're happy to give you an early discount. Just contact us at support@graphcdn.io.
Right now only outside requests count as a request, no matter if cached or not. Purging calls might also count in the future.
We use Apollo client, I’m worried about 2 specific features of it:
Batch queries [1], we’re currently using it.
Persisted queries [2], we’re planning to using it.
Are those compatible with GraphCDN?
1. https://www.apollographql.com/docs/react/api/link/apollo-lin...
2. https://www.apollographql.com/docs/react/api/link/persisted-...
I’ll try to keep myself updated on your changelog.
1. Cache invalidation 2. Decent project name
Joking aside, this is great. Traditional CDNs + GraphQL always felt like an impedance mismatch.
One of them actually recently looked into building their own solution. However, they realized, that in order to create something really valuable, you need at least a couple of months of dedicated engineering efforts of people who really understand GraphQL.
Our automatic and powerful cache invalidation - currently only possible with Fastly, a whole GraphQL-specific Analytics solution and a Security suite is nothing you can quickly clone.
Anyone of course can, but you need a highly GraphQL-specific product with many workflows like CLI workflows to upload your GraphQL Schema etc - it requires quite a bit of thinking and engineering to make that work.
IMO if you get traction, you would be wise to take an early acquisition offer from Fastly or Cf. I was an early double digits engineer at Cloudflare and I can tell you, you REALLY don’t want to build a CDN to try to compete. Love them to death but look at ImgIX as a case study of how not to bizdev this same business model.
Also, don’t entertain the idea that you can compete with a DIY VM-based “Cloud CDN”, it’s really not comparable. Look at the many dead “mobile first” CDN startups of 2014-2018, as case studies.
Lastly, if you guys create a caching query planner on the edge, you might have a head start on a completely unmatched (afaik) product as an intelligent graphQL “gslb”. As a customer, if I can move my user data geographically closer to my users, that’s a big win for many reasons.
Best of luck!
Also to seed an idea for you, it'd be great if you were somehow able to provide subscriptions dynamically based upon queries and mutations being performed.
Acting as the middleman, you can see the freshest data and so therefore know when something updates. If I could hook that up to my existing GraphQL API and not have to worry about eventing and subscription services for every single object that would be a huge value add for me.
"Automatic subscriptions" is definitely something we've considered offering, but isn't on the immediate roadmap for now as we want to focus on the "peace of mind" first before expanding from there.
Really don't want to maintain a server. I love firebase for that reason, although querying it is a pain sometimes. Would love to make graphql queries so I can fetch multiple things in a single call.
the parallel "||" calls (push or pull) are great and the power to switch off a call you send with ? after the function name means all your calls during dev can be in a single text page and you just switch them on/off at will.
To be fair, we use the Tailwind system, but a lot of customizations on top.
If so what are the potential read/write/latency speeds we can expect per object?
The purging takes ~150ms in the same datacenter that the mutation passes through, how long it takes globally depends on Fastly's bimodal multicast system — they've written a fascinating article about it that I'd highly recommend reading: https://www.fastly.com/blog/building-fast-and-reliable-purgi...
> our 58 data centers worlwide
*worldwide