Scaling to 1M active GraphQL subscriptions on Postgres
github.com
github.com
We evaluated a lot of stuff recently before settling on Hasura. Prisma, aws appsync, postgraphile, roll our own socketio backend etc. I had a few hesitations about Hasura. In particular I was hoping to lean on Postgres RLS for the authorisation, and to be able to customise the graphql endpoints a bit more. After hearing the rational from the guys about how their system is better than RLS and seeing the patterns they have for implementing the logic I’m really happy we made the leap.
We’re still in early days, but it’s been really solid and a real pleasure to use.
(No affiliation, but happily paying for their support licence)
In a previous project, I used Postgraphile and enjoyed it. Sure, RLS and PLPGSQL were a bit awkward, but I knew they were well-tested and had lots of eyes on them, so I felt comfortable that if an RLS policy blocked access to a row, then that particular role wasn't getting access. I also enjoyed working with a text-based format that more clearly fit within my other workflows of, well, working with other text-based formats. :)
Another part of it is that I'm blind, and Hasura seems to make doing things outside of its web interface somewhat painful. Sure I can write a YAML migration by hand, but the YAML migration format seems more machine-friendly than user-friendly. Yesterday I needed to create a function, and the SQL web interface wasn't showing me the line number on which my function had an error. Ultimately I dropped to the psql command line, cut-and-pasted the function in manually, and right away got back the line number and fixed the issue.
Please don't get me wrong. I'm not trying to hate on Hasura. But the fact that I can't just drop to a non-YAML text-based format and throw down a few checks to secure my tables has been an endless source of frustration for me. So if there is a non-NIH reason for abandoning RLS in favor of a separate security system, then maybe knowing that might help me be a bit less annoyed. :)
In terms of the setup and migrations, I also have the same reservation about it being harder to do with text. So far I’ve left it to one of the other guys to implement so it hasn’t been painful for me, yet!
Edit to add: I believe they built their version of RLS before Postgres had it. So it was more a case of Not Invented Yet
You can also add sql migrations as a sql file and that’s something we should document better! https://docs.hasura.io/1.0/graphql/manual/migrations/referen...
Coming back to RLS, doing what we’re doing to fetch multiple results for different clients is not easy. Connection variable vs being in the same sql statement essentially.
https://github.com/hasura/graphql-engine/blob/master/archite...
Further, owning RLS helps us target more hosted Postgres vendors, and other Postgres flavours that don’t support RLS well. RLS on Heroku doesn’t work: https://devcenter.heroku.com/articles/heroku-postgresql#conn...
Owning RLS does have a few other advantages. Having a unified experience in bringing that authz experience to “remote schemas”, for example.
(Typing on my phone, apologies for typos etc)
As a blind person who works more productively in a text editor than a web interface[0], here are a few things that might help: * Give me the ability to enter JSON directly in the permissions screen. I'm sure the query-builder is visually nice, but I can accessibly edit JSON strings better than you can accessibly visually render them, so let me just give you a string and have your interface validate it. Maybe just parse it, then transpose the values directly onto the form inputs in the visual interface. * Give me a command to validate YAML migrations without applying them. This could even be a --dry-run flag to `hasura migrate` telling me what SQL would run and the JSON for my new/changed permissions. * Some basic documentation on hand-coding migrations would help. I can sort of reverse-engineer them by pouring over the file formats, but the time taken to work more slowly in the web interface hasn't quite inspired me to take more time learning the undocumented format. :)
I'll see about filing these as issues soon if I can remember to do so. :) Either way, thanks for a cool product! Even if it frustrates me to use, tools like Hasura really hit a pain point for me in building apps.
0. Obviously that isn't universally applicable, but it is easier for me to solve logic problems than "why is my CSS all weird?" problems. Tools like Postgraphile and Hasura make it possible to test more of an app's functionality by writing just SQL rather than lining up my ORM with my language with my SQL schema. I guess just as Node supposedly reduces context-switching by placing more logic in a single language for front-end developers, tools like Hasura and Postgraphile let backend developers focus mostly on SQL when building out the app logic. And while I can't make something look good, damn can I think through how it should work and how someone might break it. :)
I've created an issue here for now: https://github.com/hasura/graphql-engine/issues/2310
Feel free to add more thoughts/suggestions!
We had to mark many internal functions as LEAKPROOF in order to get them to perform. It was not possible to move our database into RDS, so in the end we had to abandon the RLS authentication layer.
1. The overly complicated schema it generates. For example, given an integer id, one can (and must) use various operators to query that id, e.g. `user(id: { _eq(1) })` rather than just `user(id: 1)`, where other operators are `_gt`, `_lt`, etc. To make this possible, hasura defines a comparison object, which results in a very long schema. However, often the only comparison that makes sense is equality, e.g. for IDs, so this should be optional.
2. More importantly, using Hasura establishes coupling between the DB and the GraphQL schemas. What if at a later stage one decides to remove information from the DB and move it to an external source? GraphQL was designed to hide such implementation details. With Hasura this is no longer possible.
We’re also making everything in the Hasura schema aliasable via metadata so this becomes a more elegant name if you’d like!
Not sure how people enjoy FP.
~Earl
Step one, which applies to most functional languages, is knowing how to read a Hindley-Milner type signature [1].
Armed with that you can often gain a surface level understanding of what some function does - and what the program does in aggregate - just by reading type signatures and seeing what calls what.
Hoogle [2] is a good resource for figuring out Haskell syntax, like what's up with all of those $ signs.
Converting OCaml to ReasonML syntax with sketch or reason-tools [3] makes OCaml less foreign.
With basic knowledge of Haskell and ML syntax you'll find many FP languages readable.
[1] https://drboolean.gitbooks.io/mostly-adequate-guide-old/cont...
That being said, my experience with Haskell is limited to working on some meta-meta-programming research in college, where every file used at minimum 3 different GHC extensions. So I may be biased.
are you kidding? most files i see have all of these
Exception: apl.
The number of times I have wished for something like Hoogle when writing Python...
I’ve noticed that a _a lot of_ stuff written about FP tries to emulate this fun and easy going style - I wonder why that is. I still refer to clojure’s “re-frame” docs as the epitome of technical literature - a text so good you can read it in its own right just for the fun of it.
I don’t want to diminish other languages’ communities of course, as this is totally anecdotal. I’ve found awesome prose in other places as well (mostly ruby/python) but I just encounter those funny gems more often reading about FP.
So like anything else?
>Not sure how people enjoy FP.
By virtue of actually having tried it?
In contrast I picked up Rust and Elixir in a couple of weeks apiece, so I don't feel like I'm particularly dense; I just suspect that it needs a certain way of thinking to get the best from Haskell :)
import qualified Data.ByteString.Lazy as BL
It's ironic that after all this obfuscation, the author of the source still finds it useful to do ASCII art and align everything neatly. Talk about misplaced effort.See The Blub Paradox [0]
I actually find their code quite readable, but that’s because I have experience writing Haskell.
This fancy "subscribe to events when the results of this query change" system boils down to just polling the database...
Future work:
Reduce load on Postgres by:
- Mapping events to active live queries
- Incremental computation of result set
Fake it till you make it.Since Clojure is pervasively built on laziness, there has been some discussions in the community about doing the same but for incremental computation. And doing it pervasively most importantly guarantees it will work at any granularity without any manual fiddling with observables and reactions.
A recent contribution to the solution field in Clojure is micro_adapton [4][5] which looks like it is to incremental computing what miniKanren is to logic programming, i.e. a tiny core for a a challenging problem.
Also, considering the literature around it it seems to be pretty tricky to do properly, especially when it comes to databases [6].
[1] https://github.com/janestreet/incr_map [2] https://github.com/janestreet/incr_select [3] https://github.com/mobxjs/mobx [4] https://github.com/aibrahim/micro_adapton [5] https://arxiv.org/abs/1609.05337 [6] https://scholar.google.fr/scholar?hl=fr&as_sdt=0%2C5&q=incre...
Only some of Clonure’s data structures are built on laziness, everything else in the language is strict. If you want something actually built on laziness, you need Haskell.
https://www.postgresql.org/docs/9.6/logicaldecoding-explanat...
Thank you to the Hasura team for all you do and in turn allow me and my team to do!
For folks that want to play around with wal and listen/notify:
https://github.com/hasura/pgdeltastream https://github.com/hasura/skor
[1] https://github.com/hasura/graphql-engine/blob/master/archite...
I wish Hasura the best and would be cheering for them from the sidelines
It seems to expose your internal data structures, and any change to those will immediatly impact that public GraphQL API. Also it seems that this approach would only work well for applications without any business logic, or with business logic solely implemented in the database. At least the applications I've been working on wouldn't fall into this category.
But in any case, impressive work and demos!
But you'd have to handle changes through another API, or at least route them differently, and then worry about consistency so probably not ideal for all business cases.
We’ve added event triggers to Hasura to kind of support this pattern via Hasura itself. So you can create “action” tables that basically have a log of request data (a mutation inserts an action). Hasura will then call an event handler which can run with the action data, user session information, related data in case there are any relationships etc. This handler can then go update the tables that can be queried from.
Ofcourse, if the eventing system is in-order / exactly once quite a few use-cases become feasible ;)
Our auth system has many roles per user, then each role has its permissions, pretty common setup. The problem is that Hasura expects a single default role per request to then evaluate against its permissions. The Hasura team has been looking into accepting an array of roles and merging permissions and whatnot, but AFAIK this hasn't been solved.
If you start with Hasura from scratch this is not really a problem, it happened to us because we had to figure out how to integrate Hasura with our current permissions.
But integration with external authorization systems is an extremely basic requirement for every piece of software, certainly on the top 5 on the requirements list for every serious solution - I can not understand how did you forget about this?
In Hasura's case, it looks like the flow of data to the authorization bits shows a simple/naive approach, although it also appears this is something Hasura are working to correct. Even as it stands, moderately complex authz logic could be implemented via x-hasura headers or in-database associations.
Hasura GraphQL subscriptions are over websockets. This post is a benchmark that should address those scalability concerns. It’s about getting to 1M concurrent active websocket connections.
What’re you looking at?
Is supporting CockroachDB something Hasura is actively working on?
Don’t want to spread ourselves too thin too early. Also working with the yugabyte folks closely to have Hasura supported with their 2.0 release.
https://twitter.com/karthikr/status/1128003465370693636?s=21
I certainly understand that, but I wonder since CockroachDB is essentially Postgres, if it wouldn't be easier to ensure compatibility from the start. It seems like it'd be easier to just avoid using special Postgres features that aren't in CockroachDB than to go back later and re-write all your code that relies on those features.
Citus / timescale / pipelinedb are much easier and we already support timescale and pipelinedb well. Citus with native support for their distribution columns is going to happen soon too! These are packaged as Postgres extensions which is definitely first priority for us before moving to databases that speak Postgres and eventually to other databases as well. :)
So you're waiting for CockroachDB to implement features before you suppoort it?