Blueprint for a distributed multi-region IAM with Go and CockroachDB
ory.dev
ory.dev
Personally, I am extremely proud of the work. I believe that in a year or two, most companies will adopt multi region IAM (hopefully from Ory as we’re currently the only ones capable of this). :)
And what could be better than hearing these kind words from the critical readers on HN :)
Cheers!
I work on an IAM system that is sub-ms p99 for our authz checks, with policies and keys pushed to each network edge instead of running a centralized system. The biggest perf hits are crypto verification and logging to the fs. We fail-closed to last known policy state when we have partitions, data loss would imply the application service or datastore proxy is lost. We measure policy deploy times in minutes though, and it’s eventually consistent.
However, if you're considering a multi-region setup, the latency will depend on the distance between the regions.That's why usually you define a preferred region (that stores primary copy of the records) or deploy in a geo-partitioned mode (when data is automatically pinned to configured regions).
That's correct. In a Zanzibar-like model, you have a global storage, but individual clusters in each datacenter/edge providing consistency-aware caching. This means p99 can be something like 25ms, but p95 or p50 is often FAR lower.
Disclosure: I'm a co-creator and maintainer of SpiceDB[0]
It’s been awhile, is the gist that Spanner’s coordinated clocks allow tighter consensus (i.e. faster writes) and caching provides read-my-write consistency?
Unfortunately, nothing is ever simple; comparing Spanner and CockroachDB is comparing apples to oranges. Two years ago, we wrote an article that details exactly how the differences matter in terms of a Zanzibar implementation[0], but I can give as short of a summary as possible: Spanner is linearizable and CockroachDB only guarantees external consistency for transactions that share rows. The post outlines how we workaround this and we've also more recently talked about how we've managed to scale that to 1M requests per second[1]. Our team focuses a lot on CockroachDB because we offer a permission systems that can span not only regions within a single cloud, but across various cloud providers. However, if you're all in on GCP, SpiceDB itself supports Cloud Spanner (which we also use in production for our GCP-only customers).
[0]: https://authzed.com/blog/prevent-newenemy-cockroachdb
[1]: https://authzed.com/blog/maximizing-cockroachdb-performance
[0]: https://fauna.com/blog/distributed-consistency-at-scale-span...
A push model is also valid if you’re heavy on policies and can accept eventual consistency. We will investigate how to generally push things to the edge (like we did with Ory Edge Sessions) or to cryptographic verification wherever staleness is acceptable.
By solving the primitives correctly in the beginning (with a multi region architecture) that job does become a lot easier, which is what we decided doing at Ory :)
Have you any documentation on your approach publicly available? I‘d love to get some education and insights from other large scale authz systems! We have a couple of ideas such as running a local replica in our customer’s stack but nothing concrete yet.
We have an additional scaling dimension though, as our permission model is richer and mutable by end users, therefore our policies are not uniform. For special hot-path services, we use symmetric keys to reduce latency further but that makes rotations complicated.
Also, very interesting project. I love SQLite and what the community is contributing to it - yours included!
The assumption that Ory does not offer support contracts for self-hosted Ory is wrong (although we did not in the past, when the team was smaller).
We are doing contracts for companies using our software self-hosted: See here and contact us if you are interested! https://www.ory.sh/support/ This way we can assign engineers to your case and work on any issues you encounter or work on any contributions or features required.
Ory releases all features for free for everyone to use. What is not free however is our time and work. To merge a PR/add a new feature/etc. a significant amount of time is needed to make sure the code lives up to standards, passes all tests, any security implications, etc. This depends on the feature of course, but the one you are alluding to is probably one of those. See the Code of Conduct on OSS support as well: https://github.com/ory/hydra/blob/master/CODE_OF_CONDUCT.md
I hope that makes it clearer and feel free to reach out to me directly in the Ory Community on github or slack.
Sometimes, PRs are not aligning with an architecture or API principle which is when they often go stale. This is why we generally require design documents for changes or additions to APIs.
Saying that the open source is second class is a false accusation in my view:
- Over 1500 PRs merged in Ory Kratos alone: https://github.com/ory/kratos/pulls
- Very active contributor and commit frequency: https://github.com/ory/kratos/graphs/contributors?from=2018-...
- A growing community and footprint
Also, we do offer support contracts for self hosted environments - this is relatively new though: https://www.ory.dev/support/
It is true though that have to balance open source work and things that people pay us for. It’s the only way to ensure that Ory open source, for which we have a deep commitment, continues for a long time.
Hope this makes sense!
A few issues on kratos that I consider relatively important are still missing / nobody from Ory is giving their input so it's hard to make progress and I would not take my time to contribute if I dont know if the owner are going to merge it.
An example that comes to mind is the OAuth email auto-verification or the search of users that is still super basic (we only recently got the filter of identifiers).
from the top of my head some of the features added in 2022: - verification and recovery codes - import of MD5-hashed passwords - integration with Ory Hydra - device information in session - session management APIs - session metadata - blocking webhooks - many improvements to OIDC mappers - session refresh - opentelemetry tracing - complete rewrite of docs - import identities including hashed passwords - custom email templates - passwordless with webauth - 1:1 compatibility Ory Network and Ory Open Source
Of course there was a huge amount of bugfixes and smaller improvements going on already. 2023 also already saw a ton of work being done on Ory Kratos including the 1.0 stable release. Of course there is still much to do, and feedback like yours also helps! If you are looking to contribute its always recommended to talk to the maintainer before you start coding - then we can let you know if its realistic to be merged or not. Search is a not a trivial thing to implement, on the one hand it is needed in some form, on the other hand Ory Kratos should not bloat too much.
Anyway, thanks for the feedback, will take it into consideration :-)
Deploying app instances across distant locations was never an issue. However, databases used to be the bottleneck. I'm glad to see that changing, thanks to CockroachDB and YugabyteDB.
My favorite multi-region deployment mode is geo-partitioned deployment. This is when a database automatically pins user data to specific locations, ensuring low latency for both reads and writes, regardless of user location. One-minute demo how it works: https://www.youtube.com/watch?v=9ESTXEa9QZY&list=PL8Z3vt4qJT...
Total side question, if anyone knows -- what tool (if any?) was used for the graphics in this article? The dot matrix looking map style stuff? I really dig it.
Our designers will love that feedback! Unfortunately it’s not a shelf product but they used Figma to design the graphs.
Even Google Cloud Spanner (NOT the same as Google Spanner - the internal DB) lacks a couple of things we needed for data homing.
We’re currently in a early stage, but more info will be available to public in the next couple of weeks. You can fill out the form, if you want to be in the known: https://tally.so/r/w2ajRb
Thanks again, appreciate your honest feedback.