Authorization in a Microservices World
alexanderlolis.com
alexanderlolis.com
> "So the logical thing to do is to implement an authorization service and everybody would be able to use that and keep your precious service boundaries, right? WRONG. Your hell, has just begun!
Drawing the right boundary for authorization is near impossible. If I want to check whether the user is allowed to see who left an emoji reaction on a comment response to an issue inside a repository belonging to an organization -- do I store that entire hierarchy in my authorization service? Or do I leave some of that in the application? I've yet to see a good heuristic for the latter.
Also, thank you to the author for referencing us :) I'm Sam, Cofounder/CTO at Oso where we've been working on authorization through our open source library for the past couple of years. Authorization for microservices has been a recurrent theme over that time, and we've been furiously writing on the topic (e.g. [1], [2], [3]).
We'd be interested in talking to anyone who is currently facing the challenges of authorization for microservices (or really just multiple services). We're building authorization as a service, which you can learn more about here: https://www.osohq.com/oso-cloud. It's currently in private beta, but if you would rather _not_ speak with us first, there's also a free sandbox to try out the product linked from that page.
[1]: https://www.osohq.com/post/why-authorization-is-hard
[2]: https://www.osohq.com/post/microservices-authorization-patte...
[3]: https://www.osohq.com/academy/microservices-authorization
It seems like building a service that tries to abstract all that (the emoji is an "attribute" of the "action" to a "object" that has an "owner" that has "rules") is doomed to succumb to the inner platform effect. The only solution I can see is to build a general rules engine, but, in my experience, those tend to just be complicated, hard-to-read ways to implement logic that should just be code.
I think it's absolutely fair to say that authorization logic is a subset of application logic! The question is whether it's possible to separate any of the authorization logic from the application.
There are two reasons you might want to do that:
1. Separation of concerns (true whether monolith or microservices) 2. You have authorization logic shared across multiple services.
(1) is still a hard problem, but you can be a little less rigorous about it. (2) is where it gets really fun.
Continuing the example, that repository-specific blocked-users lists is _probably_ going to be needed across every other service.
I don't think any of that contradicts what you're saying, just clarifying that "in the application" might still mean you want to extract the shared logic into a service.
But to get to your last paragraph: yes it's definitely hard to build a service that's both sufficiently generic to handle all the kinds of authorization use cases people typically need to do.
You'd be surprised though at how many use cases fall into very similar patterns. We have some internal (soon to be external) documentation where we managed to get to about 20 distinct patterns -- from the common one like roles, and groups, to less common ones like impersonation and approval flows. And most of these share the same 2 or 3 distinct primitives (user is in a list of blocked users for a repository, emoji is in the disallowed-emoji list, etc.).
So you can build something less abstract than a general rules engine, but that still saves you the work of building from scratch for the nth time a roles system or a deny list.
One concrete example of completely separated auth is AWS. It would be quite the nightmare if each AWS service had its own authorization system. Centralizing that management in IAM makes it... manageable (and only barely at that).
Rather than a service acting as a shared upstream for multiple microservices, you might want to put this kind of generalized authorization into an API Gateway service that sits in front of multiple (internal or external) microservices.
Compare/contrast: blocking malicious origin IP addresses in a firewall appliance. But substitute "IP address" with "API key", and "firewall appliance" with "load balancer."
And also, to be clear, the services would still do domain-object policy-based authorization themselves. The point of such a multi-microservice API gateway is to optimize universal, pre-authentication, static-credential-based denials (e.g. blocking specific API keys, rather than blocking specific users) out of the critical path, such that users can't DoS your backend with 403-generating requests.
I look forward to seeing that pattern list. I'm in an adjacent space (an authentication server offering RBAC) and am constantly amazed at the intricacies of many orgs authentication needs.
I suspect that common patterns can take care of 95% of authorization needs, but I imagine there'll need to be an escape hatch for the 5% that are really super business case specific.
Each task in that system has an accompanying set of roles that are authorized to execute it, and essentially that authorization happens at the controller level and is "at the edge" of the application instead of in the business logic itself.
Is there a benefit I could gain from Oso? I will read what you've provided.
Where we normally see people considering Oso is when the model gets more complex, or the data requirements get larger. E.g. when you introduce new features like user groups, projects, or sharing, then the amount of logic + data you need to share between the services grows beyond what's sustainable.
If you're current system is working for you, I'm not going to try and tell you otherwise! If that changes though, let me know ;)
Is the request composer responsible for checking the authorization data? Like what roles/permissions the user has?
There’s a similar OSS implementation (OPA) targeting mainly k8s but allegedly useful generically that uses Datalog.
Or would the acquired claim be communicated towards the service in the request? Which begs the question, how does the service communicate which claim is required.
Not trying to be critical by the way, genuinely curious.
The concepts are simple but the implementation can be very difficult.
Authorization and Business Logic are two entirely separate domains. You have to start there. They are orthogonal.
From that follows that requirements involving who you are get implemented in the authorization logic/service/etc, and other requirements get divided up into the appropriate domain logic/services/etc.
If there are requirements around access to emojis then that involves the authorization service.
Sometimes data needs to get duplicated across service boundaries and that's when you need application concepts like sagas to manage this. That's where the implementation starts to get difficult.
I would absolutely love to learn more about this, I feel like I’m unable to conceive an appropriate solution to these requirements.
Doubt it will answer all your questions but I found it interesting.
The more I come across different systems, the more I realize authorization in large distributed systems isn't ever one approach: it has to be tailored for each use case with different tradeoffs in mind. It's often directly coupled to the problem domain you're trying to solve for. It has to be integrated with the data access pathways, _and_ also has to be tailored to the authentication system it deals with, _and_ it has to be tailored for the data-locality model of the overall system.
The more authorization problems I solve, the more I realize that my dream of coming up with a generalized authz SaaS service that helps me grease out VC money & a billion dollars probably doesn't really exist. It's different from authentication, because authentication has less dimensions of coupling, and less tradeoffs (Auth0 sold for 6b, Okta worth 18b, both authentication offerings)
Maybe I'll figure it out one day. Or maybe this is one of those problem domains that is only solved by an army of engineers
There are a bunch of companies who popped up in the last few years to solve this problem. We're one of them -- Oso (I'm the CTO).
It's definitely a fun/challenging problem to work on. So if building a generalized authz SaaS is the dream, come join us!
There are (at least) three of us startups on this thread that are trying to tackle this :) (disclaimer - I'm a co-founder of one of these - Aserto).
Good read, but I am not really a fan of that stance. The "separate schema per ms" pattern is an overhead which has to be justified and even paid for by customers (small 50k projects and enterprise grade projects of million $). First and foremost the db is a shared data structure which CAN be used for communication and in a one actor writes and 1..n actors read scenario (especially when all connecting services are controlled by the same team) is typically the simpler and better approach as a start. Schema evolution will also hit you with messaging and APIs which can be done properly in smaller databases as well.
I’d do it again, if I had the guts.
Thanks for explaining the problem space so well and referring to Cerbos[1] as contextual solution.
We had to build these systems in the past so many times and we were tired of reinventing the wheel. With Cerbos we aim to turn this problem space into configuration rather than code and enable enterprise-grade access management for any application.
I couldn't agree more that the question of how to get the data to the policy decision point is one of the most interesting and hardest challenges for this scenario.
You don't want to replicate the entire world in both your application's data store and your authorization system. But you also want to follow the principle of separation of concerns as much as practical.
There may be scenarios where your authorization system has to make calls to an external system, but that can cause availability and latency concerns. Ideally the decision engine has all the data it needs to make an authorization decision, without having to query another system.
For filtering scenarios that involve returning too much data to the client (which the client would need to filter locally), you could also imagine the authorization system doing a "partial evaluation" and return you an AST which you can walk and attach as "where clauses" to your query. Here's an interesting read on how to do this with the OPA decision engine. [1]
[0] https://www.aserto.com [1] https://blog.openpolicyagent.org/partial-evaluation-162750ea...
I've had success doing this.
Elasticsearch has a general purpose authorization system as well, based on the concept of "application privileges": https://www.elastic.co/guide/en/elasticsearch/reference/curr.... It's an interesting concept I never really considered at previous jobs, to piggy-back the authorization system of some infrastructure already in your stack. On one hand I can see some team members finding it to be a bit of a "smell" since it's pretty far from the original intended use case for something like Kafka or ES, but on the other hand it can free you up from having to build this type of thing yourself from scratch or libraries.
Seriously, you're claiming to "solve" one "problem" by building an entire new system which is inevitably massively more complex and difficult to manage than the "problem" you're trying to solve.
Then again, I don't work in enterprise-level environments but to be honest if this is what it's like, I'm glad I don't.
Everything is difficult with scale. We can’t expect every service owner to implement authz correctly, but if we can expose and build tools that help standardize and abstract as much as possible the difficulties of authz then service owners can focus their energy on other things.
Here's a blog post [0] about the challenges we faced when using OPA for application authorization.
[0] https://www.aserto.com/blog/the-challenges-of-using-opa-for-...
If two services is chatty, they should be merged. This can be on several levels, data being one level.
Authorization services should not be chatty with other services. And why would they since you ask for claims upon an authentication request. Yes, I meant authentication.
And I believe an authorization service is something that has been solved for a long time. It doesn't have other traits than any other services you must collaborate with, that migh be your mistake to look at it like so.
The authorization service will give you the claims for the user and, each service will dictate whether the claims are valid and sufficient within the service boundary.
Don't look at it ad a special kind of service, because it's really not.
Welp, I guess I'll just merge with stripe then, thanks for the tip!
If schema need to change you just update the view.
For accessing one resource that's fine, a cached call to any k/v store would work. But when you want to list every items for a resource in a table that contains millions of rows it starts to be more difficult.
Without a JOIN you either need to do a costly WHERE IN or filter after the fetch that can results in scanning everything for nothing.
Is there any blog post on that?
article answer is 'user role [..] attached to a JWT' but that only really applies if you control your distributed microservice system, if you need to scale to etherogeneus identities you need to get into the magic world of federated authorities
and that is where the pain really is.