IAM Is the Worst
matduggan.com
matduggan.com
Just some of the things that make it challenging:
1. There are permissions at various layers. If anything along the chain doesn't line up, permission denied.
2. You need deep understanding of each service's specific IAM setup. It's not enough to write a policy that will grant you read access to a DynamoDB table. Your application probably also needs to grant access to the GSI/LSI indices created.
3. Ancillary permission requirements are not obvious if you're not familiar with the details of how a service works. Want a Lambda function to have logs and traces? Make sure you have the relevant CW and X-Ray permissions set on it.
4. Permission related failures do not make the root cause immediately clear. Your S3 get operation may fail because you're missing permissions to the related KMS key. The usage of the ancillary KMS API calls here is not obvious unless you inspect the configuration details of the resource.
5. Secrets related permissions are especially tricky. To be able to read a cross-account secret, you need to grant the IAM identity permissions to get the secret value, grant the identity permissions to decrypt the associated KMS key used for the secret, grant the related account identities permissions to decrypt the key in the KMS resource policy, and grant the related account identities permissions to get the secret value in the secret resource policy. This is assuming there's no other things like SCPs and permissions boundaries mucking it up.
6. The out-of-the-box managed policies are too broad and will often have you granting much more permissions than you need if you use them.
Low level IaC tools like CloudFormation and Terraform suck for this. They leave too much complexity to the end user to get right. CDK does mitigate the issue somewhat with it's grantX methods, but even those are fairly limited and require you to write manual policy statements for many use-cases.
A lot of times in the “developer guides” AWS includes the correct policies as a role buried in the docs somewhere. But those guides are often not tailored to work with Terraform and the like so if you go the IaC route you need to figure them out, often by trial and error.
- I am shocked that you don't seem to find Deny By Default the best thing in the world... (looking at you Azure...)
> You need deep understanding of each service's specific IAM setup.
- Color me shocked...
> Ancillary permission requirements are not obvious if you're not familiar with the details of how a service works.
- Imagine...Having to understand how stuff works to be gainfully employed....
> Permission related failures do not make the root cause immediately clear.
Cloudtrail is your friend...
> Secrets related permissions are especially tricky.
- Define the complaint....
> The out-of-the-box managed policies are too broad and will often have you granting much more permissions than you need if you use them.
At least for AWS, you are not supposed at any point in time to use out-of-the-box managed policies. Instead, you should use them as templates for your own policies or create your own Customer Managed Policies from scratch.
"...Another best practice is to create a customer managed IAM policy that you can assign to users. Customer managed policies are standalone identity-based policies that you create and which you can attach to multiple users, groups, or roles in your AWS account. Such a policy restricts users to performing only the AWS Private CA actions that you specify..." - https://docs.aws.amazon.com/privateca/latest/userguide/auth-...
If most people find it to be difficult to use correctly, it is difficult to use correctly. Maybe that’s the best we can do but it’s still bad.
The problem is not deny by default, but the complexity of setting "allow just the things I need". This is not easy.
> Cloudtrail is your friend...
Having to dig into the data of another service (that hopefully your org permissions allow you to read) instead of just being able to see a clear error message is not great DX. There are maybe valid security or performance reasons for not returning clear error messages, but there is a trade-off to usability made here.
> At least for AWS, you are not supposed at any point in time to use out-of-the-box managed policies. Instead, you should use them as templates for your own policies or create your own Customer Managed Policies from scratch.
Right, but because they're so broad, the templates themselves are overly broad. Even just using them as a reference, it's difficult to pare down to just what you need. You will inevitably go too far and have to play around with combinations until you identify the real need.
---
The rest of your comments essentially boil down to saying "skill issue"/"git gud". I think that downplays just how hard these things are to get right. I worked at AWS for almost 8 years and have used it for several more years as a customer since then. I still wind up with runtime errors due to permissions issues that I need to debug. I still find myself needing to spend lots of time shuffling through official docs and blog posts people have written about how to setup specific combinations of AWS services. I've seen other engineers within AWS struggle with this. I've spoken with many founders at startups who've struggled with this. The biggest challenge comes up when first learning and getting acquainted with a service. You don't even know what you don't know and there are many hurdles that can pop up along the way.
I mentioned it in my last comment, but CDK is probably the single biggest improvement to DX in the space here.
The comment about out-of-the-box policies is true, I suppose, but hard to take seriously. Almost every policy example you encounter in the AWS documentation is insecure by default. They've gotten better over time noting this and pointing to better examples for different use cases. But it's still pretty bad.
Sure, we could just put everything on one VLAN and hand out . credentials, but you can do the equivalent in the cloud too.
One to add to the list is that IAM conditions[0] are extremely powerful but there's no good way to know which conditions to use in which scenario and troubleshooting is very difficult.
For instance if you look at the EC2 CreateNetworkInterface action[1] you'll see that there are three possible resources (network-interface (required), security-group (not required), subnet (required)) and each of those resources have several possible condition keys associated.
What's not obvious is which condition keys will be available in any given request. I've run the same CreateNetworkInterface request with the same parameters and IAM role twice in a row and by looking in the "encoded authorization message" that was returned with the failure in each case I found that in one case the resource was a security group while in the other case it was a subnet. Depending on the resource type different condition keys are available in the context. So if you want to allow CreateNetworkInterface but only if the ec2:SecurityGroupID is 'abc' it might or might not work.
An extra challenge is the encoded authorization message is truncated in CloudTrail so if you're using CloudFormation you don't actually get to see what the context was if a call fails. Then you have to find a way to make the same call CloudFormation made using an SDK so you can get the full text of the encoded authorization message.
There's no easy way to just say "try this API call with this role and tell me exactly what the context would be and what part of the IAM policy hits it if any"
0: https://docs.aws.amazon.com/IAM/latest/UserGuide/reference_p... 1: https://docs.aws.amazon.com/service-authorization/latest/ref...
For example, I recently needed to allow some EC2 instances to push a private IP around between those. I would have assumed I can create some policy along the lines of "Yeah, VMs with this role can push 10.20.30.40 around between their network interfaces". I haven't been able to find any way to restrict these IP addresses, so now I have the smallest policy I could create: "This role can assign fuck-any internal IPs to these interfaces, let's hope for the best." Doesn't really feel the greatest.
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Action": [
"ec2:AssignPrivateIpAddresses",
"ec2:UnassignPrivateIpAddresses",
"ec2:AttachNetworkInterface",
"ec2:DetachNetworkInterface"
],
"Resource": "*"
}
]
}A start could be this even if does not address your scenario of calling twice in a row. I will discuss that one further below in the comment.
aws iam simulate-principal-policy --policy-source-arn arn:aws:iam::ACCOUNT:user/Paul --action-names "ec2:CreateNetworkInterface" --context-entries ContextKeyName="ec2:Subnet",ContextKeyValues="subnet-12345678",ContextKeyType=string --resource-arns "arn:aws:ec2:REGION:ACCOUNT:subnet/subnet-12345678" > simulatedIAMOutput.json
> I've run the same CreateNetworkInterface request with the same parameters and IAM role twice in a row and by looking in the "encoded authorization message" that was returned with the failure in each case I found that in one case the resource was a security group while in the other case it was a subnet.
Well EC2 would process these requests by first verifying subnet-related permissions before moving on to security group permissions. Variations in the error messages could reflect the point at which the request encounters a permission issue?
Kidding around though I'll try that if I face a similar issue in the future. It has been improving quite a bit lately.
> Well EC2 would process these requests by first verifying subnet-related permissions before moving on to security group permissions. Variations in the error messages could reflect the point at which the request encounters a permission issue?
I would think the context would be deterministic in that case but I verified calling the API with the same parameters using the same role twice in a row ended up with different 'resource' values in the context. It was almost like under the hood boto3 or something else was changing the order of the parameters in the API call which was changing the way the context was created. I could've put in a support case but had bigger fish to fry.
Some security group at a company will have a "review" of your permissions. Occasionally they will run a "sweep" and yank permissions out from under you.
Instead, here we have an API that the security team can actively manage and PROVIDE A SOLUTION. Should a developer on some project have domain knowledge of IAM to make a perfect bespoke (and it WILL be bespoke) least-permission policy?
No, of course no. that domain knowledge should be a service in any substantive AWS org where they provide it to you, and much more importantly, DEBUG it for you when it doesn't work.
Because here's the deal: IAM may be a bit ugly and have some cruft and evolution, and I believe S3 permissions are another entire headache atop IAM, but this is what an extremely fine grained permissions model looks like: detail hell.
ALL detailed permissions models will look like this. Defining perfect names (by definition coarse grained) to communicate the precise multidimensional n-brane border of a policy is basically impossible.
Here is another issue: in my last job they were obsessed with short-duration tokens and TOTP. Ok great. Hey wait, if I need to run an automated cluster-wide job that will take hours (backup, cleanup, log analysis, etc), what do I do then?
Security team didn't care. Automation? What's that? Just sit there watching the log and manually refresh the keys.
So I end up using a software TOTP generator and hacking it that way. I should not be doing that. It is likely a security hole. The security team should have heard my requirements, accepted them as a necessity (they are) and provided me a solution.
Security should be a solution and service.
The problem with IAM isn't the functionality it offers; it's almost exactly what you want. I mean look at what it actually does, isn't it literally exactly what you would do? Sure, maybe the terms are confusing or something, but on the whole... it's hard to argue that what it offers isn't basically what you want. You want to allow identities, to do things. You group those things into roles. And so on.
That said, the more powerful and granular permissions and ACLs get, the more grand the architecture you have to craft to make good use of it, and I think at some point, your own IAM rules become a work of engineering themselves. You wind up having to specifically engineer around and for IAM. This is not unique to cloud IAM; when doing complex NixOS setups, I have occasionally realized that my webs of plumbing secrets through SOPS to systemd units, and setting up group permissions for UNIX domain sockets between services, winds up getting quite complex quite quickly, and that it is basically, yes, engineering of its own sort. And if you add cgroups and network namespaces and nftables rules and seccomp, god help you, it's even worse than IAM! And that's just on a single machine...
Typing that out it's really the massive gulf between the abstraction "User wants to invalidate a cache" vs the implementation of 87 granular grants with obscure nomenclature.
Both that solution and the one in the article miss one point though: If you use the AWS Console at all it makes hundreds of calls to all manners of AWS service in all available regions. Because of this you can't just assume the calls made by a role intended for interactive use over some period are the "correct" privileges for that role because someone just clicking around in the console will generate thousands API calls to many different services.
IAM is not that complicated. The example diagram for AWS is showing all the features available. You can use only what you require.
Do two things from the beginning:
1) Least Privilege - Only the permissions required: You need to do that, because due to the constant zero days around the different software vendors, you are statistically guaranteed to have a security event. Establishing security boundaries will allow to contain the damage.
In combination with....
2) Temporary Privileges - Assume the roles with the permissions you require only when you required them: This will make it harder on an attacker, since they will need to compromise you while you are holding those elevated privileges. When you finish your task, either manually or programmatically, detach from the role or assume a different one with lower or different permissions.
And if you still really want to do that, you don't need AWS roles as a separate concept for this. You can just use temporary membership in groups.
AWS IAM model is overcomplicated compared to GCP/Azure. It becomes clear when you try to migrate the project with multiple envs (projects in GCP aka accounts in AWS) with cross-env accesses.
I agree. I think the benefit of this is quite low. If someone takes over your machine, things are lost anyways. If they take over your machine but for some reason cannot access your password manager (or so) or your 2fa to increase priviledges, you at least gain some time before the attack happens since the attacker has to wait - but it's not a major win.
The only case where this really helps is if the attacker gains only temporary access to you (maybe a temporary vulnerability in the browser) but can't "persist" it. In that case, you can reduce the blast radius.
https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_us...
It's different if you use a different machine for the priviledged account. Then an attacker has to take over that second machine too. IMHO this is a mucher better concept, but also increases friction significantly.
What you say (that people should only have the permissions that they need) is orthogonal to that.
A realistic attack that compromises a low privilege session may not be able to leverage that into higher privileges. Therefore, limiting the use of high privilege to a smaller window of time definitely reduces the attack surface.
I think you’re talking about something different. AWS session tokens let you use your SSO to request session tokens that have a short expiry. So you can do API/console actions but if an attacker takes the creds, they expire. It also lets you generate session tokens that only have the subset of your allowed perms that you need for that workflow.
But also it’s a complex and very important problem space.
I had an exhausting discussion on Reddit about why storing UTC is not always sufficient from commenters who continually proved they hadn't read the article or the rest of the the comments.
You don't define the privileges you need, by running around with full privileges. The article is pitching some kind of tool the author developed, but you can do the same at least for AWS, for a long time. And if you don't know what policies you need you should talk to either the vendor or the application creator, as you will never be able to exercise all the compute paths...
"Generate policies based on access activity" - https://docs.aws.amazon.com/IAM/latest/UserGuide/access_poli...
And the lack of roles of roles is beyond idiotic, it’s downright dangerous.
GCP's lack of "AWS role" concept is great and straigtforward. As well as its lack of both identity-bound policies and resource-bound policies at the same time.
The rest of those are either manual or too broad.
Instead I just remove it automatically - I now automate GCP's advice. https://github.com/james-ransom/auto-apply-gcp-iam-recommend...
Not saying your point is not true, I met guys who did it just because too. But it's not always malice or incompetence on their part.
Hopefully not the same ‘thing’ being targeted of course, because then it will get really bad, but yes conflict is inevitable.
This meant that the janitors had the same security responsibilities as the CFO.
It was ... invigorating.
The consultants that recommended it made a lot of money, though, so I guess it's all good.
This is a bit of a shameless plug, but I hope since it's an open source project it's okay. I'm working on a suite of tools called Otterize (otter and authorize, get it, haha :) that automates workload IAM for Kubernetes workloads.
You label your Pods to get an AWS/GCP/Azure role created, and in a Kubernetes resource specify the access you need, and everything else is done by the Otterize Kubernetes operators so that your pod works.
It's a lot simpler than all the kungfu you normally have to do, but it's not magic, honestly, it's just the result of limiting scope and having an opinionated view of what the development workflow should look like. Basically, instead of maximizing on capabilities, it trades some capabilities to maximize on developer comfort.
Check it out if you're keen on contributing, or just think IAM has a tendency to devolve into a mess ridden with politics.
github.com/otterize/intents-operator and docs.otterize.com
Not that I blame Amazon. I think they’re a victim of their own success in this regard and it was a solution that was devised ad hoc reactively as they ran into authorization problems rather than something that was architected top down. When you do that you always end up with a mess, but they may not have had a choice.
But boy as soon as they started adding IAM they took all the fun out of deploying my personal shit to any cloud.
The more complicated stuff starts happening when you have many services that need to access each other's resources directly. Then you need to think a bit more about the architecture and how you expose resource names, managed policies and roles between services. It's no longer just simple role definitions within the CloudFormation stack, but you have to pass around account identifiers, regions and resource ARNs in the configuration.
They should merge them together into one big happy standard, and call it Worldwide Enterprise Access Rights Environment (WEARE).
Maybe if they signed up enough celebrities to sing a song about it, it would make the world a better place...
If your devs are thinking about IAM roles/permissions/whatever, your security dept failed.
what are you using?
Plus the documentation is out of date in so many places, describing actions to take that have long since changed.
It's ok for single users (simple single-user use cases) and large organisations that can afford to allocate the time and effort to administer the thing; anywhere in between is a just-say-no.
(ranting a bit because I just lost the whole morning adding a policy to allow a single user to do a limited set of things).
Example, even with something as abstruse as SELlinux, you can literally just attempt whatever you're trying to do, and then pipe in the failure log into `fail2allow` and the permissions will be set to least privilege automatically and then it just works in most situations.
You can also pipe in the failure log to the policy advisor and it spits out a bunch of advice in plain text as well as a copy / paste command to implement it.
IAM is horrible if it can take design notes from SELinux.
I only associate other permissions through AWS instance roles, because I try to not give out specific IAM roles or keys to users or developers AT ALL -- only to the instances they're logging into.
Obviously this probably can't work for all companies and is dependent on how you have your environments set up, but we even run development on EC2 instances, and with Userify we get a color coded view for who can log into which instances, and then those instances already have the correct custom role.
And there, no IAM to individuals at all.
GCP sucks so bad as a product, that the only way to tell what IAM policies apply to your service account, is to run some kind of analysis query thing exported to a BigTable (which will cost you money).
You'd think you could just go into the console and click on the service account and it'd show you which policies are linked to roles are linked to your service account? That would make sense, and be convenient. But this is Google we're talking about. Engineering principles will always trump customer experience.
It's much worse than that of course. The default roles give too many permissions, for nearly anything you want to do. Often you are limited by what you can control, to only at an Org level, or Folder, or Project. Yet making a custom role is often difficult, leaving you to usually just slap on the default roles, making your resources insecure. Much of the time, a user must have an Admin-level permission over all VMs in order to SSH into them with GCP creds. Kind of defeating the purpose of having IAM to begin with.
I think the only reason we haven't heard of more GCP accounts getting compromised due to the shitty default policies is, thankfully, GCP has few customers.
You create a role called "cleanup user servers", which requires a scope of "read:users" and "admin:server"³ (i.e. a role is bundle of scopes), and then define the services that role has, such as "garbage-collect".
You then generate a new service that names "garbage-collect", and associate an API key with it, and boom, that API key has the permissions to read user metrics, and administer the servers and is tied to a specific service.
Want to add another service to that role (e.g. "berate-users")? Just add it as a service in that role definition, and you can use the exact same API key.
1: https://jupyterhub.readthedocs.io/en/stable/rbac/scopes.html
2: https://jupyterhub.readthedocs.io/en/stable/rbac/roles.html
3: or you can just shoot for the moon with the "admin" scope which encompasses everything.
It ingests log data and permission/role assignments then reconciles them to then allow you to create custom roles that only have the permissions that someone actually uses.
MS bought them and calls it permissions management: https://learn.microsoft.com/en-us/entra/permissions-manageme...
Palo Alto has a tool called PRISMA. https://www.paloaltonetworks.com/prisma/cloud/cloud-infrastr...
The sector of tools is called CIEM. Cloud Infrastructure Entitlement Management.
Here's the thing though...PA and MS charge PER MANAGED RESOURCE. It's crazy. This is something that should be core capability, but its an added charge. Its a space that is screaming for open source tooling to make it less rent-seek-ish.
Cloud security, especially from AWS, is (as described elsewhere in the comments) byzantine, but I've always felt that the real underlying problem is generic advice that only serves to protect the backside of the cloud provider. What's really needed are crystal clear patterns that cover >95% of actual use cases.
In one mode, it proxies all your app's API calls and tells you exactly what permissions you need.
- Use AWS Organizations to organize your teams into Organizational Units
- use SCP to limit permissions of the OUs.
- let the OUs create new aws accounts for every project/workload
- now you have permissions and costs organized per project/workload
Don’t be afraid to create many AWS accounts, this is encouraged and considered best practice.
This is exactly what AWS Organisations do, available for many years already. This now comes with IAM Identity Center that gives you SSO into multiple AWS accounts. So the setup I use is one management account running IAM Identity Server with users and groups. Then each product gets one or few AWS accounts that they own. Super simple and effective for small organizations. For larger orgs you would need to also use SCP and maybe AWS Control Tower or similar.
This way, your database guys can access one big folder with all the database stuff, your server guys can access server stuff and your frontend guys can deploy to an S3 bucket, but not much more.
This is the level of granularity you need in the real world. i don't believe that any organisation, no matter how big or sophisticated has an employee that can have roles/datastore.backupsAdmin, but not roles/datastore.backupSchedulesAdmin.
Believe it. Based on past experience as CTO and head of Security Engineering at one of the biggest orgs, this split is used and necessary, unless you want to inject yet another approval loop somewhere.
The first one lets someone get, list, or delete the backups, the second one lets someone make backups happen or not happen. I can absolutely see forcing regular backups to happen (a regulatory requirement) being a different person than whoever is using the backups, even different from the admin who can delete those backups.
(Delete means the backup admin can make it as if a backup didn't happen by deleting, but that's not what the compliance regulation covers, it has to happen in the first place, which is what the scheduleadmin covers.)
Article would communicate better if it started out with the idea (which is cool) rather than a long, detailed, easily misinterpreted rant
Store secrets directly with the client so attacks only compromise the data of one user and not all of them. Want if they lose the key or what if multiple groups need access? Shamir’s secret sharing. What if we might not trust some of the k of n group members? Require interaction with a public ledger that provably logs secret access as a part of the secret sharing scheme. What about machine learning on a massive amount of user data? Well, homeomorphic encryption isn’t quite there yet, but how much sensitive info do you really need for your training data?
We’re not going to eliminate security flaws in systems without provably correct programs by default (which is probably never going to happen), and even with that, you have the whole social element of security, which means you still need to design the system in a way that limits the damage one or a few people can do. Which requires a different type of identity management than IAM.
But that's true for anything with granular permissions.
Answering "what permissions are needed to perform X" is hard in any context. In AWS it can be overwhelming because AWS is big and interconnected.
When data isn't being lost to the void, undo-ability grows. And, having perfect undo-ability is genuine "bulletproof" security. Security, in the traditional practice, then becomes needless undo prevention. That's a lot simpler to tackle than disaster averting prevention.
You want any kind of overly thin privileges, for whatever reason ? You can too.
IAM has no complexity per se: the only complexity come from what you want to do.
Saying that "IAM is bad because we can do lots of things" is like blaming a programming language because you can implement whatever you want with it.
Things like Prisma (Palo Alto) and Permissions Management (MS) are the tooling in this space.
While we're at it, publish these IAM engines as public, open-source projects so that we can write proper unit tests. I don't want to call a policy simulator API.