Amazon EC2 Enhances Defense in Depth with Default IMDSv2
infoq.com
infoq.com
The problem is that many companies have very rigid IT departments so even though in theory it can be done somewhat easily, in some places it can be a months-long process to actually get it done.
I have a government customer running 150 web apps in about 20 vCPUs. If you naively multiply 150 apps by the minimum scale for HA, non-prod, and the 2 vCPU minimum for cloud hosts you end up with 450-600 vCPUs! That’s expensive even for governments.
While it's technically possible to run something like an old ASP.NET Web Forms app in a container, it'll be painful.
You should have put that in your first comment, I've removed my downvote!
Az functions might not be the most optimal in terms of per request cost, but for the volume of traffic we are pushing it doesn't scale into the equation.
Operationally, they're basically magic. TLS, auth, etc. all handled. GitHub deploys with zero drama. Whatever margin we sacrifice at the request billing is more than compensated by how well our engineers are sleeping at night.
B. Enterprises run a ton of commercial off the shelf software hosted on VMs
C. Ever heard of Windows? It’s still big in the enterprise
There's also shared vs dedicated and burstable options if you need to go lower. Which clouds don't (in practice) offer 1 vCPU?
GovCloud may be different for security purposes, not sure.
They do, it's called a container or a function, which have their own IAM/auth/role.
Google invented Kubernetes, Apple runs Kubernetes and previously was heavily into Mesos, likewise for Netflix. And I've seen a lot of Engineering tech blogs and worked on AWS for a decade and never have I ever seen anyone spinning up thousands of micro-VMs in that way.
Not only are there all sorts of account limits you would immediately hit and have to negotiate but the networking would be problematic to say the least.
Meanwhile you can spin up ECS, EKS, Fargate which allow you to run thousands of containerised apps without any headaches at all. Plus they support you having unique IAM roles for each container.
It's functional, but I think it's not as polished as the rest of Kubernetes which is why Kubernetes has a multi tenancy SIG that spawned the hierarchical namespace controller (https://github.com/kubernetes-sigs/hierarchical-namespaces) and virtual clusters (https://github.com/kubernetes-sigs/cluster-api-provider-nest...)
"Multi-tenancy in Kubernetes: implementation and optimization" - https://www.cncf.io/blog/2022/11/09/multi-tenancy-in-kuberne...
"...As we can see from the above discussion, the sharing of Kubernetes clusters is not an easy task; multi-tenancy is not a built-in feature of Kubernetes and can only be implemented with other projects’ support in tenant isolation on both control and data planes. This has resulted in considerable learning and adaptation costs for the entire solution. As a result, we are seeing an increasing number of users adopt multi-cluster solutions instead..."
While ECS/Fargate is easier, it’s still not headache free and you’re not taking into account how many large enterprises run lots of commercial off the shelf software on VMs.
I am confused why you just could not use the AWS_WEB_IDENTITY_TOKEN_FILE. You could just produce that file for every user, refresh it before expiration and set the environment variable for each user. I believe that should work with the default SDK just fine, no?
I mean, sure, it is not exactly zero effort, but still seems fairly doable...
Bit of a head-scratcher. If you run multiple services on a VM, why not issue each of them identities of their own, rather than relying on the IMDS identity? If you need separate identities on the same machine for your security posture, then the complexity of setting that up is largely handled for you by IRSA: https://docs.aws.amazon.com/eks/latest/userguide/iam-roles-f...
Even the logging service should not be using the IMDS identity. The IMDS identity should be very much locked-down, something like AmazonSSMManagedInstanceCore if I recall correctly.
https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_pr...
Because the easy ways of doing this are not nearly as secure as the short-lived credentials provided by the metadata services in these clouds.
With 5-min credential expiry you can know that if a developer logs into the VM (for debugging or troubleshooting or whatever) after 5min they can no longer invisibily impersonate that role. This makes audit logging and similar much more powerful.
I've frequently seen long-lived credentials being used as an alternative and this way anyone who have ever logged into any VM with that key can invisibly impersonate that service forever.
You can do this by writing your own credentials manager but this is much more work, harder to make as reliable as the one provided by the IAAP provider and you still need a root of trust somewhere. Relying on the metadata service moves the root of trust to something ephemeral and managed which is highly valuable.
We've looked a few times at different ways that IMDS could vend different credentials to different user-ids. We documented how to use local firewall rules, which you've linked to. This gives single uid restrictions, similar to filterd. We also have a BPF based tracer tool, https://github.com/aws/aws-imds-packet-analyzer which can monitor which user ids and processes are calling IMDS (and which version they are using).
Our next thought was to expose IMDS as a filesystem. That way ordinary POSIX filesystem permissions could be used to control which user ids could read which credentials, and it would even work on Windows. But our research found that security issues (in applications and libraries that customers run) that grant local file reading privileges are even more common than SSRF issues (in part this because many SSRF issues allow "file://" urls, so they become a subset).
We've looked at interfacing with the kernel keyring and the TPM 2.0 interface (and we now have Nitro TPM) ... but both are difficult to call to user space. The latter doesn't get user ID granularity, and the former is hard to coordinate with from the outside for revoking and rotating credentials.
Absolutely, long-lived credentials are worse. The only advantage to long-lived/file-sourced creds right now is granularity. Being able to lock them down to specific processes/UIDs and have a role per service. I guess what we want is the best of both worlds - short-lived per-service creds.
It appears to be a difficult problem to solve. I was wondering if something could be done with kernel modules.
Are you still working on this? Would really like to hear about future developments.
This solves so much problems: security is trival (set file permissions), audit is trivial, no more timeouts, deciding whether to expose credentials to container or not is trivial (using directory sharing...), the client code is dramatically simplified.
[1] https://www.cyberbit.com/uncategorized/aws-imds-v2-secures-s...
Over years, they moved from alerting to changing the default and will later remove IMDSv1 for new instances. Even then, they will not remove it for the previous generations and those may stay for years, the current oldest is M4 which I think is from 2015.
Prior to this change, IMDSv2 could be enforced on an account, machine image, and launch API call level.
If you're interested in this topic, you might enjoy reading https://steve-yegge.medium.com/dear-google-cloud-your-deprec...
now they are changing the default modes for instances launched from aws console using quick start process. and sometime in 2024 they will give an option to customers to control the default value for all run instances API calls from an account
for example if you run below curl request on ec2 instances launched from aws console before nov 6th then you will get a response
```sh
curl http://169.254.169.254/latest/meta-data/profile
```
but if you run the same command on ec2 instances launched from aws console after nov 6th, then you will get an error that auth token is missing
https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/instance... has a good example.
The 2019 blog post[1] has fairly good explanations on how it works and the threat model.
[1] https://aws.amazon.com/blogs/security/defense-in-depth-open-...
tl;dr - misconfigured reverse proxies allowed cloud metadata URL access across the bigger cloud providers.
[1] https://github.com/opsdisk/cloud_metadata_extractor/blob/mas...