EC2's most dangerous feature
daemonology.net
daemonology.net
Think about things like services that accept user submitted URLs, crawl them, and display results...
edit: not sure why the downvote, I use Google Cloud, so honest question :(
It's well documented and very useful feature in aws and it also exists in gcloud, although gcloud documentation is not as good^1.
It only becomes a vulnerability if you don't read the documentation AND don't follow best practices^2.
I don't know if azure provides this feature.
^1 I use gcloud now and have used aws in the past.
^2 Don't run unknown code without proper sandboxing
Like Colin pointed out in the blog post, this completely subverts the permissions model in modern operating systems.
Calling this a vulnerability is akin to calling the existence of `rm` a vulnerability because it can delete files.
I allow usage patterns similar to what is being described, so it is a vulnerability in something, be it my fault or not.
169.254.0.0/16 is link-local range which you should be flitering along with publicly routable ip ranges that might be very upset if you access them like .mil reserved ip ranges. Go as far to also only allow DNS names instead of arbitrary ip, keeping in mind dns names may resolve to non publicly routable ranges or ranges you may not wish to access. These are all standard dangers of making queries on a user's behalf.
Good list of ipv4 ranges you should not allow: https://github.com/robertdavidgraham/masscan/blob/master/dat...
Some limited instance info about your host ID, fault domain, and update domain: curl http://169.254.169.254/metadata/v1/InstanceInfo
Poll this regularly to find out when your VM is about to go down for maintenance: http://169.254.169.254/metadata/v1/maintenance
Those are the only two endpoints I know of. I encouraged them to require a special header for access to this API before releasing to the public but it looks like it was not included. A special header would help prevent your app from being able to access this from a user specified URL.
Google requires a header to help prevent this: https://cloud.google.com/compute/docs/storing-retrieving-met...
I try to limit the possibility of abuse (restricting protocol to http, https, //, data, or ftp) but didnt know about this metadata issue.
I updated my product to account for this issue too.
it looks like google's metadata doesn't leak any important secrets (unless I had custom metadata, which i do not) but better safe than sorry!
I've described it previously as "Kerberos for the AWS Cloud" (which will make any self-respecting crypto nerd squirm) but hopefully it conveys the general idea. Yes it was designed for mobile & browser use, and yes the API isn't pretty, but it's there.
Mind posting your entry for us iptables impared folks?
iptables -t filter -I OUTPUT -d 169.254.169.254 -m owner \! --uid-owner 0 -j REJECT --reject-with icmp-admin-prohibited
What they have reduces the duration of a vulnerability so that if you know someone had access to your machine at some point in time, you cna figure out from there how long their keys would have lasted and scope down the timeframe to start digging with in cloudtrail
He _is_ right in his first criticism that the IAM access controls available for much of the AWS API are entirely inadequate. In the case of EC2 in particular, it's all or nothing--either your credentials can call the TerminateInstances API or they can't. I'm sure Amazon is working on improving things, but for now it's pretty terrible. But in practice it just means you have to take care in different ways than you would if his tag-based authz solution were implemented.
That said, while it's certainly frustrating to an implementor, it's not "dangerous" that limitations exist in these APIs. We're talking about decade-old APIs from the earliest days of AWS, and while things have been added, the core APIs are still the same. That's an amazing success story. But like any piece of software, there are issues that experienced users learn how to work around.
You can bet that the EC2 API code is hard and scary to deal with for its maintainers. Adding a huge new permissions scheme is likely nearly impossible without a total rewrite... I don't envy them their task.
The biggest issue with the IAM instance profiles is that they trade security for convenience.. and it's not a good trade.
In a perfect world, any service would cleanly interoperate with any other service. Unfortunately we don't live in a perfect world. If you want to take advantage of `advanced` features in a given platform, you have to understand the drawbacks and limitations of those features, and what it means when they aren't available on another platform.
To me, the greatest tragedy in the way EC2 operates is that it looks/tastes/smells like a `server`, but it's far more akin to a process.
It is a full virtual server with its own Linux kernel and operating system: it has to be updated, secured, and maintained just like any other Linux server. Most Linux distributions on an EC2 instance have dozens of processes already running out of the box.
I understand your point -- that ideally a single instance can be treated as a single functional point from the point of view of the application, and I agree, but not from a point of view of security. As you know, in any larger environment, there are likely many additional support applications running on that server: things like app server monitoring, file integrity, logging, management, security checks, remote data access or local databases, etc. Those must not all be treated with the same levels of security and access. (i.e., why would rsyslog or systemd need access to all objects in our S3 bucket or be able to delete instances or any of the other rights that might legitimately be granted to an instance via an IAM instance role?)
To treat security for all of these processes as if they're all part of the same app tosses out decades of operating system development and security principles and places your single function app, as well as that of your entire environment, at grave risk. I.e., there's a reason why a typical Linux distribution has about 50 accounts right out of the box and everything doesn't just run as root.
If you are developing or deploying microservices or containers and don't want to be burdened by the security requirements, then there are alternatives at AWS like ECS and Lambda that you should seriously consider.
It really sounds like we have vastly different ideas about what kinds of processes belong in an EC2 instance, as well as the ideal life-cycle of an EC2 instance. I tend to adopt a strategy of relatively short-lived EC2 instances that get killed and replaced frequently. Persistence that depends on a single instance surviving is avoided at all costs, in favor of persistence distributed across a number of instances (or punted out to Dynamo/S3/RDS).
You're absolutely right that there is a reason why the typical Linux distro has 50 accounts out of the box -- it was built with traditional multi-user system security models in mind. I sure as hell appreciate it on workstations and traditional stateful hosts. That said, eschewing the traditional security model in favor of an alternative model does not make your environment inherently more or less safe -- there are going to be pros and cons to both approaches (in terms of both security and functionality).
> either your credentials can call the TerminateInstances API or they can't
Note that you can restrict the inputs to this API using IAM Policy in semantically meaningful ways. Three controls I'm familiar with that are useful for restricting inputs are the resource type for instances and the conditions for instance profile and resource tags [1]. The latter two are most flexible.
An instance profile restriction allows you to express a concept like, "This user may only terminate instances that are part of this specific instance profile"; in that way, the instance profile characterizes a collection of instances that can be affected by the policy. The resource tag condition can be used in a similar way. [2] is an example of a policy restricting terminations based on instance profile. The key fragment of it is:
"ArnEquals": {"ec2:InstanceProfile":
"arn:aws:iam::123456789012:instance-profile/example"}
A role with this policy condition can only affect instances that are part of the specified instance profile.This allows you to create roles or users that have access to instances that are part of a certain instance profile only. If you wanted a group of instances to be able to manage (e.g. terminate) themselves, then the role on those instances could be access-restricted to the instance profile of those same instances. By assigning different fleets of instances different instance profiles, you can control which users or roles can access each fleet by restricting access to the fleet's Instance Profile. Similar restrictions are possible with resource tags on instances.
That said, though, I agree that there's room to improve the access control story. Managing instances through their full lifecycle sometimes involves accessing other resources like EBS volumes too, and it's not easy to construct a policy container that sandboxes access to just the right resources and actions while allowing the creation of new resources. Colin called out some of the gaps in his post. If you do not need to allow the creation of new resources then the problem is a bit easier. For example, you can avoid the need to create new EBS volumes directly by specifying EBS root volumes as part of instance creation using BlockDeviceMapping.
[1] http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ec2-suppo... [2] https://gist.github.com/jcrites/d6826fc57b17c3c0ac50cae1fc9b...
The fact is that Amazon provides a commodity service and it's not a standard thing that most people expect to have an internal HTTP service that exposes potentially sensitive information to non-root users.
I actually disagree with the OP where he says they should use Xen Store for metadata. If I were Amazon, there is no way I would want to commit to using an option that is specific to one hypervisor technology. What if Amazon wants to switch to KVM?
1) IAM instance roles have no security mechanisms to protect them from being read from any process on the instance, thus completely eliminating them from all Linux/UNIX/Windows permission systems. (The real reason for this is that instance metadata was a convenient semi-public information store for things like instance ID, but it was extended to also provide secret material, which was, at best, an idiotic move.) As the author points out, Xen already provided a great filesystem alternative that could be mounted as another drive (or network drive) to be managed with the regular OS filesystem permission system. (reading an instance ID is just a matter of reading a "file")... for some reason, AWS didn't leverage this and instead just added the secret material to its local instance metadata webserver.
2) the API calls are not fine grained enough and/or there are big holes in their coverage -- so, for instance, if you want to use some other AWS services, you can end up exposing much more than you intended.
So yes, there are probably lots of services which get this wrong; but they did at least get a warning about that particular failure mode.
We redid all our policies to be extremely restrictive in response, if the instance did anything based on user input. Anything more admin happened on a different machine.
1. Resolve hostname and remember the response
2. Verify that the response does not contain any addresses in a private IP space, or any other IP that is only accessible to you
3. Use the IP from step 1 when establishing a connection
With other solutions, you might end up being vulnerable to DNS rebinding attacks.
Bonus points for doing all your URL fetching in some sort of sandbox that enforces these access rules.
Remember that even ELB's in AWS have IP's that change all the time, and this itself is actually a source of vulnerabilities from apps that don't respect DNS TTL's (as has been seen in the forums repeatedly -- apps get connected to the previous IP instead of the new one). It's probably safer to retrieve and verify the IP for each request, and just cache if the IP is 'safe'. (And just doing IP subnet calculations is non-trivial in most less-common languages.)
Also, request throttling should be maintained and HTTP verb checking, to prevent being turned into a proxy for other attacks.
Actually, any decision to accept an arbitrary URL should be carefully examined in light of how hard it is to do safely.
Pentesters have been using this trick to pivot from unexpected backend web proxies to (e.g.) management consoles, LOM servers, JBoss interfaces, &c for 15 years or so already.
There is actually an example of this in the IAM documentation [1], although the source VPC parameter doesn't work for all services, and I can't see a list of services that support this parameter. This would ensure that the requests actually came from instances within your VPC.
[0] http://docs.aws.amazon.com/IAM/latest/UserGuide/reference_po...
[1] http://docs.aws.amazon.com/AmazonS3/latest/dev/example-bucke...
Doesn't this become more complicated when you think about EC2 offering Windows instances? Even with straight UNIX file writing, what writes this? Where does it write this? Which user has read permissions?
In Windows, I'm guessing that this would be exposed as a network drive.
I still think they should just disable it by default, so you have to "opt-in" to the potential security risk and plan accordingly.
If you want roles to work for other users via the meta data store you can intercept requests with a proxy and then grab the temp credentials STS assume role. This is how kube2iam works. Depending on your use case you'd have to write the proxy, automate the mappings and firewall rules, etc etc etc. PITA but probably doable.
On a different note, I agree with just about everything you had to say about the PITA that is IAM. Properly scoping permissions is much harder than it needs to be. Not all resources support tags, and even then almost nothing outside of ec2 supports tag conditions in IAM. This leads to many naming schemas and wildcard resource conditions :( AWS should have created the concept of resource groups. This would have greatly simplified giving users permissions to subsets of an accounts resources. I chalk this up to AWS's VERY poor collaboration between service teams. Nothing that came out seemed coordinated. This appears to be getting better (shrug).
> Credential Isolation: A container can only retrieve credentials for the IAM role that is defined in the task definition to which it belongs; a container never has access to credentials that are intended for another container that belongs to another task.
If you do firewall the instance metadata service and want to get credentials into individual processes, then you could do that using one of the credential providers in the AWS SDK. I haven't worked with every language SDK, but service clients in the SDK for Java take an AWSCredentialsProvider as input, and you can pick from a number of standard implementations [2] or define a custom one.
> An admin user could fetch these subsets of the metadata and leave a copy of them in the local filesystem.
So if you wanted to take this approach, an admin agent could periodically copy the role credentials as property files into the home directories of users that need them, and then applications could load them by configuring the SDK with ProfileCredentialsProvider (which can refresh credentials periodically). The admin agent could perhaps be a shell script run by cron that `curl`s from the instance metadata service and writes the output to designated files.
[1] http://docs.aws.amazon.com/AmazonECS/latest/developerguide/t... [2] See ProfileCredentialsProvider and DefaultAWSCredentialsProviderChain
Single-purpose doesn't mean single-user. Lots of services divide their code into "privileged" and "unprivileged" components in order to reduce the impact of a vulnerability in the code which does not require privileges. As far as I'm aware, there's no way to have an sshd process which is divided between two EC2 Containers...
Use AWS be pushed towards an architecture based on containers and services. AWS is the OS, not any individual machine.
The instances hosting our users go a step further and null route Metadata service requests via iptables.
This does mean that everywhere else you'd have to have explicit service accounts and such, but that seems like a reasonable "workaround" until or unless we make metadata access more granular (I like the block device idea! Would you want entirely different paths for JSON versus "plain" though?)
Considering the amount of unpatched Docker containers out there, that's a bit scary. It also effectively prevents GKE from being usable in any scenario where you want to schedule containers on behalf of third-party actors (think PaaS). (GKE also doesn't let you disable privileged Docker containers, but that's another story.)
On AWS you can run a metadata proxy to prevent pods from getting the credentials, but I don't know of a clean way to accomplish the same thing on GKE.
It's a balance between security and convenience.
The access is controlled by source IP (and namespace). I wonder if it's possible to spoof the IP and access Metadata of other servers/users.
Not that I mind, but getting your HN submission in within 30 seconds of my blog post going up is very impressive if you're not a robot.
(I got better.)
And to your question, dwaxe is not a bot (there are comments associated with the account too), and this has happened before (apparently a lightning fast submitter):
https://news.ycombinator.com/item?id=12218797
Of course he/she could still be running a script.
I'll probably do that next time. I didn't think it would matter! (And really it doesn't -- it's not as if I need the karma points.)
It's the principle of the matter! Does HN allow bot submissions?
Compare these timestamps (hover over date): https://hn.algolia.com/?query=dwaxe%20eff&sort=byPopularity&... To the RSS feed: https://www.eff.org/rss/updates.xml
Humans don't post everything from an RSS feed within 10 minutes!
Presumably to farm karma for some other purpose.
I once observed that everything that shows up here eventually shows up in a particular cluster of subreddits and vice versa. Usually with a lag of 24-48 hours.
I half-jokingly floated the idea that one could write a karma arbitrageur bot which cross posts between HN and those subreddits.
I was told that I would be banned. Not just the bot: me.
Side question: I wonder when AI pass Hacker News Turing test. So bot can trick us into being human by HN comments.
Your domain seems to have been added recently [1]. Congratulations!
So whether you choose to blog about tarsnap or anything else, chances are that it'd be posted to HN before you're able to.
Cyborg. Story submission is clearly robotic, but there's a lot of charmingly humanesque entries under dwaxe's comments:
SELECT AVG(score) AS avg_score, COUNT(* ) AS num, REGEXP_EXTRACT(url, r'//([^/]*)/') AS domain FROM [fh-bigquery:hackernews.full_201510] WHERE score IS NOT NULL AND url <> '' GROUP BY domain HAVING num > 10 ORDER BY avg_score DESC
which returns a list of domains with more than ten submissions sorted by average score. This turns out to be a list of some of the most successful tech blogs on the internet, as well as various YCombinator related materials. Out of the domains with over 100 submissions, daemonology.net has the 9th highest average score per submission. I manually visited all the domains with more than about 30 submissions, found the appropriate xml feeds, and saved them. I added a few websites like eff.org whose messages I think everyone should read anyways.
Then I jumped into python and started trying to figure out how to post to Hacker News. It was a little more complicated than I anticipated [3], but an open source HN app for Android helped me figure it out.
I set up a cron job on my $5 Digital Ocean that runs the script every few minutes (pseudocode):
If you can reach http://news.ycombinator.com, Check all feeds for new entries, Post a new entry to hn, Sleep for an hour before posting another
[1] https://bigquery.cloud.google.com/dataset/bigquery-public-da...
[2] The only difference on Reddit is the subreddit system.
[3] After you send a POST request to send to the login screen, Hacker News gives you a url with a unique "fnid" parameter, and you send another POST request to another url with the appropriate "fnid".
There are many more reasons why this isn't a good thing for HN. For example, it's better for submissions from popular sites to be distributed across a wide range of accounts. That gives more users a chance to feel like they're making important contributions, and gives the community (and authors) a clearer sense of the audience.
There are lots of ways to write software to interact with HN, and lots of users with the ability to do it, so we really depend on the good will of the community only to do that when it serves the whole.