AWS Nuke – delete all resources associated with AWS account
github.com
github.com
Found this out after using AWS nuke
You're not the only one to complain about the documentation though, so it's easy to see how aws-nuke apparently missed it!
It is easy to screw it up.
Though 'not at all documented' makes that harder of course. I haven't used it.
If you male a mistake setting up the Batch ComputeEnvironment using cloidformation and it rollbacks then the cloidformation error message is useless (resource failed to stabilise) and no trace is left behind for you to check.
Two, Batch needs to use servicelinkedroles. If you create an AWS thing that needs a servicelinkedrole via the console it gets creates for you automatically. If you created via CLI or Cloidformation it does not. So if someone has created a batch env in your account before it will probably be there but if you are setting up in a new account you might have no clue whu things arent working.
Compute Environments and JobSpecs that use Fargate are configured differenrly, requiring a extra IAM role, from EC2 batch envs.
Edit : I've ran into this but it's considered a feature not a bug!
If you detach an EC2 instance from a virtual tape library and destroy it you can't delete the tape library any more. Even AWS support couldn't delete it. This is fine until you have 60TB of tapes online and are paying for it.
Fortunately we found an EBS snapshot of the VM.
You can also try updating the cluster to have a new role.
Remember boys and girls: Being "cloud scale" means data corruption and referential integrity violation!
If the role is deleted, it can’t do this.
In fact, being able to delete a role without removing any associated resources is a feature, not a bug. And how would you even ensure referential integrity in this case - you would achieve the same effect by modifying the assume role policy but keeping the role around.
You could craft a policy that only allows Batch to assume the role on a Tuesday for example.
Because it didn’t handle the fact that the role it’s using might be deleted or otherwise rendered un-assumable for a variety of different reasons at any point in time.
Which is a feature. Not a bug.
It’s a feature because there are plenty of cases, such as a role being compromised, where you don’t want to cascade-delete every single associated resource without any need.
If you want this then you can opt into it by using cloudformation.
Asserting this doesn't actually make your argument for you.
Why would this be a problem?
So I think it’s up to you to explain how this would work when there is no difference between a role being deleted and a role being inaccessible.
What exactly would you do if I block instance with the ID “xyz” from assuming the role I have assigned it? How would you detect that I’ve done this in every single case?
Again you keep just asserting this isn't possible: why? AWS are aware you need the role to exist to delete the instance (as noted up in the OP), why apparently is it completely inconceivable that the IAM system would check for this condition before executing an action - which again - irreversibly wedges a delete operation for a resource?
Unless you have some deep knowledge of how AWS IAM is implemented which makes this literally impossible, then you're asserting fluff.
In short: it would be a confusing mess for nebulous gains that doesn’t pass any kind of smell test. Instead they should just… fix the AWS batch service.
The better solution is to provide an API and console tab to show you what services last used the role, when they used it and how they used it. Which is what they do.
Not my idea of a good time, but whatever floats your boat. Hopefully you had a backup?
1. Create resource. Success.
2. Attempt to use/reference first resource in another resource/call: Failure: Referenced resource does not exist. Odd.
3. Create resource again? Failure! Resource already exists.
Our scripts have so many retry loops and arbitrary pauses mixed in to account for garbage like this. Distinguishing "did the call fail?" from "or is the system just lost in the land of eventual inconsistency?" ugh.
And yeah. shout out to Azure AAD where I can have a role assignment that is granting permissions to "unknown". We call them ghosts.
This is unfortunately your fault, not Azure's. Their API is explicitly designed around this eventual consistency and weak references, so your client side must take this into account. Typical bash, Python, or PowerShell scripts are the Wrong approach with a capital W, and you will forever be tearing your hair out if you persist on using them. (Or any similar imperative deployment mechanism)
The only robust method is ARM Templates, or better yet, Bicep templates[1]. The latter simply compile down to ARM, so they're essentially equivalent but terser and with nicer tab-complete.
Compared to scripts, templates have key advantages:
1. Built-in incremental / differential deploy capability. A partially deployed template can be simply redeployed[2] to "fix it up", without requiring client-side logic for every corner case.
2. Can deploy multiple changes that would fail if deployed step-by-step. For example, App Gateway can have intermediate configurations that won't validate on the way to a valid final configuration. This is madness to unravel with scripts. Templates generally just take you to the final configuration in one step.
3. Inherently parallel. Anything that can be deployed concurrently will be. Anything. No need to write complex and error-prone parallel loops on the client side!
4. Largely immune to temporary failures like missing reads after writes.[3] The template engine has a built-in retry loop for most (all?) fallible steps. You'll see it has "failed"... and then "succeeded" anyway.
[1] https://docs.microsoft.com/en-us/azure/azure-resource-manage...
[2] Most of the time. All resources should be idempotent to redeployment, but many aren't because this is not mechanically enforced. IMHO, this is just shoddy, shoddy engineering and everyone involved should be ashamed. Being nearly idempotent is like being nearly pregnant.
[3] You still need a few tricks up your sleeve for robust deployments. Anything outside of ARM, such as Azure AD groups and RBAC tend to be a PITA. Generally you want to wrap your deployments in a script that takes the object GUID of the created group and feed that into the template. That works, because the GUID can be used even if the full object is not fully replicated around yet.
Most of the failures we see are with AAD, which, AIUI, ARM templates do nothing for.
Even within ARM, as I understand them, ARM templates cannot handle deletes or changes. (They are deployments of new resources.)
And even if one could use templates, that just abstracts the same problem: how long do you wait for the template, if it has finished? (we see changes in ARM take >30 minutes to effect, and even after completion, it can take more minutes for things to "settle", i.e., successive requests to reliably return the same result. It just bottles it all into one highly inconsistent box, maybe, and requires me to learn an entire language on the side.) if it hasn't?
> That works, because the GUID can be used even if the full object is not fully replicated around yet.
Interesting. I'll keep that in mind.
> Generally you want to wrap your deployments in a script that takes the object GUID of the created group and feed that into the template. That works, because the GUID can be used even if the full object is not fully replicated around yet.
In the case I have today (wanting to perform an action on a newly created application), this trick doesn't work. (We make the call by specifying the application by ID, too.)
"Bad Request […] It looks like the application '[ID]' you are trying to use has been removed or is configured to use an incorrect application identifier."
It's not removed, of course, and the ID isn't incorrect.
Edit: actually, it is worse than that. So, there's the above error, and that's essentially a failure to have read-your-writes.
Our scripts retry on that, because our scripts have become accustomed to AAD's shit. But we eventually hit this sequence of events:
1. Create the app
2. Grant admin-consent
[other necessary setup]
3. Create an AKS cluster: <this fails>
And it fails because the App need to have admin-consent granted on it. But we do that, in step 2, and my logging is now good enough to show that we not only retry it after a read-your-writes failure, but that the command eventually succeeds, but the UI doesn't end up reflecting that. This is not a read-your-writes failure, this is a lost write!So far we haven’t had any issues or bad surprises, although we have setup some aws billing alerts just in case.
Feel free to make them responsible for cost and resources and you’ll be surprised how well they can manage their own account.
How many developers work at your organization?
My experience with these has been decidedly mixed. As in, you define them and never, ever see an alert.
We always get the alerts in time with thresholds set to 70% of the wished value.
Wow, this is horrible. I understand responsability but this is too much. Are other employees responsible if the company loses money for their actions?
Lessons learned for everybody, it’s a win-win situation.
I’ve had to do it a couple of times for personal and profesional accounts, and I’ve never had any rejections from them
Giving people responsibility and autonomy also comes with some responsibilities by the providers in a shared responsibility model is all I’m saying and every policy works out fine until it doesn’t.
Other accounts and environments (including dev) require everyone to follow a streamlined process: read only access to the account, a fully documented solution design and a corresponding terraform project in GitHub. terraform project checkin triggers a pull request for a review and an approval. Once the pull request has been scrutinised and merged, CI/CD runs terraform to provision resources in the account.
Not `destroy` - the opposite - destroy/nuke what I don't have in my config.
But the existence of aws-nuke makes me think (I haven't looked into what it's doing yet) there must be a better way of discovering used services/resources. Through billing perhaps?
Link to the repo in my HN profile.
terraform is best used on a per solution basis: one solution, one dedicated terraform project that will manage its state and its own state only. Multiple projects and people work carry out work in the same cloud account in parallel, and their terraform projects are not meant to interfere with each other. It works best for solutions that make use of fully managed cloud services.
Then there are also platform or connectivity level cloud resources (e.g. Transit Gateway and subnets that are mapped into the internal organisational network address space in AWS) that a random terraform project ought not to manage.
Lastly, if there is an actual need, a resource that terraform does not know about can be manually imported into the terraform's project state. This works best when infrastructure level resources were manually created a while ago and now have to be refactored into and managed by a terraform project. It is a tedious process that has to proceed with a lot of caution.
The caveat. It gets somewhat tricky when a non-serverless cloud resource requires an explicit subnet range allocation within an existing and managed CIDR or similar. There is no one solution to fit it all but containing such projects to their own dedicated VPC and setting up the VPC peering between a solution specific VPC and the main account VPC usually works satisfactory. That is, for example, how Kafka (AWS MSK) can be introduced into an AWS account without affecting existing CIDR mapping.
> Then there are also platform or connectivity level cloud resources (e.g. Transit Gateway and subnets that are mapped into the internal organisational network address space in AWS) that a random terraform project ought not to manage.
Not a 'random' one sure, personally I'd still want it somewhere. I suppose the hypothetical command might want an optional whitelist of non-tf-managed stuff to ignore though. (But then, you could whitelist it just by writing the terraform and importing it?)
> Lastly, if there is an actual need, a resource that terraform does not know about can be manually imported
The hard/annoying part that I'd like this command for is discovering these resources. i.e. it's not just that terraform does not know about them, it's that I probably don't. Or at least I don't realise they're not captured in terraform.
A very easy one to overlook is security group rules: unless you define them inline in terraform (i.e. ingress/egress blocks on a security group resource) then adding additional rules outside of terraform does not cause a diff. So you might be testing them out by manually poking around, and then you forget to terraform them/remove them, and they're left there forever with terraform blissfully unaware, and if you ever happen to notice it might not be obvious whether they're needed or not.
Essentially, it'd be useful for enforcing that terraform's used for everything; maintaining 'IaaC' hygiene.
Now has an API (with some caveats)
However, you can only close them from Control Tower at the rate of 2 to 3 per month, due to a hard limit quota which cannot be changed, even if you request it. Needless to say, this sucks when you've followed AWS's own best practices and created lots of accounts using Control Tower's "vending machine."
AWS's archaic account model is one reason we've switched to GCP.
Maybe it was quicker to implement with a hard limit, or there is some internal service that can't easily handle large volumes of account removals. But if it was Amazon losing money from this limitation instead of the customer I imagine it would be fixed pretty quickly.
https://docs.aws.amazon.com/organizations/latest/userguide/o...
A co-worker made this when we worked together on a project that ran a large number of terraform configurations in CI against real IaaSes. Each account was a sandbox, so we would run this at the end of a pipeline to clean up any failures or faulty teardowns.
Though, I do feel the project reaches out to quite a lot of IAAS's for its level of current maintenance, which seems to be somewhat zero.
aws-nukes gets contributions here and there and maybe that's because it's quite focused.
That's a real shame, the implementations for each provider look to be great, standalone by themselves.
https://cloud.google.com/docs/terraform/best-practices-for-t...
> After you run the terraform destroy command, also run additional clean-up procedures to remove any resources that Terraform failed to destroy. Do this by deleting any projects used for test execution or by using a tool like cloud-nuke.
> Warning: Don't use such tools in a production environment.
It seems like this tool only supports AWS though, at least nowadays.
I kind of feel bad for the GCP people here. Damned if you do (link to an external project, which might change), damned if you don't.
Besides this thing, I really enjoyed reading this article though! AWS is missing this kind of content.
https://cloud.google.com/resource-manager/docs/creating-mana...
All because some unknown AWS service that is in some zone that I cannot find at all and neither can AWS support.
Feel free to ask any questions.
I've found that the alternative to this is either a CI run that takes hours, or simply not testing the edge cases and hoping for the best. The former is a nightmare for velocity, the latter is a nightmare for having more than one person working on the team ever.
Lately, I've been fortunate to not be working in the "general cloud provider" space (where everything is a black box that can change at any moment, and documentation is an afterthought of afterthoughts) and have only focused on things going on in Kubernetes. To facilitate fast tests, I forked Kubernetes, made some internal testing infrastructure public, and implemented a kubelet-alike that runs "containers" in the same process as the test. For the application I work on at work, we are basically a data-driven job management system that runs on K8s. People have already written the easy tests; build the code into a container, create a Kubernetes cluster, start the app running, poke at it over the API. These take for-fucking-ever to run. (They do give you good confidence that Linux's sleep system call works well, though! Boy can it sleep.) With my in-process Kubernetes cluster, most of that time goes away, and you can still test a lot of stuff. If you want to test "what happens when a very restrictive AppArmor policy is applied to all pods in the cluster", yeah, the "fake" doesn't work. If you want to test "when my worker starts up, will it start processing work", it works great. (And, since the 100% legit "api machinery" is up and running, you can still test things like "if my user's spec contains an invalid pod patch, will they get a good error message?") Most bugs (and mistakes that people make while adding new features) are in that second category, and so you end up spending milliseconds instead of minutes testing the parts of your application that are most hurtful to users when they break. (And, nobody is saying not to run SOME live integration tests. You should always start with, and keep, integration tests against real environments running real workloads. When they pass, you get some confidence that there are no major showstoppers. When they fail, you want the lighter-weight tests to point you with precision to the faulty assumption or bug.)
Anyway... it makes me sad that things like aws-nuke are the kind of tooling you need to produce reliable software focused on deployment on cloud providers. I'd certainly pay $0.01 more per VM hour to be able to delete 99% of my slow tests. But I think I'm the only person in the world that thinks tests should be thorough and fast, so I'm on my own here. Sad.