My Amazon S3 Mistake
devfactor.net
devfactor.net
1. Use IAM roles. AWS keys should only have access to the specific functionality they need. Best practice is to never use your root credentials for anything; always use IAM users and use unique keys for each use case, so that you can invalidate and replace them easily. Every time documentation or a tutorial asks you to insert AWS keys, your response should be "let me go create a role for that" rather than "let me go look those up".
2. More importantly, if you ever accidentally publish or leak credentials, don't try to clean it up by deleting those commits. Invalidate the credentials immediately and re-issue them.
3. Always, always `git diff HEAD` before committing. Know what you're about to push up. This isn't just a security concern - the number of small, stupid things you'll catch that you'd otherwise end up fixing 15 seconds later is substantial. As a bonus, this incentivizes you to keep your commits small and atomic.
You might talk to Amazon directly - they've been known to forgive debts like these in these kinds of circumstances.
They did:
> Lucky for me, I explained my situation to Amazon customer support – and they knew I wasn’t bitcoin mining all night. Amazon was kind enough to drop the charges this time…
I prefer to always "git commit -v", which shows me the diff while I edit the commit message.
Could you explain what parts of it you don't like? Also, could you link to git gui? I tried to Google it, but didn't really get good results (shockingly, all the other Git GUI programs came up...).
It just seemed so confusing compared to the other tools I mentioned, and it would try to do all this extra stuff for you automatically (i.e. adding weird extra arguments to the git commands and I had no idea what they did)
My workflow is primarily raw CLI (no aliases, no github extensions, etc). I have vim configured as my primary editor and my .vimrc has the vim-git plugin (https://github.com/tpope/vim-git) loaded so I can enforce 50-char summary and 72-char line-wrapping.
I use other clients for special purposes:
Fugitive (https://github.com/tpope/vim-fugitive) primarily for git blame while working.
SourceTree for visualizing branches and for staging individual lines from a hunk in a finer grained fashion than `add -p` allows. Occasionally for branch maintenance when an interactive visual list is useful.
gitk for loading partial histories (eg. --author options, pickaxe option), visualizing complex branch arrangements that make SourceTree choke, anytime I need performance to search back many months.
The only reason Linux is nicer for certain kinds of programming is because Bash and the GNU utils are so great. But why bother when there's a world of people making things so easy and smooth for Windows?
I like Sourcetree for history and diffs, but it's CLI all the way for commits.
Seems like the best lesson learned is keep personal work personal, use private repos if using GitHub.
That said, I'd still argue that you should practice defense-in-depth; plan for "this is how I'll limit damage when someone gets ahold of these credentials" rather than "Nobody will get these, so I can do dangerous things with them".
It's the same argument behind "don't do things as root unless you have no other choice"; trading security for convenience works really nicely right up to the point where you get your teeth kicked in.
and use least privilege to make sure that if you do screw up it's as minor as possible.
And personally, I've found that it's just good practice to exercise those kinds of precautions for "throw-away" projects regardless of their "throw-away" status. It keeps you honest.
I used to be an AWS fanatic and got burned. Unless you need to scale up on the spot (and have the time to code management of that) it's a huge waste of time for no good price. EC2 instances have dismal performance.
Just get decent servers somewhere. For most scenarios it's hard to max out a server nowadays, if you program decently. You have to do that for AWS 10x anyways, IMHE. The cloud is a black hole of developer time and problems.
On top of that it's stupid to put data in US when you are a foreign company, you are open game for 3 letter agencies no questions asked.
BTW, tweaking around with servers can also be a major time sink.
You learned the wrong thing :-P The lesson here is "always check what you're committing"
I realize I haven't always followed this advice myself. So the reminder from the OP story is useful.
In essence, "committing" then is the act of pushing your changes to your local repo. Or you could, you know, just look at the changes on disk if really required.
I've simply stopped using Amazon for anything tinkery, because the costs of making mistakes can be tremendous. At least when I make mistakes on my own colocated server, I know it can never cost me more than the $100/month I pay to host it. And, storage is practically infinite (4TB hard disks), and I can spin up more VMs than I would ever need for tinkering in 32GB of RAM on our "spare" server.
And, when I have needed to use VM with some cloud host...Linode and Digital Ocean and similar may have dramatically smaller toolsets for managing virtual resources than Amazon (probably unworkably so for large deployments), but my mind has a much easier time predicting costs than with Amazon. After being surprised on more than one occasion with a ~$300 bill from Amazon (for running nothing but personal pet projects with no economic return), I turned everything off.
Hmm.
Logging into AWS and checking the Cost Explorer showed that I had been charged for a t1.small EC2 instance I had running, but when I logged into the AWS console, all I saw was a single "stopped" instance. There's no way I was being charged for cycles I'm not using, right?
Turns out that the instance I was being charged for was in us-west, and I had only been looking at us-east- and there is no indication within the actual EC2 panel that you have instances in other regions!
I was a bit perturbed, to say the least...
For me, when I saw a $300 bill for one running EC2 instance, a couple of halted ones, and a few hundred GB of storage for something I thought of as a "toy" project that I didn't want to invest serious time or money in, I knew I was done with AWS. A small colocated server could readily provide those resources for vastly less money (and I work on tools to manage cloud and VM resources, including a reasonably good API for spinning them up and down and such, so I don't really miss the Amazon API or UI). This was after I'd already gotten a shock from an automated backup gone wrong that ran up a huge bill. So, it took me a couple of times getting burned.
Of course, it's always been my fault for not understanding how Amazon bills for things, what services cost money (i.e. a down instance still costs money), how much things cost, and sometimes how the API works (or at least confirming that it's doing what I think it's doing, in the case of removing old backups). I'm not blaming Amazon. I'm just saying, I don't trust myself to use Amazon for anything that I'm not going to spend a lot of time and energy on, because I'm obviously not capable of using it without making mistakes when I treat it like a toy. I readily admit I shot myself in the foot; Amazon just provided the guns.
Years later I'm using AWS professionally, and it all seems easy now. But when you are first introduced to AWS and want to kick the tires on a toy project, it can be confusing and overwhelming.
http://docs.aws.amazon.com/awsaccountbilling/latest/aboutv2/...
tried googling but it doesnt seem so straight-forward... i like the idea though.
There are cheaper providers and more expensive providers...sometimes, it can be hard to tell whether you get what you pay for, without actually trying it. But, most of the big colo providers have been discussed online at Web Hosting Talk or other places. When getting new service, I usually check WHT first for deals in my local area in a good data centers. Sometimes, for low bandwidth (where low bandwidth may mean something still kinda huge, like several terabytes of transfer per month) a reseller may be the best deal. Sometimes, going straight to the provider is the best choice. I ended up a CoreXchange by way of a deal from ColoUnlimited.
The good thing about going this route is that ongoing costs are somewhat lower, generally speaking, though up front costs are much higher (~$3000 for a server, for example). And, costs can be predicted with high precision. You can simply say, "I want this much bandwidth." and that's what they'll give you and bill you for. You're pretty much just paying for power and bandwidth when you're in a colo; all the other stuff is up to you (hardware upgrades, replacements, etc. either need to be done via shipping gear in and out or by going on-site to make the changes).
The bad thing about going this route is that it's all on you. This isn't "managed" hosting (though they do have staff on-hand, and you can usually pay for "remote hands" to handle system upgrades and such, it tends to be very pricey for anything more than simple reboots). And, scaling is non-trivial. In the bad old days, if you had a site get crazy popular overnight, you had to figure out how to get more servers online on short notice...maybe missing an opportunity to grow and blowing your shot at a good first impression on new users. AWS and other cloud services are much more readily scalable. But, there's no reason you can't build with scaling in mind, and use cloud services for elasticity while using colocated servers for your baseline service. That's how we do for our business stuff. If we suddenly need more servers online, we spin them up via Amazon or Google Compute Engine (the software I work on supports several cloud providers, as well as building cloud infrastructure out of heterogeneous servers, so this is not much different than managing VMs on our own hardware and we can move websites and such back and forth across our own machines and AWS VMs reasonably easily).
If you're running AWS, I highly recommend for all developers and sysadmin to attend AWS free training[0][1]. While you're at it might as well get yourself certified[2].
You might have senior expertise with system operation and application deployment but sometimes, AWS approach things differently. The essence of the training is to always implement best practices, not just solving problem.
Also use the opportunity of AWS event to network with their Solution Architect. Trust me on this one. This worth more than AWS Enterprise Support.
[0] http://aws.amazon.com/training/course-descriptions/architect...
[1] http://aws.amazon.com/training/course-descriptions/architect...
Hopefully I'm mistaken... because I'd love to take it if it's free!
I attended 2 free session in Kuala Lumpur and Singapore back in June. I always assume its free for all. Sorry
https://www.hipmunk.com/flights/WAS-to-KUL#!dates=Jan12,Jan1...
I haven't used it much, but it looks like you can be very specific in what you allow, such as only allowing access to a single bucket with S3, or a single domain with SES.
I use a specific key for only allowing DNS updates to be applied to "example.com" and no other domain, for example.
As others have mentioned the billing alerts are also a useful thing, along with the IAM Simulator.
All I had setup was a micro S3 instance I'd been using for some toy craigslist scraping & hadn't touched for months.
Then out of the blue I get an "urgent please check your account" email from amazon. Go check the AWS console, and what do you know -- maximum number of maxed-out instances instances churning away with 100% CPU usage in every region on earth. The charges were already about to $50,000 when I turned everything off.
I wrote a very, very apologetic email to amazon, and they forgave all the charges, for which I was very grateful.
Definitely a learning experience.
Or maybe just have a script on local github pre-commit hook?
Also how do you differentiate an AWS root key (bad) vs an IAM Key (good)?
In my circles, at least, it's standard practice to use environment variables.
But I would think clearly it'd be an option.
Do you have articles discussing the cons of AWS keys in private repos?
We deploy our systems on vanilla EC2 instances, which are configured by using a server orchestration system (Ansible). So for any env variables to get set, we'd have to put them in config scripts, which are currently checked into github.
To make it clear, we only check in our IAM keys that are AWS service specific, like SES.
Bots having tons of false positives doesn't really matter (except to the bot maker, maybe). But GitHub having tons of false positives means customers get annoyed by false alerts, locked data, whatever.
You might be right if it really is a ton, but then you work on your algorithm. I think the problem is so big that there really do need to be warnings for these kind of issues.
I was very surprised/impressed.
Lesson 2: When using AWS then use Billing Alarms[1].
It takes about 1 minute to setup and enables e-mail or SMS notifications at dollar-thresholds of your choice.
[1] http://docs.aws.amazon.com/awsaccountbilling/latest/aboutv2/...
Root credentials are deprecated, but can still be used and if this guy used them then yes, his billing alarms could have been disabled.
There's a difference between an admin iam user (can't do billing stuff) and the root credentials (just as powerful as username / password).
1. Used IAM roles (but he didn't lock down it's permissions nearly enough clearly)
2. Used two factor authentication.
http://snaketrials.com/2014/11/11/espionage/
http://snaketrials.com/2014/11/12/espionage-update/
I should note Amazon forgave the money his account owed, given it was his first time making this mistake on their systems. Amazon told him the servers were probably being used to mine bitcoins.
edit - Oh, and on his blog he says it was $2000 over 12 hours. It turned out to be $5000 over 2 hours!
If I want to learn rail, I'd start by installing it and running it on my local PC. What's the point of relying on heroku, aws, etc for just everything? What's the point of using github as the main hub for everything? I'm synchronizing stuff to github explicitly and carefully, my casual, day-to-day operations are all running on my own hardware, network and services. Even my most simple websites run first at home, then are sync'ed to the web after I've checked everything works fine.
one of the big things about Java that at first I thought was tedious but now have come to love is the build process. By creating the "built" project in a new folder, it is like you are guaranteeing that only stuff you manually specify to be in the production code actually makes it in.
i know .gitignore is supposed to handle this but it just seems dangerous to rely on source-control related filters when there is really sensitive security info involved.
in node i've used grunt & really like that model. The idea being that you have the project run a build then check that into a "build" branch in git & deploy from there. it's also nice because i can uglify etc. if i'm not worried about preserving the source.
I guess for distributed coding this is tough because you want your peers to have access to the raw source + build instructions, but if I were working on a really security-sensitive project (I do enterprise so I'm starting to understand the need for extra precautions) I would probably only distribute code that goes through a (perhaps even thin/transparent) build process & then figure out a way to integrate changes back into the source branch.
I know it must sound like overkill to the open source world & especially for Rails projects, where there ordinarily is no build stage, but it seems like there is some level of danger in distributing a source control repo where you are relying on you + any collaborators to properly configure .gitignore and never accidentally check anything sensitive into the repo. I am decent with git but no expert, i still sometimes accidentally stage things I shouldn't when initially checking in a .gitignore -- and i have much more faith in myself than some of my peers, hah.
This doesn't at all solve the OP's problem, you shouldn't be putting your build under source control in lieu of your source code.
i just appreciate the security builds add. In enterprise coding the paradigm is very simple -- pretty much keep whatever u want in source, but only built artifacts are ever deployed to the servers. In a closed-source setting this is perfect.
I understand what you mean though it doesn't really solve the issue of distributed, open-source development. (I guess no one would really make a public repo unless they wanted to share, huh?)
I am a big advocate of self hosting - not because services like github suck (in fact I'm a big fan of github, bitbucket and other services) but you have more control over your own code (that and you can setup private repos).
Here is a small list of self hosted solutions (I forked indefero into srchub - many don't like the google code feel but I personally like it for the simplicity and the fact I can easily fix/modify/add on to it): http://softwarerecs.stackexchange.com/questions/3506/self-ho...
For most individual projects - they are probably good enough. However, the thing to keep in mind is that for free they usually limit features (I think bitbucket limits the number of developers) or remove features all together (google code removing downloads support, github did at one time but reintroduced it). I'm not saying that I think they should be offering everything for free and shouldn't take away features - they have to make money - but in my opinion always have backups and a plan B in case you need to jump ship.
I was a huge google code fan (the design isn't Web 2.0 with social integrations - but I prefer functional over design) until they pulled the plug on downloads support (which for most Linux people isn't a big problem - a make && make install and it's compiled AND installed - for Windows it's not that easy). Which forced me to self host - I do mirror many of my public projects on github but in the event github management decides to remove/reduce a feature I won't be struggling to find a new "home" for my ~100+ projects.
Also - if you look at the issue tracker for google code it's pretty obvious that Google is no longer supporting or even monitoring it (several spam issues have been appearing). I don't have a crystal ball - but my best guess is that Google will pull the plug on Google Code as well. And it would make sense - most people have already "fled" Google Code for github/bitbucket etc.
With that said - let me put on my tin foil hat and say that even if we ignore any potential future issues there could be potential privacy issues if you upload your code to a private repo hosted by someone else. I'm personally not the type to snoop, but that might not stop some lowly intern from getting curious. Or even some bug on their site that makes private repos public (for a short time anyways). Did you hear about the time that Dropbox made passwords optional for four hours [1]?
[1] - http://techcrunch.com/2011/06/20/dropbox-security-bug-made-p...
Why doesn't an account automatically come with, let's say, a $100/mo default cap. And also a $25/day default cap. Someone playing around would just leave those defaults alone and thereby not wind up looking like a total fool.
Using Wikipedia for a guide, a micro instance only costs $0.013 per hour. An instance with 30 GiB memory and 8 cores only costs $1.00 per hour. So the cap numbers I suggested would work fine for many people.
For some reason Amazon allows people to shoot themselves in the foot too easily. Yes, they usually? waive the charges, but it seems like such a waste of resources, overall. It wouldn't be so bad if bitcoin mining were more profitable. But $1000 spent at Amazon probably mines $10 or less in bitcoin. What a waste of energy.
This blog genre would then be "Amazon deleted all my stuff and ruined my business because they didnt think I'd pay $26". As the story notes they tried to reach out multiple times using multiple contacts. There are also existing tools for setting up specific billing alerts if you really want to take some action at $26.
I ended up helping give a talk about my experience at the Amazon Summit in Sydney. I hope I made a good cautionary tale to the devs/ops/managers attending.
If your project is reading credentials from a file, rewrite it so it reads them from environment variables.
Most IDEs make it very easy to do that, and Python's virtual environments can do that work for you. Yes, it takes more effort, and sometimes it will be a little convoluted[1].
However, it's well worth the effort as you will have a system that you can put your faith in rather than having to double check every time in order to make sure you're not about to inadvertently commit your API keys.
[1] Example of my own: Pycharm does not read variables stored in a virtual environment's configuration, so I have to set them twice.
I had the situation in reverse. I found the Lynda.com rails tutorials first (in 2011) and I found it lacking compared to Rails Tutorial. I'm not sure what's the situation now, from Lynda.com website it seems the author (Kevin Skoglund) has updated the tutorials for Rails 4.
I have learned that if this is the case, I'm doing something wrong with my AWS setup.
E.g., "Buy our engineer intelligence subscription! Find out if your job applicant is a ninja or noob!"
I suppose there's an argument that the data is already public and anyone could apply that analysis, which is true. But the difference is that Github is not profiting from this; their feature development will continue to focus on private repo owners as their customers (whose interests are pretty much aligned with public repo owners).
$100/mo cap would save a lot of these accounts from being hacked. My mistake was not unique, and if you browse around on Google you will find other authors who have had the same issues.
I have not used OVH, but a quick perusal of the OVH API[1] suggests that you could invoke plenty of commands that would incur costs.
I wonder how often this happens. Are they mining bitcoins?
To illustrate this, EC2 GPU instances have a NVIDIA Kepler GK104 with 1536 cores, so that is like the GTX680. It gets 120 MHash/s on bitcoin, which translates to 2 cents per month.
It gets 207 kHash/s on litecoin, which translates to 32 cents per month.
I think on vertcoin you could get 90 cents per month.
If I remember well you could make a few hundred bucks a month with just one high end ATI graphic card.
Perhaps one lesson here aside from keeping keys outside of public repositories is to learn how an API works (IAM, arns, etc) before using it.
Can you really use the S3 API to spin up EC2 instances? or is the guy just mixing the fact that the credentials he used for S3 can be used for other AWS APIs ?
No. My guess would be credential with allow:* service access.
Adding environment variables to a file is just a convenience.
Fortunately it was forgiven.
I'm not sure how much clearer GitHub could make that.
I've never read nor seen that page.
shouldn't it be "its"?
Do I want to buy and manage servers myself? Would I pay for an additional computer to host my little toy Rails app? Will my ISP allow me to route network traffic to my personal IP with a standard plan? Most, at least that I've seen, prevent you from hosting.
For a little project something like the author of this article described, using virtual hosting is significantly easier (and more realistic). If you're going to run an enterprise operation, bringing that sort of hardware in house might make sense, but a lot of larger companies are still pushing that hosting out to places like Amazon as well.
Now, if you're asking why Amazon would allow (new) customers to scale up to $3000 worth of hosting overnight, maybe that's a separate issue. But how should Amazon judge that? There are probably users who want to jump straight up to that sort of scale. And evidently they did detect that the author didn't seem to fit that use case -- they actually called him about it pretty much right away.
EDIT: this got downvoted, but I stand by it, plus it's a question, so you could reply and answer it. In my thinking it's the same reason there's a daily ATM withdrawal limit set by default. You can lift it, but it's there to reduce incentive (payoff from trying to see your PIN and then stealing your ATM card.) the current policy is like the bank calling you and saying, "ummm, I hope you know you've already withdrawn $7,000 and seem to be continuing." Given that bitcoin is (literally) cash, it seems to me saner to not run bitcoin mining instances by default, unless you authorize it specifically. or can they not tell?
Amazon doesn't have access to the data on the instance, or a list of processes running inside of it, or similar.
The CPU or GPU pattern of bitcoin mining must be completely unambiguous and trivial for Amazon to detect on EC2 instances. Or am I wrong for some reason?
For instance if you look for a magic number. Instead of putting two numbers on the stack, and perform an addition you now have to have a conditional jump before the addition.
Maybe I'm wrong, but at least I think it would be virtually impossible.
If they want to be neutral and not do introspection like this without permission, they can ask the user on sign-up if they want "Protection from mining processes" where sudden activity spikes will cause them to do introspection and shut down an instance if it seems to be bitcoin mining.
EDIT: plus, bitcoin mining is the same operation over, and over, and over, and over again. You don't have to "catch" it doing a particular operation. A brief sample taken at any time will show the VM doing the same exact thing. (tons of SHA hashes.)
I'm not an expert though so perhaps I'm missing something!