Twilio: Someone broke into our AWS S3 silo, added 'non-malicious' code to JS SDK
theregister.com
theregister.com
They've invested all their resources into building out the infrastructure, and have spent very little if any time on building the UX around it. You have to be fairly technical, never mind painstakingly detailed-oriented, in order to manage their services, a big chunk of which is hidden behind black screens and various control nobs.
Amazon's logic is probably to provide the piping, and leave the rest of the end user, but at this point it's not enough. In my opinion, they need to heavily invest into the customer-facing layer of their applications that shows all the status, gaps and holes, or these things keep happening.
Then again, they can continue as is, and the incidents will keep as is, and as long as nothing implodes and no big customer leaves, financially it makes sense to do nothing. Just thinking out loud.
Looking forward to finally try Pulumi in a project.
Ever had CloudFormation yell at you because some function like "Sub!" only accepts alphanumeric characters? Created complicated constructs because there are no variables, and very limited conditionals? Copy-pasted templates like an insane person because there is no way to abstract anything?
Not my idea of fun. All of this is solved by using a proper programming language.
The biggest paint point with something like Terraform is the fact that it can't actually validate cloud provider specific problems (for example declaring an Compute instance with a compute type (string) that doesn't actually exist). Does Pulumi try to address this in some way (granted its much easier to write your own checks like that in language as opposed to writing a go module for Terraform providers).
Yes, the AWS Pulumi SDK package includes a discriminated union of valid instance types, and the type of the property `aws.ec2.Instance::instanceType` is that union.
(Disclaimer: I am a former maintainer of Terraform and a contributor to Pulumi).
I worked for a company where our CTO identified the same thing (unauthorized modification of JS SDK served via S3) as a potential vulnerability. It took us half a day to get the bucket locked down to where only the root AWS IAM user—secured by physical 2fac—was permitted to change the SDK, and then only after "unlocking" the configuration changes. Since the SDK was not updated frequently, having a little manual ceremony was fine. If it had been a high-touch path we would have had to do a little more engineering to keep it safe, but not that much more.
That said, I don't blame Twilio for not catching this; shit happens at a company with years of legacy built up. I do blame anybody who says "S3 is hard" and throws up their hands. You're a professional! Read the docs, they're dry but quite extensive! Play around with it in a test bucket! It's no harder than learning any other moderately complicated professional system. If you can learn Rust, you can learn S3's API. If you can learn Kubernetes, you can learn S3 ACLs and the IAM authz model. It may not be as interesting as the topics you want to read about, but it will probably be even more useful.
To do anything in AWS there's at least 3 or 4 vaguely connected services, including IAM and RAM as a completely separate UIs. Usually with S3 you also have Cloudfront or some other CDN. You probably also have Cloudwatch logs and CLoudtrail event tracking. You might have VPCs involved. Some or all of these things, in an org like Twillio, could have entirely different teams managing them.
The problem isn't that S3 is hard. The problem is, as OP suggests, the UX of actually doing anything non-trivial across half a dozen services is somewhere between abysmal and war-crime.
If you’ve got one team managing your logging who don’t talk to whatever team who manages your s3/cloudfront setup, then that’s your problem.
I love Terraform, but the fact that you require an entirely different company's orchestration tools to make AWS "not that hard to orchestrate" is evidence of how poor the UX of AWS is.
Most of the stuff available with terraform is available with CloudFormation minus a thing or two. (takes them a minute to catch-up to other teams new services/features)
AWS console was only complicated as a complete noob. Once you understand how services work you actually realize is decent.
Thanks for the giggle.
I do.
Twilio is a multi-billion dollar company and there is no excuse for them not having proper security processes to catch stuff like this early. Even if we take the "S3 is hard" arguments at face value, this wasn't a 0-day or some complicated unpredictable exploit. This was an extremely basic misconfiguration on a mission-critical part of their architecture that would have been caught by even the simplest out-of-the-box penetration test or audit. For a company like Twilio to not be doing basic, fundamental stuff like that is a big deal and they certainly should be blamed for it.
better tooling makes this easier to do.
And you should have this done by somebody who knows it well, or who learns it well in the process.
No one knows S3 better than the AWS S3 team.
It took us half a day to get the bucket locked down to where only the root AWS IAM user—secured by physical 2fac—was permitted to change the SDK
Or a small feature team could spend a month or two improving the user experience so that literally every S3 customer doesn't have to have the technical acumen, knowledge of risks, and time to get this right. 0.5 days * a million users (to be conservative) is 500,000 days.
And they definitely haven't been doing nothing. They've drastically redesigned the permissions UX around S3 over the years. They introduced GuardDuty to allow automatic threat detection (such as S3 bucket having open permissions) which can be applied across your organization (not just an account). They definitely have to keep doing more, but implying they are doing nothing just isn't true.
In AWS you need to explicitly set a bucket as publicly accessible, which requires you to go through a couple of dialogs where you're expected to say "yes, yes I want to allow others to access this S3 bucket."
Afterwards you still get a bucket that does not have write access by default. To allow write access you need to specifically set bucket access policies that explicitly allow others to write to your bucket.
So, from the confusing UX to the multiple permissions settings, I can easily see how a company could get in this situation. In fact, the problem is so bad that AWS has revised the whole S3 thing several times, with each new iteration more secure than the last. Unfortunately, that doesn't help the company that configured their buckets 5 years ago and haven't touched them since.
But today if you make an S3 bucket, they take you through a wizard that makes it secure by default. They also run audits for free and send you the results.
They also have Lightsail that just does everything for you securely.
They have lots of user friendly options.
FTFA:
> During our incident review, we identified that this path was not initially configured with public write access when it was added in 2015. We implemented a change 5 months later while troubleshooting a problem with one of our build systems and the permissions on that path were not properly reset once the issue had been fixed.
> { "Sid": "AllowPublicRead", "Effect": "Allow", "Principal": { "AWS": "" }, "Action": [ "s3:GetObject", "s3:PutObject" ], "Resource": "arn:aws:s3:::media.twiliocdn.com/taskrouter/" }
Obviously you have to be using CloudFormation, so if they're using i.e. Terraform they're out of luck.
AWS Config is easy to setup and can be used to warn for that type of access-control issues. It's a good layer of additional defense.
The new S3 UI is also much better and won't let you do insecure things by default.
AWS also sends emails when they find an open bucket. I don't know exactly when those are triggered but I got them in the past.
However, I’m not necessarily sure that it is the wrong approach. You can use AWS at almost any level you like and if you go deep into the woods with any service you’re eventually going to find yourself buried in documentation.
Essentially, people only ever want one of two things:
- A private FTP bucket
- A public static asset website (with built-in CDN and HTTPS!)
Thing people never want:
- A public upload bucket
Seriously, no one wants that in any circumstance, ever. It's a huge liability. I literally cannot imagine anyone wanting anything even similar to it. But many, many companies have accidentally made one with S3.
While they're at it, S4 should also actually use IAM and all that other "modern" stuff that S3 doesn't have because it's so ancient.
Second Simple Storage Service.
Thus is outright wrong. Public upload s3 buckets are AWS's go-to basic technique to handle file uploads. AWS's API Gateway has a max file limit, and AWS advises developers to leverage S3 and lambda triggers to implement that usecases.
[1] https://docs.aws.amazon.com/AmazonS3/latest/API/sigv4-post-e...
[2] https://docs.aws.amazon.com/AmazonS3/latest/dev/PresignedUrl...
This requires a Bucket Policy that is written explicitly to allow this. I don't want to make assumptions because of course we don't know all the information and I personally know very little about infosec, but that sounds very strange.
Why would you do that? I assume the SDK gets uploaded from a secure environment that has the credentials to upload objects, so there's no legitimate reason to have writing permissions open to the world.
> [...] a 2FA verification code sent to your phone (via a call, SMS message, or an authentication app like Authy).
PSTN-based 2FA in 2020 is not a vote of confidence in the security team's competence.
A lot of companies go with email and/or PSTN based 2FA for customer convenience and to get them to use 2FA for their accounts.
I try to avoid using 2FA for accounts that only offer PSTN or email based 2FA because it's less secure than my randomly generated long password. It would be nice if more services offered 2FA based on client side TLS certificates.
1) When not implemented properly, it is treated as an alternative authentication mechanism, not an additional required mechanism. My ultra strong password is useless if it's enough to steal my phone number and convince customer support that the attacker simply forgot the password.
2) It gives a false sense of security. I already have 2FA enabled! What do you mean I should be using TOTP/FIDO2/TLS client certificates?!
3) It reinforces an unfortunately widespread security anti-pattern, which reduces security for Web users in aggregate. So what if it's not a great idea? Everyone is doing it! Or put another way, MD5 password hashing is better than storing passwords in plain text. But please, please, please do not use MD5 password hashing! There are much better ways to accomplish that goal without introducing more flaws.
It seems this incident has reminded Twilio to audit all their access right, which is a net positive.
Actually, imagine you were an employee who wanted to get this done, but couldn't get management approval? This is one way to fix things ... :)
a. adopted the resource into a Cloudformation stack & a1. Enabled drift detection.
b. Use an AWS config rule to monitor (appropriate) s3 buckets for any public access.
That's not a AWS problem. That's a "our team failed to manage the service" problem.
This includes monitoring a publicly accessible bucket.
AWS is getting really confusing, IAM is not simple, which is bad.
But S3 is a bread-and-butter service, and the policy is such that anyone using it should understand. I mean it says there in 'almost plain English' than 'anyone' can write content.
This is a pretty major blunder, because it hints that maybe they don't have any security reviews at all.
If I, a completely non security-knowledgable person were asked to do an audit of Twilio, I would find this issue. It's way on the lower treshold of 'obvious problem'.
Twilio is not 'Medium Blog' - they provided IT plumbing and therefore it's doubly problematic.
"Specifically, the modification added code to the end of the TaskRouter.js v1.20 SDK that made an HTTP GET request to hxxps://gold.platinumus.top/track/awswrite?q=dmn and followed the URL returned in the HTML by that request."
Translated: we can't possibly know it's non-malicious.
TFA even has this to say:
> And judging from the URL involved, it appears to be an attempt to install a payment-card skimmer – RiskIQ has spotted the same URL in other S3 buckets targeted by miscreants.
Details in the linked blog post[0].
So, very much malicious.
Why the hell they included "non-malicious" in the title (granted, in quotes), I don't know. Readers could have easily dismissed this as actually non-malicious.
[0] https://www.riskiq.com/blog/labs/misconfigured-s3-buckets/
If you hit it directly they will redirect to a random "Spin the wheel win a prize" type site in an attempt to hide the malicious nature of the payload.
A failed drug dealer (who got caught), by your definition, should get less time because they never managed to sell the drugs they were transporting over the border and got caught :)
The impact of these actions matters far more to the victims than whatever consequence the perpetrator may face.
If your website was affected by this hack, surely you’d care about how your customers were impacted.
Lots of people here keep repeating the claim that this was magecart, that’s just not true.
A common S3 use case is: Use a Bucket for Read-only static content, (js, html, imgs). From a dev perspective ideally it will work like a protected folder + Web Server, whereas the webserver will have read-only privileges, but reality is far more complicated.
I can understand why was easy for Twillio to have this infosec issue, since we've done it ourselves many times when using S3.
If it’s cross account then allow assuming to other accounts, but no reason to bother with the bucket ACL.
Leave the bucket policy blank and you’ll never have to worry about an open bucket. Better yet make a deny rule to everything but a single role, and you don’t have to worry about rogue roles exposing the bucket.
Perhaps some of the issue is that permission can be managed on both IAM and a bucket policy?
You already have a role in AWS that you use. Go into the console and say “this role can do S3 to this bucket” and you’re done.
I'm sure an expert finds it all very natural. Security is always hard and has to be done right the first time. That's a domain that's always going to deter novices, and perhaps the only way forward is to hire an expert to ensure you're doing the right thing.
Still... I've been programming for four decades and it felt a little odd to be quite so out of my depth. It's not the same as anything I've done before, but it felt like it shouldn't have been quite so bewildering.
But, as others have said, the documentation at times is poor at best, and it can be easy to make mistakes.
Some preventives every ops team can take are:
1. Never make manual permission changes in the console. Everything is done via SCM’d code.
2. All permission changes are peer reviewed.
3. Create logical separations of buckets so it’s more obvious what the permissions should be of a bucket. You can go so far as to have separate accounts that are for public uploads, separate accounts for public assets, etc. Avoid granting elevated permissions at a folder level, which is harder to understand as a developer without access to view the ACLs.
Good luck everyone!
The first time I heard “infrastructure as code” I thought it was some dumb new buzz words, but no, it’s good advice. Your cloud formation, serverless, terraform whatever should be stored in GIT a stack replace is how changes should be made.
I consider the console as read-only.
The web console should be just another UI for editing the configuration in source control. As an example of why that's important, a UI can give you a warning and use affordances to guide beginners away from dangerous setups. In the case of S3, a slider for basic access level can use colors and show a list of objects that would become world readable or writable at a given level (for that matter why is world writable even an option?). A giant YAML file looks the same whether it's safe or dangerous.
If you really don't want to use a UI, then maybe abstracting a linting layer that feeds the exact same warnings to the UI as to your CI system and text editor would help.
Should. But isn’t.
So until it is, IaC for production. Console or CLI for development.
I do the same, and sometimes I need to do something manually so that I can grok what configurations I need to define in terraform / whatever other tool I'm using. POC'ing is different than a production or even staging service.
It would be much harder make that mistake with Google Cloud Storage, at least in the web console. GCS also supports disabling bucket ACLs permanently at bucket creation time, and that is the option they recommend[1].
AWS does scan and alert for some public bucket issues. But with all the public S3 issues, I think it is clear there needs to be more prominent warnings in the AWS console, the command line, and with tools like Terraform.
S3 does too. There's an entire page during the creation wizard that is dedicated to blocking any and all public access, even causing the bucket to ignore ACLs or other settings that would otherwise expose the bucket publicly. All public access is disabled by default, and enabling it actually requires the user to actively uncheck 5 different checkboxes, each of which explains that unchecking it will open the bucket up to public access, and then even requires an additional attestation, in a large orange warning box, that says you acknowledge that the options you chose will result in the bucket being public.
I'll be the first person to tell you that AWS is overcomplicated and hard to use, but this isn't that. IAM and bucket policies are a pain to work with, but if you screw up with those, the worse you will do is expose your bucket internally. But exposing a bucket publicly to the internet is an entirely different act, and there's not really any excuse for it other than just not reading the directions.
However, neither this option in S3 nor the option you linked in GCS would have served Twilio's use case. They wanted their bucket to be publicly accessible, just not publicly writable. That's an entirely different access management issue.
Before using AWS I was in the 'what morons!' camp whenever some company got breached via leaving things wide open on AWS. Now I'm in the 'ok I understand why its so common' camp.
You'll even see this in their sales pitches. AWS will spend so much time talking about the most advanced use cases and how they are possible, but if you ask them a simple question about how to run a single bucket that doesn't also involve using all of their ML tools and global accelerator and cloudfront, they'll be blindsided. It's almost as if they consider the common use cases to be "too simple/obvious" and therefor they never bothered to create any documentation or put any thought into it.
But with that said, one of those layers of defense is also using tools that are easy to keep secure so that the vulnerability doesn't appear in the first place. While Twilio isn't "a lone developer", I wouldn't be surprised if the person who did create that bucket on behalf of Twilio was just a lone developer fumbling around in the console. And again, that's no excuse, but it still is an area for improvement.
Now, as has been stated before, reading any cloud vendor documentation is like reading the Oracle manuals of years ago, not a happy experience.
Granted, these clouds now do things much more powerful and target more difficult enterprise scenarios. One can miss Heroku, and AWS has Elastic Beanstalk for that, but EB gets much more difficult fast.
I understand someone brand new to S3 being in the dark but someone in your PRs approval chain should have enough Ops skill to know to check these things if you're not going to hire an actual Ops person. And no, none of this is hard for anyone who has ever touched S3. Ops teams have known about these issues for over a decade. It's practically a meme at this point.
The number of heads on my project is ...4. We're it. There is no place or person to just push the work off to. There is no PR approval chain. We regularly meet to sync and discuss plans and thats it.
Until very recently all of our infrastructure was in house, right down the hall from me. I (or whoever was in that day) took care of making sure it was running smoothly.
Cloud is new to us, and there is no budget to hire a guy who manages AWS full time. That is the case at an insanely high percentage of software shops. Heck I would wager most people outside of 'new tech (sorry, I can't think of a better term)' shops are using point & click in the web UI to manage AWS. There is no code in source control that manages our infrastructure. I'm jealous of the people that get to do it all the 'right way' but for many there isn't enough manpower or time in the day to do it.
Now why we're in this position in the first place is another debate about why management thinks we need to go cloud at all. But its reality for many.
I wish MORE small dev teams brought in ops people and didn't wait until they had a staff of 20-30 or whatever it typically winds up being. But I know that's a pipe dream and you're basically losing a developer salary to hire a devops/ops person which isn't the direction they want to go.
edit: typos
This is anecdotal, of course, but it seems to me that any company that didn't start cloud-native is going to have flaws like this, and it's often not at the fault of developers but at the fault of top-down initiatives to "get to the cloud" ASAP instead of correctly.
(I work for Twilio)
This incident report should really put to bed all of the "It's AWS's fault for making things so complex" complaints. (To be clear, it won't... but it should.)
Even a cursory look at that bucket policy should tell you something named "Allow Public Read" should NOT be associated with anything named 'Put'. This takes 0 AWS knowledge to figure out.
And stating to the press the clearly malicious payload is "non-malicious" (assuming TFA didn't lie about Twilio's statement)? That's ridiculous.
They owned it. That is more than can be said about other large incident reports that I've seen regarding AWS.
Report: https://any.run/report/e8c581d52ca3527b61c7ddd9aba5dfeaf9440...
Original URL: h t t p s : / / gold.platinumus.top/track/awswrite?q=dmn
I only know this because a friend talked about a case where they rotated a car in-place, and there was a question about is that theft? its center of mass is still where they left it, so it wound up not being theft. Moving furniture in a house is a weird one. Probably leans on ill intent, moving a chair to block a door seems like it's pointed to giving the owner a hard time. moving a chair to perform cpr has much purer motives.
So anyway, there's my ill-informed view.
It's a really simple service that fetches a list of script URLs that you provide, and notifies you with the before/after diff whenever the file changes.
I found GuardScript through a Show HN a while ago (https://news.ycombinator.com/item?id=20265141) and have been very pleased with it.
https://developer.mozilla.org/en-US/docs/Web/Security/Subres...
For example, here is Stripe's script tag:
<script src="https://js.stripe.com/v3/"></script>
SRI only works for static files whose contents never change. In this example, it would mean that Stripe could never push minor bug fixes to the SDK.
The workaround would be to use static versioned files instead of a generic script URL (or self-host the script), so that the value of the SRI checksum never changes.
Edit: In the case of Twilio, their docs say "By linking directly to our CDN in production, any patch-level (non-breaking) changes made to the SDK will automatically be applied [...]" meaning that if you link to Twilio's CDN, you can't expect the file to be static, and therefore can't use SRI.
Subresource Integrity is supported by every evergreen browser and 80%+ of market share (no IE11/pre-Chromium Edge).
[0] https://www.twilio.com/docs/taskrouter/js-sdk/workspace/work...
https://developer.mozilla.org/en-US/docs/Web/Security/Subres...
However, on Twilio's documentation site, they do not include the integrity attribute in their examples: https://www.twilio.com/docs/taskrouter/js-sdk/workspace/task...
Even putting the security arguments aside you can bet a lot of customers would be really irritated having to constantly update this stuff on their site. If a competitor didn't have that restriction it would become a selling point. IMO the ideal is that Twilio provides both options: a "sdk-latest.js" that's a moving target, as well as "sdk-v1.3.js" that is frozen and can be locked with subresource integrity.
<script src="..." report-integrity="https://cdn.stripe.com/report">
in which clients compute the integrity and asynchronously ping a lightweight payload of the integrity to the third party? You would then seed your report service with a list of known hashes and set up alerting for when new hashes were observed.Regardless of the severity of the bug, the only-case scenario is that all the sites you have pulling from that CDN break until you recompute the hash. How annoying this is scales directly to how frequently your libraries have to release vital security bugfixes.
Until you recompute the hash and communicate that new hash to them and they implement it on their site. It’s not nothing from an implementation point of view.
In reality, no you cannot.
You have multiple layers of cache between the S3 bucket and rendering, unless you disable caching entirely at a massive increased cost. Some of these caches are poorly behaving (e.g. intermediary caching).
The correct way of doing this ALREADY is to increment the version number in the URL (e.g. /3.1.0/My.Lib.js to /3.1.1/My.Lib.js) and to re-point the pages to use the updated library. This is reliably cache breaking, and will assure end users are getting the bug fixed version faster (or ever in some cases).
Once you're already doing it correctly, Subresource Integrity is a freebie. There's a reason why "sdk-latest.js" is largely dead concept from a bygone era: it is a huge anti-pattern and anti-feature.
Heck allowing upstream to blindly push you changes without being informed is nuts regardless. This argument is essentially: "I'm doing it wrong, and Subresource Integrity would stop that, so it is a non-starter." Without considering that what you're doing is bad practice before Subresource Integrity joined the party.
> There's a reason why "sdk-latest.js" is largely dead concept from a bygone era
Is it, though? From a quick check, it's what Google Maps does. It's what the Facebook SDK does. We already know it's what Twilio does.
> This argument is essentially: "I'm doing it wrong, and Subresource Integrity would stop that, so it is a non-starter."
Zero dispute with that characterisation from me. But it's how things already operate in the real world. I'll join you in shouting from the rooftops that people shouldn't be doing it, but that doesn't really get you any closer to actually stopping them. For a great many people the flexibility to quickly push up changes is a feature, not a bug.
Nope. Google Maps loads an uncachable JS file that itself points to version specific sub-files that are cache breaking.
> It's what the Facebook SDK does.
Nope, they have versions in the fbAsyncInit configuration.
> We already know it's what Twilio does.
Indeed, and if you want to copy a company that almost had a massive security problem then go right ahead.
Allow admins to "down feature" the account. Ie, simple mode.
Basically, allow admins to toggle OFF S3 ACL's and even S3 Bucket Policies - zero those out. Then let users use IAM policies for access which may be a more familiar / more used mental model.
"Someone" was able? If the bucket was unprotected and world-writable, then _everyone_ was able.
> 'non-malicious'
> And judging from the URL involved, it appears to be an attempt to install a payment-card skimmer – RiskIQ has spotted the same URL in other S3 buckets targeted by miscreants.
In what world is that non-malicious? Non-immediately-damaging _maybe_, but it seems clear that there was a malicious chain of reasoning involved. Twilio got super lucky if this is all that happened.
Time to find an alternative.
They might as well have said "we are not smart enough to figure out the nature of this code. In unrelated news we are hiring senior JavaScript engineers, experience in opsec is a bonus"
And you wouldn't get much sympathy, but this is a legally viable complaint. When my car was stolen, the thief was charged for theft, but his girlfriend who was joyriding with him was charged for criminal trespass. Just being in the car when you don't have permission is illegal.
It isn't necessary to actually "break" anything.
Edit: Of course, it's still not a great metaphor because it implies there was at least some kind of barrier to entry that had to be pushed aside - but from the sounds of it, there wasn't.
* the cost of auditing would be spread across all customers equally, punishing competent customers and benefitting incompetent ones..
However I'm not sure if there are any built-in tools to find weird nested policies attached that could cause a similar permission to be granted to all..
If you ask for help from AWS, AWS will provide it. It may not be free, but it's available.
Even if AWS were to start proactively auditing customer setups, how in the world is AWS supposed to know what a customer's usecase is? Nevermind the fact it's a breach of the customer's privacy to just go rooting around in the customer's account without permission.
But let's assume AWS is going to take responsibility for customers' configuration decisions and violate customers' privacy by proactively auditing their accounts. Would AWS auditing Twilio's configuration here work?
The default is for S3 buckets to be private. The customer has to take specific, affirmative steps to give s3 buckets public access. You really have to jump through hoops to make a bucket public accessible.
Since the Twilio chose to make their bucket public, AWS auditing Twilio's setup wouldn't be helpful. AWS would just assume the Twilio knows what they're doing. How is AWS supposed to know there is a misconfiguration? Because Twilio clearly decided to make their S3 bucket publicly available.
I don't bother because it seems pretty clear what buckets are public.
That said, quick tip to make your life easier.
DON'T use S3 ACL's DON'T use S3 policies.
If I was AWS and didn't have so many customers I'd probably just create one mental model (IAM policies probably) as the place to manage things and block the rest.