Amazon S3 will no longer support path-style API requests
forums.aws.amazon.com
forums.aws.amazon.com
To put it simply, right now I could put some stuff not liked by Russian or Chinese government (maybe entire website) and give a direct s3 link to https:// s3 .amazonaws.com/mywebsite/index.html. Because it's https — there is no way man in the middle knows what people read on s3.amazonaws.com. With this change — dictators see my domain name and block requests to it right away.
I don't know if they did it on purpose or just forgot about those who are less fortunate in regards to access to information, but this is a sad development.
This censorship circumvention technique is actively used in the wild and loosing Amazon is no good.
If you're looking for a similarly robust and scalable alternative, Google Cloud Storage is accessible via S3 API when enabling that in the bucket's configuration and it supports path-style access (at least back when I tested the different S3-compatible services).
Just as a counter argument, one of the things we tried to do at a previous employer was data exfiltration protection. This meant using outbound proxies from our networks to reach pre-approved urls and we don't want to mitm the TLS connections. This leaves a bit of a problem, because we don't want to whitelist all of s3, the defeats the purpose, so we had to mandate using the bucket.s3 uri style, which is a bit of a pain for clients that use the direct s3 link style, but then we could whitelist buckets we control.
I don't want to say this use case is more important, but I can see the merits of standardizing on the subdomain style, and that this might be a common ask of amazon.
It was allowed to go on simply because noone knew about it. Could a skilled attacker spread a card card number across three lines and get past the system? Absolutely. Is exfiltration protection pointless? Absolutely not. Once you scale a certain number of users, you'll find someone somewhere that completely ignores training (which they did have) and decides they don't see the problem with something like this. And you won't know about it until you put a suitable system in.
Instead of investing in i.proved tools and productivity for your workers so they dont have to do stupid shit you dont want them to do, you instead made it harder for everyone to do their job.
I'm not saying your business is going to crash and burn. I'm saying you will never be as successful as you could have been. You're literally wasting resources, leaving needs unfulfilled, and giving up ground to your competitors.
You can make it harder to do on accident, or to prevent someone from doing it for convenience (e.g. someone copying data to an insecure location to have easier access to do their job), but you can't stop a malicious actor from getting the data.
They could tunnel over DNS, they could use a camera phone and record the data on the screen, etc. The possibilities are endless.
That's a major point of exfiltration prevention, both because accidents are a real problem and because reducing the opportunity for accidents makes it easier to establish that intentional exfiltration is intentional, which makes the ability to impose serious consequences for it greater (especially against privilege insiders with key contractual benefits that can only be taken away for cause.)
Technical safeguards aren't standalone, they integrate with social safeguards.
But that’s not how it is advertised. Usually they claim is to catch hackers and mal intending employees.
Why are you discounting these as valuable use cases?
In my experience, they're far more common than the determinedly-malicious actor. And they're far, FAR more common than the malicious actor who also (a) knows that the exfil monitor is there, and (b) has the technical prowess to circumvent it. (The analog hole is only _trivially_ usable for certain kinds of data.)
(I am continually frustrated by the number of people who claim that protection is worthless if it's potentially circumventable. In most situations, covering 90% of attacks is still worthwhile.)
- end up forcing you to lock it from the inside, and then crawl out the window
- have your friend who is visiting request a door-opening-token 24hr in advance through a JIRA ticket
- cause the power to go out once it's locked, also for security reasons
- force you to replace the keys with 'special' plastic ones from a new third-party vendor
- leave you stranded outside for a few hours because the door-opening system is having an outage
Those are the kind of trade-offs that will be made, not simply the act of locking the door.
'arcbyte over at https://news.ycombinator.com/item?id=19827012 does have a point - a lot of potential exfil risk is caused by companies doing their best to make it as difficult as possible for their employees to do their jobs.
People are saying that because it's misguided and potentially harmful.
Doing so is security theater, where the solution is scoped down to something incomplete but easier, and then everyone walks away happy they solved 90% of the smaller problem they chose to attempt.
Particularly with things like data exfiltration, this is potentially harmful because then you've organizationally blinded yourself.
Nobody wants to poke holes in their own solution, and so they stop looking.
But, hey, we're catching the odd employee accidentally sharing confidential documents via OneDrive.
Fast forward a year, and an entire DB gets transferred out via an unknown vector, nobody finds out about it for a couple months, and it's all "Oh! How did this happen? We had monitoring in place."
Go big, or run the risk of putting blinders on yourself.
Network security is a balancing act between prevention, detection, needed user and network capabilities and cost. If I have unlimited money or no limitations on hindering network usage I can make a 100% secure network - it's not even that expensive, just unplug it all.
It doesn't cover most attacks. That's why it's so misleading. It mainly just protection against incompetent people from accidentally sending out data.
This is not a good general principle, since it can be easily applied in contexts that (I would predict) many same individuals would vehemently disagree with. For example:
- Personal privacy is pointless, there are simply too many ways for governments/corporations/fellow citizens to find things out about you
- Strong taxation enforcement by governments is pointless, there are simply too many avenues for legal tax avoidance and illegal tax evasion
- Nuclear arms control is pointless, the knowledge of how to make a bomb and enrich uranium is widely available (I mean, if NK could pull it off, how hard could it be?)
Maybe data exfiltration prevention isn't a good policy, but I think you need a more nuanced argument than 'there are ways around it'.
They are not a guarantee that an incident cannot happen, though. They can only lower the rate at which incidents (privacy violations, tax evasion, nuclear proliferation) occur.
Same with exfiltration.
Here are a few more (somewhat) related to this topic:
1.Joe Grand, “Advanced Hardware Hacking Techniques”, Defcon 12 http://www.grandideastudio.com/files/security/hardware/advan...
2.Josh Jaffe, “Differential Power Analysis”, Summer School on Cryptographic Hardware http://www.dice.ucl.ac.be/crypto/ecrypt-scard/jaffe.pdfhttp:...
3.S. Mangard, E. Oswald, T. Popp, “Power Analysis Attacks -Revealing the Secrets of Smartcards” http://www.dpabook.org/
4.Dan J. Bernstein, ''Cache-timing attacks on AES'', http://cr.yp.to/papers.html#cachetiming, 2005.
5.D. Brumley, D. Boneh, “Remote Timing Attacks are Practical” http://crypto.stanford.edu/~dabo/papers/ssl-timing.pdf
6.P. Kocher, "Design and Validation Strategies for Obtaining Assurance in Countermeasures to Power Analysis and Related Attacks", NIST Physical Security Testing Workshop -Honolulu, Sept. 26, 2005 http://csrc.nist.gov/cryptval/physec/papers/physecpaper09.pd...
7.E. Oswald, K. Schramm, “An Efficient Masking Scheme for AES Software Implementations” www.iaik.tugraz.at/research/sca-lab/publications/pdf/Oswald2006AnEfficientMasking.pdf
8.Cryptography Research, Inc. Patents and Licensing http://www.cryptography.com/technology/dpa/licensing.html
Copying files from one folder into another could do the job.
Oh, this was a fascinating read! Thanks for encouraging my perusal.
There is also research about using modem lights (even in the background of a room) to figure out what people are doing on their dialup internet connection. Those RX and TX LEDs are actually blinking at your data transmission rate.
It's not really the same threat model as people living under dictatorships, but it might just work.
[√] absolute dependence on authorities for food, shelter, clothing, transportation, money.
[√] curfews often in effect for you and your social circle, especially if suspected of deviance.
[√] 24/7 electronic or in-person monitoring is possible and largely accepted.
[√] social circle often molded by authorities.
[√] not allowed to vote or generally exercise political agency (and when allowed it's dismissed).
[√] not allowed to leave your workplace or home without permission from authorities.
[√] possible to flee and seek asylum but it means leaving everything behind for an uncertain future.
[√] indoctrination is so effective you're extremely likely to continue the system when allowed to be an authority.
Good thing it's a benevolent regime.
Not that there aren't problems with both P/S education and higher education or public discoure and media generally, though your analysis misses a few key salient aspects and presents numerous red herrings.
J.S. Mill affords a longer view you may appreciate:
https://old.reddit.com/r/dredmorbius/comments/6x7u6a/on_the_...
AWS S3 will only provide SSL validation if your bucket name happens to not contain "."
Which is a practice encouraged by AWS. [1]
So anyone that has www.example.com as the bucket name can no longer use HTTPS.
[1] https://docs.aws.amazon.com/AmazonS3/latest/dev/website-host...
Still, seems kinda lazy from their part, they could just generate a custom cert.
More seriously; it's going to be a nasty migration for anyone who needs to get rid of the dots in their domains. At minimum you need to create a new bucket, migrate all of your data and permissions, migrate all references to the bucket, ...
I mean, if Amazon wanted to create a jobs program for developers across the world, this wouldn't be a bad plan. :(
$0.005 / 1,000 copy requests...
ref: https://blog.cloudability.com/aws-s3-understanding-cloud-sto...
Also you will likely want to use some sort of parallel operation. I used this eons ago: https://github.com/mishudark/s3-parallel-put
The only way it would be a lot of money to move to a new bucket is if the bucket is hardcoded everywhere. Moving data from one bucket to another is not expensive, and a configuration change to a referenced URL should be cheap, too.
They probably got an ultimatum from Chinese authorities to either stop allowing this or get blocked entirely.
They just blocked wikipedia last week.. no one is too big to get shutdown in China
Thinking about it the other way round, how likely would have Amazon been the target of similar attacks?
1: https://arstechnica.com/information-technology/2015/04/meet-...
Because the S3 buckets are virtual-hosted they share IPs so there is deniability if you can hide the DNS/SNI.
You need TLS 1.3 because in prior versions the certificate is transmitted plaintext, but eSNI itself is not part of TLS 1.3 and is still actively being worked on as https://datatracker.ietf.org/doc/draft-ietf-tls-esni/
If a hypothetical tyrantical government was willing to block all of Amazon S3 this change doesn't affect anything.
China has blocked GitHub and Akamai before. https://www.latimes.com/business/technology/la-fi-tn-great-f...
There are like <1% websites in China relies on S3 to deliver static files. Blocking AWS as a whole has happened before. There is simply no freedom was "collateral". Freedom has to be fought hard and eared.
Edit: Why am I getting downvoted, it's a legit answer, CDN hides your origin.
Chinese government will just ban the whole s3.amazonaws.com domain. Same as facebook.com, youtube.com, google.com, gmail.com, wikipedia...
However letting them banning sub-domains will actually make S3 a useable service in China. It's a huge step forward.
(Of course "who is politically right" and "who has the most technical expertise on their side" are at best tenuously related, but that's a different and longstanding problem. If you believe you're politically right and you have technical expertise on your side, use it.)
Viagra solved tiger poaching political problem.
Isn't there everything to gain from encouraging everyone to use encryption so that there are too many targets to process?
In particular, a political solution requires that people be able to communicate (in order to work together), and technology can be a component of that.
https://www.dropbox.com/s/zzr3r1nvmx6ekct/Screenshot%202019-...
There are so, SO many teams that use S3 for static assets, make sure it's public, and copy that Object URL. We've done this at my company, and I've seen these types of links in many of our partners' CSS files. These links may also be stored deep in databases, or even embedded in Markdown in databases.
This will quite literally cause a Y2K-level event, and since all that traffic will still head to S3's servers, it won't even solve any of their routing problems.
Set it as a policy for new buckets, if you must, if you change the Object URL output and have a giant disclaimer.
But don't. Freaking. Break. The. Web.
That seems very inconvenient, and is pretty inline with my experience with aws: I guess their services are cheap and good, but oh boy! The developer experience is SO bad.
- So many services, it is very hard to know what to use for what - Complex and not user friendly APIs - coming with terrible documentation
I'm pretty sure they'd get a lot more business if they invested a bit more in developer friendliness - right now I only use aws if a Client really insists on it, because despite having used it a fair amount, I'm still not happy and comfortable with it.
There are other providers out there like Digital Ocean Spaces, Wasabi, and Backblaze that offer storage solutions for much cheaper than S3 now.
Digital Ocean Spaces and Wasabi in particular actually use the Amazon S3 api for all their storage. This means you can switch over to either of those solutions without changing the programming of your app or the S3 plugins or libraries that you are currently using. The only thing you change is the base url that you make api calls to.
Backblaze has their own API, but they also offer a few additional features not offered under S3's api.
I don't think that's true ... we (rsync.net) try to very roughly track (or beat) S3 for our cloud storage pricing and we've had to ratchet pricing down several times as a result.
I don't think we just imagined it ...
I have a DO box for myself and their docs/admin panels are better imo.
But my comment about aws is not really a comparison, just more a comment about my experience as a non-devops engineer, and how I hate having to read their docs.
The takeaway is that for those of us that do still wish to uphold those values, we can let this serve as a lesson that we should not publish assets behind urls we don't control.
Now, it seems, this is a big problem. V2 resource requests will look like this: https://example.com.s3.amazonaws.com/... or https://www.example.com.s3.amazonaws.com/...
And, of course, this ruins https. Amazon has you covered for * .s3.amazonaws.com, but not for * .* .s3.amazonaws.com or even * .* .* .s3.amazonaws... and so on.
So... I guess I have to rename/move all my buckets now? Ugh.
e.g.
> The name of the bucket used for Amazon S3 Transfer Acceleration must be DNS-compliant and must not contain periods (".").
and as you mentioned
> When you use virtual hosted–style buckets with Secure Sockets Layer (SSL), the SSL wildcard certificate only matches buckets that don't contain periods. To work around this, use HTTP or write your own certificate verification logic. We recommend that you do not use periods (".") in bucket names when using virtual hosted–style buckets.
AWS Docs have always been a mess of inconsistencies so this isn't a big surprise. I dealt with similar naming issues when setting up third-party CDNs since ideally Edges would cache using a HTTPS connection to Origin. IIRC the fix was to use path-style, but now with the deprecation it'd need a full migration.
Wonder how CloudFront works around it. Maybe it special cases it and uses the S3 protocol instead of HTTP/S.
It's worse than that. You can't rename a bucket. You will have to create a new bucket and copy everything over.
https://aws.amazon.com/blogs/aws/new-amazon-s3-batch-operati...
Hmm, I was going to say something about the _cost_ of getting/putting a large number of objects in order to 'move' them to a new bucket. Does the batch feature affect the pricing, or only the convenience?
Sadly neither batch operations nor replication is free.
How so? cross-region replication doesn't replicate existing objects, only new ones.
I set this up some time ago using our domain name and ACM, and I don't think I will need to change anything in light of this announcement.
1 - https://docs.aws.amazon.com/AmazonS3/latest/dev/website-host...
2 - https://docs.aws.amazon.com/acm/latest/userguide/acm-overvie...
You could still put CloudFront in front of your bucket but CloudFront is a CDN, so now your bucket contents are public. You probably want to access your files through the VPC endpoint.
With a CloudFront proxy you’d have to open up access to all of CloudFront’s potential IP addresses to allow the initial request to complete (which would then redirect to S3). Plus the traffic would need to leave your VPC.
Suddenly we had two options. Use CloudFront with hundreds of SSL certs, at great expense (in time and additional AWS fees), or change the names of all buckets to something without dots.
But aaaaah, S3 doesn't support renaming buckets. And we still had to support legacy applications anf legacy customers. So we ended up duplicating some buckets as needed. Because, you see, S3 also doesn't support having multiple aliases (symlinks) for the same bucket.
Our S3 bills went up by about 50%, but that was a lot cheaper than the CloudFront+HTTPS way.
The cynic in me thinks not having aliases/symlinks in S3 is a deliberate money-grabbing tactic.
Now one would need to hook the cert validation and ignore dots which can be quite tricky because deeply hidden in an ssl layer.
https://docs.aws.amazon.com/AmazonS3/latest/API/RESTObjectPO...
First to allow them to shard more effectively. With different subdomains, they can route requests to various different servers with DNS.
Second, it allows them to route you directly to the correct region the bucket lives in, rather than having to accept you in any region and re-route.
Third, to ensure proper separation between websites by making sure their origins are separate. This is less AWS's direct concern and more of a best practice, but doesn't hurt.
I'd say #2 is probably the key reason and perhaps #1 to a lesser extent. Actively costs them money to have to proxy the traffic along.
For core services like compute and storage a lot of the price to consumers is based on the cost of providing the raw infrastructure. If these path style requests cost more money, everyone else ends up paying. It seems likely any genuine cost saving will be at least partly passed through.
I wouldn't underestimate #1 not just for availability but for scalability. The challenge of building some system that knows about every bucket (as whatever sits behind these requests must) isnt going to get any easier over time.
Makes me wonder when/if dynamodb will do something similar
A more optimistic view is that this allows them to provide a better service.
By sharding/routing with DNS, the client and public internet deal with that and allow AWS to save some cash.
Bear in mind, S3 is not a CDN. It doesn't have anycast, PoPs, etc.
In fact, even _with_ the subdomain setup, you'll notice that before the bucket has fully propagated into their DNS servers, it will initially return 307 redirects to https ://<bucket>.s3-<region>.amazonaws.com
This is for exactly the same reason - S3 doesn't want to be your CDN and it saves them money. See: https://docs.aws.amazon.com/AmazonS3/latest/dev/VirtualHosti...
Anycast will pull in traffic to the closest (hop distance) datacenter for a client, which won't be the right datacenter a lot of the time if everything lives under one domain. In that case they will have to route it over their backbone or re-egress it over the internet, which does cost them money.
https://twitter.com/colmmacc/status/1067265693681311744
Google Cloud took a different approach based on their existing GFE infrastructure. It does not really seem to have worked out, there have been a couple of global outages due to bad changes to this single point of failure, and they introduced a cheaper networking tier that is more like AWS.
I don't think that's true. Route53 has been using Anycast since its inception [0].
The Twitter thread you linked simply points out that fault isolation is tricky with Anycast, and so I am not sure how you arrived at the conclusion that you did.
[0] https://aws.amazon.com/blogs/architecture/a-case-study-in-gl...
It's a good choice for DNS because DNS i a single point of failure anyway, see yesterdays multi hour Azure/Microsoft outage!
Their DNS nameservers which resolve those subdomains do of course.
S3 isn't designed to be super low latency. It doesn't need to be the closest distance to client - all that would do is cost AWS more to handle the traffic. (Since the actual content only lives in specific regions.)
You have a sole ip address. All traffix routed to nearest PoP. The PoP makes the call on where and how to route the request.
Lookup google front end (GFE) whitepaper. Or thd google cloud global load balancer
That front end server that lives in the PoP can also inspect the http packets for layer 7 load balancing.
https://cloud.google.com/load-balancing/docs/load-balancing-...
They _do_ use anycast and PoPs for the DNS services though. So that's basically how they handle the routing for buckets - but relies entirely on having separate subdomains.
What you're saying is correct for Cloudfront though.
Raw data could flow from a different PoP that's closer to DC.
Aka user->Closest PoP-> backhaul fiber -> dc->user
Currently all buckets share a domain and therefore share cookies. I've seen attacks (search for cookie bomb + fallback manifest) that leverage shared cookies to allow an attacker to exfiltrate data from other buckets
https://developer.mozilla.org/en-US/docs/Web/API/document/co...
With s3.amazonaws.com, they need to have a proxy near you that download the content from the real region. With yourbucket.s3.amazonaws.com, they can give an IP of an edge in the same region as your bucket.
s3.amazonaws.com subdomains are as distinct from each other as co.uk subdomains.
Absolutely mind boggling with as much as they pay people they do something so stupid and haven't changed it after so long.
Edit: Could they have found a better place to announce this than a forum post?
S3 has been around a long time, and they made some decisions early on that they realised wouldn’t scale, so they reversed them. This v1 vs v2 url thing is one of them.
But another was letting you have “BucketName” and “bucketname” as two distinct buckets. You can’t name them like that today, but you could at first, and they still work (and are in conflict under v2 naming).
Amazons own docs explain that you still need to use the old v1 scheme for capitalized names, as well as names containing certain special characters.
It’d be a shame if they just tossed all those old buckets in the trash by leaving them inaccessible.
All in, this seems like another silly, unnecessary, depreciation of an API that was working perfectly well. A trend I’m noticing more often these days.
Shame.
I guess after this change, the cors configuration will finally do something!
On the flip side, anyone who wants to list buckets entirely from the client-side javascript sdk won't be able to anymore unless Amazon also modifies cors headers on the API endpoint further after disabling path-style requests.
This could be just as disruptive.
Difficult to say that they will actually follow through, as the only mention of this date is in the random forum post I linked.
This is a great way of introducing breaking changes. Imagine that for example, ipv6 would be at near 100% adoption if "new websites" were only available over v6.
Something weird is going on if they don’t keep path style domains working for existing buckets.
Edit: autotypo
It’s not like I’m super outraged that they would change their API, the reasoning seems sound. It’s just that if I have to touch S3 paths everywhere I may as well move them elsewhere to gain some synergies with GCP services. I would think twice if I were heavy up on IAM roles and S3 Lambda triggers, but that isn’t the case.
> We asked our investors and they said you're very excited about it being less good, which is great news for you!
The scale at which different libraries, tools, and systems depending on hard-coded S3 urls will break by this change is insane.
ag -o 'https?://s3.amazonaws.com.*?\/.*?\/'| awk -F':' '{print $1, $4}' | sort | uniq | cut -d'/' -f 1 | sort | uniq -c | gsort -h -rk1,1
For anyone interesting in finding out the occurrences in their codebase. (Mac)Also: see photos of your favorite celebrity walking their dog and other news at 11.
Over a million results (+250k http). This is going to be painful.
Migrate
from: s3.amazonaws.com/<bucketname>/key
to: <bucketname>.s3.amazonaws.com/key
no later than: September 30th, 2020
<bucket>.s3.amazonaws.com
Do I need to change my origin to be, Origin domain name: s3.amazonaws.com, Origin Path: <bucket>
This is a sneaky one that will bite lots of folks as it is NOT clear.
<bucket>.s3.amazonaws.com is the V2 url formula.
How does forcing customers to rewrite their code to confirm to this change, improve customer experience?
That's the exact opposite of good customer service.
It also helps people understand why the bucket name is restricted in it's naming.
How does it do that? You can host a private bucket at foo.s3.amazonaws.com just fine.
If it's https://s3.amazonaws.com/foo/ you could believe that it's based on your cookies or something, but if it's https://foo.s3.amazonaws.com/ it's more obvious that it's a global namespace in the same way DNS domain names are (and that it's possible to tell if a name is already in use by someone else, too).
In this case, the most highly improved experience I can think of eould be that of sundry nefarious entities monitoring internet traffic.
What changes is that the bucket name must be in the hostname.
I'd have found it really useful. :-/
Is there even a UI option to make a bucket public anymore? I always edit the bucket policy and add the JSON to make it public read only.
How hard is it for 99% of the developers and technical leaders here to search your codebase for s3.amazonaws.com and update your links in the next 18 months?
I've got a number of hobby projects, some hosted on AWS, that I built ages ago. I have no idea how this change will effect those projects because ... I just frankly don't remember the codebases. I built them on a weekend, set them up, and now just use them.
It isn't the end of the world. But I'm not really excited about having to dig up old code, re-grok it, and fix anything that changes like these might break.
I suppose that's just the nature of a developer's life. But I think many of us long for a "write once, run forever" world. Horror stories about legacy software aside, it was nice to be able to write software for Windows and then have it work a decade later.
Well, I think AWS developers are in the same boat right? Here we are.
An architectural decision that many years ago was the approach now needs to be rethought and updated.
So, surprisingly hard, but doable. And from the customer's perspective, a huge pain in the ass, just to save Amazon some pennies on bandwidth.