How to win at CORS
jakearchibald.com
jakearchibald.com
This allows me to type a staging/production URL into Chrome, and get the frontend and the backend from either my local machine or staging/prod. I can mix and match any combination by checking/unchecking a box in Proxyman.
This means there is no need to whitelist localhost for CORS, and other hoops. Another advantage is that you're experiencing the app with SSL, so you may notice bugs that you would miss if you're used to work with HTTP locally. I've had these bugs which "only happen on production" in the past, and it's a nasty thing to deal with because it will be in a rush, since it happens as a surprise, and impacts users immediately. It can also bypass QA if the QA environment is also using workarounds and not a prod-like setup with SSL and the likes.
I recommend giving it a try. It's a workflow I haven't seen promoted anywhere before.
1. Let testing/staging environments present the same, valid certificates and keys as your production domain. Simply keep a copy for your internal use.
2.1 modify /etc/hosts so that production aliases localhost
or
2.2 let an internal DNS service do the spoofing work.
You'll still need some action to switch between environments. Full ownership of certificates is a recommendation. At your own risk.
Additionally, I run my own internal CA for various reasons and Firefox/Android seems to have stopped recognizing it. Irritating but it's a good razor for whether to upgrade internal or experimental apps to a Let's Encrypt cert.
They let you use hostnames that redirect, so for example: https://foo.127.0.0.1.sslip.io will redirect to your localhost.
There are many ways to lose at CORS and this one is my story.
Getting Vary right isn't just important for Chrome, it's important for CDNs too.
Hmm, I think I'll add a section to the article on this when I'm back at my laptop.
Thanks for prompting this!
S3 is used as an example, as it does not include a Vary header for non-CORS requests. However, the same would be true of some other origin which isn't correctly inserting a Vary header.
For us it got triggered because an image on the site appeared both as a video "poster" attribute (which loads with CORS) and as a regular image. So depending on which image the user encountered first you would see a CORS error. But it would be gone after a reload, so devilishly hard to reproduce until you realise what's happening.
Also took me ages to figure out that Cloudfront didn't include the Origin in the Vary header if it wasn't in the original request.
(@jaffathecake perhaps that Cloudfront behaviour warrants a special mention. Great article by the way, I learned a lot)
Incidentally, while you're here, I'm planning on building something which will require use of SharedArrayBuffer. Will adding the (now) required Cross-Origin isolation COOP/COEP headers cause any issues to existing CORS setups?
:)
I usually use NodeJS, but it turns out the HTTP library they use turns the HTTP method into an enum, so only a subset is supported (https://github.com/nodejs/node/blob/d798de1c653efa5ec0015d44...). This restriction only exists in their HTTP/1 library, their HTTP/2 library supports any method.
Anyway, I couldn't use that, so I used Deno via Deno Deploy. Their HTTP library supports any method, and the APIs they use are very similar to web APIs, so it was really easy to get started. Here's the server code: https://github.com/jakearchibald/cors-playground/blob/main/i....
(These are my recollections from investigation I did back in 2013 when I was writing the first serious Rust HTTP library.)
On the other hand, deno's HTTP stuff is built on top of Hyper, a Rust library https://github.com/hyperium/hyper
But hey, with so many impostors around - they're not capable of fathoming how golden this advice is.
Sad truth is, since "developers" don't use this, they genuinely don't understand browsers and HTTP and that's what 's dangerous.
Some aren’t aware that the trailing slash is useless in the HTML syntax, simply being ignored by the parser and not doing anything. (Except for in inline SVG and MathML content, which switch the parser into a more XML-like mode where the trailing slash behaves as in XML.) I know you’re not in that category.
Some hold that the trailing slash should be encouraged because it reminds the reader it’s a void element.
I can imagine some recommending it for XML compatibility (which is related to the original purpose of the ignore-the-trailing-slash behaviour, though slightly inverted in direction), but I don’t think I’ve ever encountered anyone saying so.
I hold that in HTML syntax the trailing slash is mildly harmful, because I almost never see it applied consistently so that the document wasn’t valid XML syntax anyway (for example, your site’s source has a <link rel="preload" as="font" crossorigin href="/c/logo-font-5449c974.woff2">) and because it misleads people into thinking that it’s a way of closing elements.
So yeah, I’m curious, because it looks deliberate. Habit? Reason? :-)
(For those that don’t know about the two syntaxes for HTML: XML syntax for HTML is still a thing and probably always will be; load data:application/xhtml+xml,<html%20xmlns="http://www.w3.org/1999/xhtml"/> in your browser as a starting point for a demonstration.)
In this first example, ASI inserts an undesired semicolon:
return
{a: 0}
This returns undefined, and doesn’t continue on to execute the block containing a statement 0 with label a. (Change it to {a: 0, b: 0} and you get a syntax error because of this reinterpretation of what was intended as an object literal.)In this second example, ASI doesn’t insert a desired semicolon:
f()
[].forEach.call(…)
This becomes a syntax error, because the [] has become subscripting rather than an array literal. (Incidentally, [].forEach is smelly anyway; prefer Array.prototype.forEach, maybe assign that to a constant if you’re doing it much.) f()
['foo', 'bar'].forEach(...)
Which is probably a type error, except if `f()` returns something like: function f() {
return {
bar: ['Not the array', 'you were expecting'],
}
}
Then you would actually iterate over the returned `bar` array, not the expected `['foo', 'bar']` array.Another fairly likely example of a desired semicolon not inserted is when the following line starts with a template string literal:
f()
`This is tagged with whatever f() returns`
This will most likely be a type error, except if `f()` returns a function, that function will then be called on the template string to do whatever. However I have a hard time imagining when you would want to start a statement with a template string literal without doing something smelly like: `some ${interpolated} string`.includes(someVar)
? sideEffectA()
: sideEffectB()For formatting, I just let https://prettier.io/ do it's thing, and it added the />. Although I do configure it to use single quotes in JS, so I guess I still have some opinion there.
In terms of HTML, how far does your "but it isn't necessary" opinion go? Lots of closing elements are unnecessary in HTML, for instance, check out the source of https://fetch.spec.whatwg.org/
For stuff I’m working on with others, I’ll act more like a normal person, though I will still prefer to drop at least <head></head><body></body></html>, and I’ve never worked with anyone that wanted to put trailing slashes on void elements (and haven’t ever used Prettier on HTML, evidently).
As for Prettier putting the trailing slash in: huh, that’s a really weird decision (and no flag for it!), given that they’re not emitting valid XML (not escaping >, at the least), so it’s just the personal preference thing, for something that was just added as an XHTML compatibility mechanism.
Yeah, I don't always agree with Prettier, but ugh, I wasted hours in my early career arguing about formatting with teammates, but now I just let Prettier do it's thing, get over it, and spend the time on something else.
Well, pleased to meet you!
I do use it for XML compatibility: it allows me to use XML editor modes (usually nxml in Emacs) which, being simpler to implement as they don’t have to encode the rules of which tags close where, are more likely to be available and to work well.
I’ve also done little customizations to nxml over the years, and this way I have them always there no matter whether I’m working with XML or HTML.
And then occasionally I use other XML tools on it, such as xsltproc. You can still pipe the HTML through tidy to translate it (although xsltproc has --html, tidy has been more reliable), but having it compatible with XML does spare me small chores again and again.
However, the competition at the time was Flash, and making <audio>/<video> so much harder than it was with Flash would have put developers off.
Fwiw, I regret that opaque responses can go into the service worker cache, since it caused quota-sniffing issues that we had to work around. But, if we didn't allow it, it would have been a feature regression vs appcache. sigh
When I'm writing some frontend that is hosted on localhost, with an API that is hosted on its domain somewhere, it always is some sort of PITA to get the dev environ started.
There's a plugin for firefox that ignores CORS which is helpful for this. It's becoming less useful for me as my APIs now usually have a toggle to add a cross origin header which allows localhost. Still useful.
I assume this would also work for CORS purposes: for some time I've not used localhost where possible, for SSL reasons. Giving the local machine a perfectly valid name that I can get a cert for via LE (or already have a cert for, I actually use a non-production name for which I maintain a wildcard cert) is slightly less faf than having my own signing cert installed as trusted everywhere I might need it. Anything I might do publicly is HTTPS-only so my dev/test environments are too.
Integrates nicely with existing JS frontend tooling.
I just want to easily let my users make API requests to an API whose CORS settings don't allow it. Instead I have to tell my users to give me their API key so I can make the request from my server. Or I'd have to tell them to run a program that makes the requests.
Both options suck for non-technical users.
You can! And it can include credentials!
You can do this with a basic <form> element, so fetch() lets you do the same. What you can't do is read the response.
Demo: https://jakearchibald.com/2021/cors/playground/?prefillForm=...
You know that nice "Access-Control-Allow-Credentials: true" header? In theory, it means that you can make authorized requests with cookies included to the cross-origin API. It actually has some extra rules that aren't obvious though. The cookies won't actually send unless they explicitly have "SameSite=None" set. Which itself isn't valid unless you are both making the request over HTTPS, and the cookie also has the flag "Secure". See here: https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Se...
The cookie logic is kind of mysterious. When everything is set up correctly and it works, it's pretty invisible to the code making the cross-origin requests. But if it isn't set up quite right, it will just not work, and it can be pretty tricky to figure out what's wrong. The whole SameSite logic is a pretty interesting kludge too. Ideally, most things would be SameSite=Strict, but that also means that links the user follows from third-parties to your site won't include cookies, since in theory those requests could include GET params doing who knows what.
Digging into all of the CORS rules and how they interact is almost an archaeology project into how the web was built and all of the weird things that both legit sites and malicious attackers have tried to do over the years.
Seems like I need to word it better though.
Perhaps a detailed walkthrough of that one specific scenario, which would enable almost all others, would be helpful.
If you load an IMG into WebGL as a texture, that's not allowed cross-domain. It's considered "processing" the image, rather than just displaying it. I ran into this when displaying slippy maps from map tiles. You can display the map tiles with JavaScript regardless of origin, but use WebGL, and you have to deal with the origin problem.
Reading image pixels isn't allowed unless the resource passes a CORS check.
I think failing to define acronyms is a violation of just about any style guide out there.
Communication is a majority of engineering so I think these things are very important.
https://addons.mozilla.org/en-US/firefox/addon/access-contro...
because properly configuring the backend is ¨too complicated¨.
Been there, done that.
But I was referring to legacy code (or those whose SPA is stored on the same domain as API endpoints).
This is an appallingly poor and misguided take, and goes against the most basic rules of writing technical documents. Docs need to be clear, unambiguous, and self-contained. The very first time a acronym is presented, it must be after the full name is presented.
It takes less than a sentence to do the right thing. There is no excuse.
I work with a client who built a web app in... Vue, I think. For unknown reasons they decided that it would be better that the APIs they need to call to live on the same domain. At the same time, the developers decided that the API microservices should not return CORS headers. Instead it was left to operations to hack in CORS headers in the webserver/loadbalancer.
I agree though that it is good practice to say what acronyms mean when using them.
I actually forgot what it stands for the other week, but it didn't prevent me understanding it. And relearning the acronym didn't help me understand it more.
It's engaging and even entertaining. Very well done!
I don't know why, but this cracked me up. :D
For dynamically generated content, the site runner should also consider open up if the content is for public consumption and not varying by user's credential.
As a pentester, I always get excited when I see ACAO or an OPTIONS request in my proxy logs. It's still really hard to wrangle and get right.
The lead prosecutor, Erin Nealy Cox, then took a job with the firm that leads Boeing's criminal defense.
It is essentially security by obscurity and protects nothing. Don't get me started how some technologies like AWS Lambda with a Gateway, when a function has an error, responds by default in such a way it makes the browser log a CORS error instead of, you know, a 500 or the actual error message.
I always do the "allow anything" setting for CORS when I make an API.
It definitely protects more than nothing :)
Also, I know quite a few public sites that serve debugging data if the request comes from the company's IP range. They shouldn't be doing this, but they do.
Maybe one day we can remove CORS for no-credential requests if we can detect that the destination isn't "internal", and we just decide that folks who serve debugging data by IP deserve to have their data leak. I've heard ideas around this for 10 years now, but maybe it'll happen eventually.
> It is essentially security by obscurity and protects nothing.
What.
I think there's been a fundamental misunderstanding of who is being protected here on your part, and what CORS is actually for.
It's your run-of-the-mill user that CORS protects, and CORS being enforced protects them when they visit e.g. an attacker-controlled site with, for example, a valid cookie-based session on your service. It prevents the attacker's site from making dangerous authenticated requests to your API service and reading the result.
> If someone is really determined they will either 1) turn off CORS with a browser extension 2) simply call your precious API from something other than an a browser
That's not what CORS is designed to protect against at all, of course you can hit an API separately and your HTTP client will ignore the CORS headers. Your HTTP client isn't a browser! It doesn't need to worry about CORS.
Similarly, users shooting themselves in the foot by disabling CORS are only hurting themselves. They are not the attacker here.
No, it isn't security through obscurity. Yes, it protects and it protects a lot.
Thank you for contributing to lack of knowledge, do keep up.
Not understanding CORS and making a comment like this is taking that ignorance to a new level though. Please read up on what you're talking about.
But no, it's not really a great idea to make every single privileged API in the world completely insecure just so the admins of public APIs can avoid adding a wildcard header to their servers.
It's not security for the provider of the API, it's security for the user of a web browser