OAuth from First Principles
stack-auth.com
stack-auth.com
One gripe from Attack #3[0]:
> The solution is to require Pied Piper to register all possible redirect URIs first. Then, Hooli should refuse to redirect to any other domain.
There's actually a another (better IMO) solution to this problem, though to date it's rarely used. Instead of requiring client registration, you can simply let clients provide their own client_id, with the caveat that it has to be a URI, and a strict prefix of the redirect_uri. Then you can show it to the user so they know exactly where the token is being sent, but you also cut out a huge chunk of complexity from your system. I learned about it here[1].
[0]: https://stack-auth.com/blog/oauth-from-first-principles#atta...
[1]: https://aaronparecki.com/2018/07/07/7/oauth-for-the-open-web
Unfortunately, most major websites end up hosting an endpoint that will redirect users to a separate URL provided as a query parameter. This means that users may easily be misled about where the token is, in fact, being sent.
[0] - https://sec.okta.com/articles/2021/02/stealing-oauth-tokens-... [1] - https://datatracker.ietf.org/doc/html/rfc6819#section-4.2.4
Is this true? Do you know of any major ones offhand? That would be surprising to me.
Thanks for sharing [0]. I found its discussion of shared subdomain cookies useful. However, I believe all the vulnerabilities in the OAuth section would be mitigated by using PKCE and not using the implicit flow, even if you leave the open redirect. Am I missing anything there?
As for open redirects in general, it is an important problem. As an authorization server, if you want to protect against clients that might have an open redirect (and as you indicate eventually one will), while still using the simple scheme I mentioned above, I can think of a few options:
1. Require the client_id to exactly match the redirect_uri instead of just a prefix. This is probably the most secure, but can result in ugly client IDs shown to the user, like "example.com/oauth2/callback". Of course clients can control that and make it something prettier if they want.
2. Strip any query params from the redirect_uri, and document this behavior. That should handle most cases, but it's always possible clients implement an open redirect in the path itself somehow. You could also check for strings like "http", but at some point there's only so much you can do.
3. Require clients to implement client metadata[1], so you can get back to exact string matches for redirect_uri. This is a very new standard, and also doesn't work for localhost clients.
[0]: https://sec.okta.com/articles/2021/02/stealing-oauth-tokens-...
[1]: https://datatracker.ietf.org/doc/html/draft-parecki-oauth-cl...
I agree some nonzero number of users would click to continue when they shouldn't while doing an OAuth flow. Thanks for the example.
Tho if you do this then clients need a distinct client id per redirect uri domain, mostly this is non consequential and a good thing but I think it has some ux implications like seeing multiple consent screens vs just 1 if you had multiple redirect uris registered against a single client id.
Language-specific domains, preview domains, multiple properties (often not fully integrated) etc.
You can always paper over things with more redirects / params (so a landing that bounces on) or multiple clients but supporting multiple redirects can save the customer work in some scenarios.
Is it standardized?
We implemented (O)Auth for 250+ APIs with Nango and wrote about our experience here: https://www.nango.dev/blog/why-is-oauth-still-hard
Kinda dense and there's bit of jargon, but consolidates person-decades of wisdom around this topic.
0: https://datatracker.ietf.org/doc/html/draft-ietf-oauth-secur...
Plaid announced this level of integration with Chase in 2018. Unsure when it was implemented.
Perhaps US needs a PSD-2 like regulation to force innovation on this front?
I gave a very similar talk at Develop Denver back in 2018: https://www.youtube.com/watch?v=7rGkz9Ib4fQ
My slides even look similar to your diagrams: https://brianschiller.com/lets-build-oauth/#/31/10
If you really want to do that, have self hosting, contribution and other things front and center and don't take VC funding.
I'll quote myself from elsewhere:
> Both Zai and I care a lot about FOSS — we also believe that open-source business models work, and that most proprietary devtools will slowly but surely be replaced by open-source alternatives. Our monetization strategy is very similar to Supabase — build in the open, and then charge for hosting and support. Also, we reject any investors that don't commit to the same beliefs.
Fortunately, nowadays a lot of VCs understand this (including YC, who has a 10+ year history of investing in FOSS companies); we make sure that the others stay far away from our cap table.
The only reason the real open source is what the world is built on, is because there is a guarantee that they will never need an exit and the community will always exist.
1. Tens of kilobytes of JS that is executed exactly once, so is not amenable to JIT optimisation.
2. A strictly sequential series of operations with zero parallelism.
3. Separate flows for each access token, so apps with multiple APIs will have multiple such sequential flows. Thanks to JS being single-threaded, these will almost certainly run in sequence instead of in parallel.
4. Lazy IdPs that have their core infrastructure in only the US region, so international users eat 300ms per round trip.
5. More round-trips than necessary. Microsoft Entra especially uses both HTTP 1.1 and HTTP/2 instead of HTTP/3, TLS 1.2 at best, and uses about half a dozen distinct DNS domains for one login flow. E.g.: "aadcdn.msftauth.net", "login.live.com", "aadcdn.msftauthimages.net", "login.microsoftonline.com", and the web app URLs you're actually trying to access and then the separate API URLs because only SPA apps exist these days.
6. Heaven help you if you have some sort of enterprise system that the IdP needs to delegate to, such as your own internal MFA system, some Oracle Identity product, or whatever.
I've seen multi-minute login flows that literally cannot run faster than that, no matter what.
This is industry-wide. I stopped using chatgpt.com because it makes me re-authenticate daily (why!?) and it's soooooooo slow. AWS notoriously has its authentication infrastructure only in the US. Microsoft supports regional-local auth servers, but only one region, and the default used to be the US and can't be changed once set. Etc, etc...
(a list of things that are specifically bad implementations)
In my demos the OAuth flow completes so fast you can't even tell it happened, you don't even see the address bar change to the IdP the second time you do a flow when you already have a session there.
At scale, you can't put everything into one domain because of performance bottlenecks and deployment considerations. All of the big providers -- the ones actually used by the majority of users -- do this kind of thing.
This argument of "you're holding it wrong" doesn't convince me when practically every day I interact with Fortune 500 orgs and have to wait tens of seconds to a minute or more for the browser to stop bouncing around between multiple data centres scattered around the globe.
> you can't put everything into one domain because of performance bottlenecks
What specifically are you referring to here?
Just about every aspect of a CDN is very different to an IdP server. A CDN is large volumes of static content, not-security-critical, slowly changing, etc... Conversely the API is security-critical, can't be securely served "from the edge", needs rapid software changes when vulnerabilities are found, etc...
So providers split them such that the bulk of the traffic goes to a CDN-only domain distributed out to cache boxes in third-party telco sites and the OAuth protocol goes to an application server hosted in a small number of secure data centres.
To the end user this means that now the browser needs at least two HTTPS connections, with DNS lookups (including CDN CNAME chasing!), TCP 3-way handshake, HTTPS protocol negotiation, etc...
This also can't be efficiently done as some sort of pre-flight thing in the browser either because it's all served from different domains and is IdP-controlled. If I click on some "myapp.com" and it redirects to "login.idp.com" then it's that page that tells the browser to go to "cdn.idp.com" to retrieve the JavaScript or whatever that's needed to process the login.
It's all sequential steps, each one of which bounces around the planet looking up DNS or whatnot.
"It's fast for me!" says the developer sitting in the same city as both their servers and the IdP, connected on gigabit fibre.
Try this flow from Australia and see how fast it is.
The answer is to threat model.
What's the risk of an access token falling into the hands of someone who has got malicious JS into your webpage?
If it is a public API or not risky, maybe implicit is okay.
There's more on the attacks and mitigations here: https://fusionauth.io/blog/whats-wrong-with-implicit-grant
If malicious code on client is a big risk, I really don't see how the flow matters at all. Specifically if the only computer I have is on the client.
- use HTTPOnly secure cookies. These will not be accessible to malicious JavaScript but can only be sent to servers on the same domain. Well, I guess if there was an exploit that let JS break the sandbox and access cookies, they could be accessible, but I think we can trust the browser vendors around this kind of security. This approach is widely supported, but does expose the token to physical exfiltration (that is, if someone on the browser opens up devtools, they can see the cookie with the token in it).
- store the tokens in memory. As far as I know, malicious JS code can't rummage around in memory. This works for SPAs, but does break if the user refreshes the page.
- bind the token to the client cryptographically using DPoP. This is a newish standard and isn't as widely accepted, but means you can store the token anywhere, since there's a signing operation tied to the browser.
All of these can work and have different tradeoffs.
I will look into the dpop thing. A bit limited as I'm using AWS cognito.
I confess my thread model is such that I think I am fine with storing token in local storage. Would like to be using standard practice, though.
Sounds like you have spent some time thinking through your threat model and arrived at what works within the constraints of your system.
Here's a good resource on general threat modeling I found: https://github.com/segmentio/threat-modeling-training
I was hoping I had missed some updated guidance on how to manage tokens. Way too much of the documentation I was finding on OpenAPI and OAuth seemed to be aspirational and referencing things that hadn't come to be, yet. It has gotten rather frustrating.
Thanks for the link, it is a surprisingly fun topic to read on.
However, I think I would trust a value that's been transmitted from memory more than one that's been stored in localStorage but not transmitted, because the latter is trivial to grab with an extension.
I'm not aware of any way for an extension to grab a variable from memory unless it knows how to access the variable itself from JS. This makes me wonder if there could be a security practice where you purposefully only store sensitive data like OAuth2 tokens in IIFEs in order to prevent extensions from accessing them. There's got to be some prior art on this.
Anyway, thanks for bringing up the question. It's been a useful thought experiment.
1 - https://openid.net/specs/openid-connect-core-1_0.html#Offlin...
They at least have decent support to guard api access using these.
If my threat model gets to where I care about this, I suppose my only real options is doing the redirect to a compute backed address. For now, glad I don't have to worry about it. :D
The "modern" code flow involves a piece of backend code that performs the exchange of the code for tokens - and you are actually not supposed to make those token available to the (insecure) browser.
If you're starting out though, probably go for a SaaS in the beginning. But be sure to have monitoring for pricing and an option to close account creation, these things can become expensive fast.
Security. You can not deactivate certain unsave mechanisms. For example, if you send it an ID token, it will not verify the aid claim, allowing Anny valid token from the same SSO provider.
API stability. We're consuming their API from a mobile app. But every major version (about five a year) changed the REST API without backward compatibility or versioning. Its fine if you use their lib and keep parity, but that's really only possible on the web.
All of this was with their self hosted offering, I haven't tried their hosted one.
If you only need minimal auth functionality and you have one app, go with a built-in library (devise for rails, etc etc).
If you need other features:
- MFA
- other OAuth grants for API authentication
- SSO like SAML and OIDC
or you have more than one application, then the effort you put into using a SaaS service or standing up an independent identity server (depending on your needs and budget) is a better solution.
Worth acknowledging that auth is pretty sticky, so whatever solution you pick is one that you'll be using for a while (assuming the SaaS is successful).
Auth0 as a choice is good for some scenarios (their free plan covers 7k MAUs which is a lot for a hobby project), but understand the limits and consider alternatives. Here is a page from my employer with alternatives to consider: https://fusionauth.io/guides/auth0-alternatives
The downside is that we only support Next.js for now (unless you're fine with using the REST API), but we're gonna change that soon.
auth0 is too expensive for new SaaS imho
If users don’t verify domains, isn’t good old phishing more effective, like the $5 hammer in that xkcd?
(SSO also reduces the attack surface of phishing, though of course then the attacker just has to phish the identity provider's credentials instead.)
Still waiting for the cease and desist :)
And why do most of these “Auth0” replacements only implement nextjs?
Regardless, the focus of this blog post is on OAuth, not Stack Auth =) I appreciate the feedback though.
If you happen to be looking for future blog posts:
One thing that would help me and probaby a lot of other people is to show the flow using API calls in Postman.
Currently I am planning to reverse engineer example code in an unfamiliar framework...
My previous post was based on the phrase “open-source Auth0” (in your blog post) and we use Auth0 (and don’t like it). But all of out apps are react, not nextjs
we’ve had lots of folks migrate from Cognito to WorkOS. Lots of more features, modern API, and better extensibility.
More here: https://workos.com/docs/migrate/aws-cognito
Probably because it's hard to support different languages. Even if all the replacements support OIDC (not a given) there are still subtle differences in implementation and integration.
That said, check out FusionAuth! We have over 25 quickstarts covering a variety of languages and scenarios. (I'm employed by them.)
https://fusionauth.io/docs/quickstarts/
We're free if you don't need advanced features and run it yourself, as documented here: https://fusionauth.io/blog/fusionauth-on-fly-io