SAML Is Insecure by Design
joonas.fi
joonas.fi
<claims>
<nameid>foo.bar@test.com</nameid>
<givenname>Foo</givenname>
<surname>Bar</surname>
</claims>
It instead encodes elements and attributes using elements and attributes, in the manner of the classic "inner platform effect" anti-pattern: <saml:AttributeStatement>
<saml:Attribute Name="uid" NameFormat="urn:oasis:names:tc:SAML:2.0:attrname-format:basic">
<saml:AttributeValue xsi:type="xs:string">test</saml:AttributeValue>
</saml:Attribute>
<saml:Attribute Name="mail" NameFormat="urn:oasis:names:tc:SAML:2.0:attrname-format:basic">
<saml:AttributeValue xsi:type="xs:string">foo.bar@test.com</saml:AttributeValue>
</saml:Attribute>
</saml:AttributeStatement>
This is roughly the equivalent of the SQL table schema where there's one table with the columns "rowid, columnname, value" along with a "schema table" that some hand-rolled code uses to validate that the data table contains only data that is valid according to the schema.It's not XML, it's XML squared: using XML to encode XML concepts instead of just using XML directly.
Btw I think your SQL example is called "EAV Schema" and does have a legitimate purpose once in a while.
A bunch of dictionaries containing UUID/data, where each dimension is its own dict, and the dicts might or might not be running on the same system?
I've always wondered why more people didn't actually issue DDL for these kinds of databases. Like if you want to let your user define their own custom fields, just take what they enter and actually CREATE TABLE / ALTER TABLE to build the the tables they need. And then more or less you could just write regular queries.
Turns out sysadmins get REALLY nervous when you tell them your application needs permission to use DDL at runtime. It required a lot more filling out forms.
I was talking about an application and its accompanying database, which would still be paid for and administered by the customer, not a database outside the application which it would be accessing.
That still raised many eyebrows.
If this is the main concern, you could have two databases (or schemas, if your database supports it).
One for application configuration/storage that the application never runs DDL on and another for user data, where the tables can be created/modified dynamically.
That separation means that, at worst, user data is screwed up if there's a bug in the DDL generation, but the application should still be able to run.
At least, that was my read.
Well, technically, the database's own system tables have tables with such column names and responsibilities.
Yeah, that's known as "entity-attribute-value model" (or just "EAV") and is considered an anti-pattern unless you actually need the flexibility it provides (which is fairly rare).
It doesn't have to be. We use this (under SQL Server) for our simple "custom attributes" mechanism (some irrelevant details omitted):
CREATE TABLE private_data.ATTRIBUTE_VERSION (
OBJECT_ID bigint,
ATTRIBUTE_NAME nvarchar(255),
ATTRIBUTE_GEN bigint,
IS_TRASH bit NOT NULL,
VALUE_BOOL bit,
VALUE_INT bigint,
VALUE_DECIMAL decimal(38, 9),
VALUE_DATETIME datetime2,
VALUE_STRING nvarchar(255),
ATTRIBUTE_TYPE AS private_attributes.ATTRIBUTE_VALUE_TO_TYPE(
VALUE_BOOL,
VALUE_INT,
VALUE_DECIMAL,
VALUE_DATETIME,
VALUE_STRING
),
CONSTRAINT ATTRIBUTE_VERSION_PK PRIMARY KEY (OBJECT_ID, ATTRIBUTE_NAME, ATTRIBUTE_GEN),
CONSTRAINT ATTRIBUTE_VERSION_FK1 FOREIGN KEY (OBJECT_ID, ATTRIBUTE_NAME) REFERENCES private_data.ATTRIBUTE,
CONSTRAINT ATTRIBUTE_VERSION_FK2 FOREIGN KEY (ATTRIBUTE_GEN) REFERENCES private_data.GENERATION,
CONSTRAINT ATTRIBUTE_VERSION_C1 CHECK ( -- At most one of the various value types can be non-NULL.
private_attributes.ATTRIBUTE_VALUE_CHECK(
VALUE_BOOL,
VALUE_INT,
VALUE_DECIMAL,
VALUE_DATETIME,
VALUE_STRING
) = 1
)
);
Since SQL Server encodes NULLs in a bitmap, having a bunch of NULLs for every non-NULL takes almost no extra space at all. But even on a "wasteful" database such as Oracle, this shouldn't add more than an extra byte per NULL.> The system does have type checking enforced by the application, but multiple applications access the data.
Traditionally, this is solved by funneling all applications through a public API of stored procedures and views/functions which enforce the desired business logic, but I understand this approach has fallen out of favor in recent years (decades?).
> attribute statement / attribute name / attribute value style
Does that even have a benefit? Is it any different that simply validating that a string is valid xml/json/something syntax?
E.g.:
<claims>
<nameid>foo.bar@test.com</nameid>
<givenname>Foo</givenname>
<surname>Bar</surname>
<extensions>
<myCustomClaim>010231239dfadsf</myCustomClaim>
</extensions>
</claims>
It's also possible to use namespaces for extensibility, and this can even be used to guarantee that the tag names don't conflict.So instead of this garbage:
<saml:AuthnContextClassRef>urn:oasis:names:tc:SAML:2.0:ac:classes:Password</saml:AuthnContextClassRef>
You'd have standard namespaced and extensible elements such as: <claims xmlns:saml="urn:oasis:names:tc:SAML:2.0:ac:classes" xmlns:adfs="urn:com:microsoft:ADFS">
<saml:nameid>foo.bar@test.com</saml:nameid>
<saml:givenname>Foo</saml:givenname>
<saml:surname>Bar</saml:surname>
<adfs:securityId>S-0-1-52534-12362341234-12312315</adfs:securityId>
<adfs:objectGuid>47e3a2d3-266a-48ad-b080-b785b4d0b45f</adfs:objectGuid>
</claims>
This would enable clients to validate the claims they do understand and ignore the claims they don't. Similarly to how X.509 does it, elements could be added to "must process" or "optional" parent elements to ensure correctness and security even in the face of a wide range of client capabilities.This is common in the XML world, but SAML had to invent its own wheels.
What we have now are groups like OASIS, the Shibboleth Consortium, eduGAIN, and InCommon, each defining the schema their members agree to support, and an explosion of competing standards.
By the way, even Microsoft can't interoperate with themselves. Traditional ADFS SAML can handle arbitrary multi-value attributes that Azure ADFS rejects.
In the context of SAML, these values are completely arbitrary metadata without fixed meaning in the protocol itslef.
You could even use attribute names that aren't valid for XML tag names.
It's the same reason why -- for example -- AWS APIs return tags as a similar entity list. You don't see
<Name>Example</Name>
<BillingGroup>prod</BillingGroup>
but rather <tagSet>
<item>
<key>Name</key>
<value>Example</value>
</item>
<item>
<key>BillingGroup</key>
<value>prod</value>
</item>
</tagSet>
---Database EAV would likewise be 100% appropriate if users could use the database to store arbitrary attributes. (Think JIRA custom fields.)
But if users are limited to certain attributes, then encode that in the schema.
Well, that's the problem though. YAGNI, KISS, and all that.
>The reason for that is that you still want to be able have an XSD to validate.
Well, you can still validate fine when you have specific fixed metadata...
Multiple exceedingly obvious vulnerabilities have been the result. One fun one was: looking at an XML signature in the document, verifying it, then ignoring the assertion it was claiming to sign and just trusting the assertion at the document root.
I tried to write a standards-based implementation and gave up. The standard is enormous, and consists of three parts:
1. The definitions of what each XML tag means in a vacuum
2. Patterns on how to assemble those XML tags into a document that means something useful
3. Protocols that exchange these documents back and forth to accomplish some authentication objective
Half the problem comes from the fact that it's meant to do anything and everything, and so you can theoretically just mix and match all the above parts to get what you want. But that also means that it's exceedingly simple to mix and match stuff in ways that are subtly (or not so subtly) insecure. The other half comes from the fact that the standard is so damned complicaed in order to handle everything under the sun that it's damn near impossible to wrap your head around it all. So people just glance at the spec occasionally and just write something that handles documents they see in the wild and hope for the best, with predictable outcomes.The whole thing is a tire fire.
Note, I last worked with it about a decade ago so I may have gotten some of the characterizations wrong.
Yup, and its complexity leads to something of a shortage of libraries in the first place so there's not much choice.
There's still no good modern maintained OSS .NET 5 library for doing SAML. There are some commercial offerings (ComponentSpace have a neat product and they weren't affected by VU#475445 either) but my worry is that folks will end up using whatever free libraries they can find without thinking too much about the consequences.
It checks that you get a signed assertion back, which is valid, and lets you pull out attributes. That's great, but there are a gajillion tags that can be set within an assertion that have genuine, important security implications and zero of them are handled by this library.
If your IdP sends those tags along with the assertion, this library will happily ignore them unless you implement support yourself. And I would bet heavily that approximating zero of this library's users write that support.
One trivial example, there's a `<saml:Conditions>` tag. This tag can have all sorts of attributes like `NotBefore` or `NotOnOrAfter` that specify a time period for which the assertion is valid. This library does not do anything to implement support for `<saml:Conditions>` tags at all, so every consumer of this library will happily process assertions as valid even if they're outside of the duration for which the IdP is claiming it's valid. If I get my hands on an assertion at some point in time, I can hold onto it and replay it forever to unsuspecting users of this library. There are other attributes that can set conditions, and there are other tags with implications on whether or not an assertion should be considered valid. None of them are handled.
Further, only minimal document validation is performed. The signature is checked, but the Issuer is not. Multiple Issuers can share a signing key, though I suspect actual cases of this are rare. At any rate, there's a billion ways users of this library can blithely chug along with an assertion that isn't valid and shouldn't be trusted, but nobody notices or cares because the happy path works.
With authentication, yes it's important that legitimate users are able to sign in. But it's as important that illegitimate authentication attempts are disallowed. And this half is completely forgotten about in literally 100% of the common SAML service provider libraries I've seen in the wild (which is admittedly not 100% of the ones available, just the ones I've looked at).
Edit: As @tptacek points out, every user of this library is trivially susceptible to AudienceRestriction authz vulnerabilities. An IdP can include a statement in an assertion that limits it to one intended service provider (e.g., "this document authenticates user FOO for Google service BAR"). Users of this library will simply see "this document authenticates user FOO" and authenticate that user, even though this document was never intended for them. Again it's been awhile so I might be mischaracterizing it a bit but that should be the gist.
FWIW, yep. It has no knowledge of AudienceRestriction:
$ grep -i AudienceRestriction Saml.cs
$(note that I don't know anything about SAML except how to configure certain applications for it).
I think it's funny because engineers are always portrayed as being grumpy and that they hate everything, but if you look at how they handle errors or security, they seem like the type of people that like everything and always imagine the happy path ;)
But I also think that latter part should be just given more weight in software engineering in general. A simple example is regex: most people tend to think about writing regexes that positively match expected input, but this generally results in expressions that match things it shouldn't. Writing them is generally easier and more straightforward when you think about not matching what they shouldn't!
Isn't this Postel's law in practice? https://en.wikipedia.org/wiki/Robustness_principle
There's a competitive advantage in handling bad input, right? I'm not saying it should always be the primary consideration and certainly not in security situations, but I think that competitive pressures are how we arrived at this situation.
A clawback of bonuses for security vulnerabilities might be an interesting idea? Or better yet bigger bonuses if there are no breaches.
Agency problem. Those who decide budgets and project timelines get the upside in such a bonus scheme, while whose who implement get the downside. Clawback bonuses, and the engineers not only get blamed, they get laid off to make the quarterly numbers for the management who want to make up for lost bonuses.
I suspect the solution lays in the spec complexity. Haven’t thought of this much, so my wrongthink knee jerk reaction is make a set of reference test harnesses part of the SOT of the spec, and that is the first step towards more robust implementations. Would like to hear how others have approached this problem space.
SAML is a way for exchanging 'security assertions' - for example, if a Alice logs into the AWS Console with her employer's single-sign-on service, the SSO service gives Alice a signed document saying "The bearer of this document can log into AWS as Alice within the next 60 seconds and access accounts B and C" which she hands to AWS. The digital signature on the document is very important.
Unfortunately, SAML uses the 'XML Signature' standard [1]
The people who wrote the XML signature standard didn't just take an XML document and whack a PGP signature around the whole thing! They thought about use cases like 'what if you wanted to only sign part of a document' and 'what if, after signing, the document was processed by something that reordered the attributes' and 'what if you have signatures from multiple sources and want to collate them into the same document'. They decided they wanted to support all of them.
As such, it's entirely legal for a document to contain a valid set of properly signed security assertions and a load of different security assertions that aren't signed.
So SAML documents have to be parsed very carefully - if you naively check the document has a valid signature then read the assertions, an attacker can add extra assertions outside of the signed part of the document, granting them broader permissions than they ought to have.
You forgot to add the part - what if someone wants to add comments randomly without it changing the signature, even when those comments (mildly) change how the xml doc is parsed!
(I like PASETO instead of JWT, but nobody else in the world does, it seems, so we're stuck with the broken thing forever. Client libraries at least special-case "none" now and make it auto-fail, which has probably saved a few people that didn't actually read the spec.)
The big problem here is that none of this matters. People do abuse OIDC as a general-purpose token for individual applications, but almost nobody does that with SAML. The point of these schemes is to federate authentication. In other words: it only matters what Okta and Google will agree to do. Your options are OIDC or SAML.
If you had to pick between the two, I guess you should pick OIDC, and think about what Ryan Sleevi said downthread about certificate verification; you'll want some kind of pinning scheme (pinning is a mess in client-server situations, but still sane in server-server applications; in fact, server-server pinning long predates browser public key pinning).
If you're not federating authentication, I don't think you should use JSON tokens of any sort.
Those use cases are valid. The second one is actually the reason why JSON can't be used for a similarly usage (there is not canonization standard, and reading then writing a file could reorder the elements order).
You can defend DSIG in a lot of different ways. I'll disagree with all of them and feel like I could probably win the argument; DSIG is a cursed standard. But that's not the bar you have to clear in this discussion; here you have to establish that DSIG is good for SAML. It clearly isn't.
The former might be important for something? But generally also very nearly impossible to do correctly. When that is the case, one should contemplate whether the actual business requirements can be fulfilled another way, instead of making things broken by design.
² "canonicalize-encode-sign" as opposed to the correct way of "sign blob, parse blob".
This is what we use in the military to provide provenance assurance, classification markings, and authorize dissemination destinations for digital intelligence products, and also what we use for manifests to enable trusted data transfers from unclassified to classified networks.
It's pretty interesting to be on a network where everyone has a PKI identity. Your client certificate gets sent to your sponsoring agency's personnel database, which returns the exact set of clearances you have, that gets sent to the server you're requesting information from, and you get back a web page with exactly what you're authorized to see and everything else redacted. No need to login and you can access any page you know the name or address for. You just may not see anything on it.
[1] https://www.dni.gov/index.php/who-we-are/organizations/ic-ci...
And frustratingly for me--trying to implement this in the course of a SASL method as a client--there's not even anything hooked up in SASL that might have hinted at the client what to do. Which is insane because the entire point of SASL is to bridge the gap of "how to request authorization given username, hostname, protocol." It makes me want to retreat back into Kerberos as a better way of supporting SSO than anything invented in the decades after it.
RFC6749 The OAuth 2.0 Authorization _Framework_
I threw out that first implementation and just got the 3 most popular servers and made it work with them.
Random example: I'd like to present directories and files differently so users can chdir to a directory. There was no single method for detecting a symlink to a directory that worked on any 2 of the 3 servers I targeted for support (other than issuing a CHDIR and checking for success, and I wasn't going to issue a CHDIR for every item found when listing a directory for performance reasons; I did have to issue a CHDIR for every symlink on one of the 3 though).
Clients can be much harder because you can any number of combinations of instructions thrown your way and you have to respect their intent. This is much harder to deal with.
Use few shot prompts with well put together examples, then iterating over every item in the specs, and finally running the generated code through a human review process - seems like you could shift a lot of effort towards validation, and in doubt so, devise a process subject to further automation.
Crawling the resulting codebase with various prompts like "this part of the specification could be improved by..." and then review a set of 20 completions for any real value.
This might work for xml,html, and so forth, but even if you only have a system that generates a baseline library for one standard, it could be powerful.
I can't personally fathom how throwing GPT at the problem will improve things, but I can trivially see how it could result in code that looks correct to the author but isn't to anyone who read the spec. Maybe that's a failure of imagination on my part.
Then again, I suppose that's the situation we're already in where well-meaning authors don't think to actually read the specs and so don't actually comprehend any of the details.
I don't know how GPT could help with that. If anything I would expect it to bias toward things it has already seen, which is the opposite of what you want when writing a new spec/library aiming to avoid past mistakes.
And as sibling comment alludes to, security implications are very frequently areas that require very deep knowledge both of the standard and of how breaks in security protocols usually end up happening. How many commenters on this article were aware of the potential for canonicalization attacks in security before reading the article or the comments here? I'd wager relatively few.
[1] If you're wondering how almost all binary searches could be wrong, the answer is they fail to account for overflow in calculating the midpoint.
Edit: I know the article talks about this, but my point is that his description of the round-trip instability isn't quite right. You can't change the bytes arbitrarily.
And somehow still manages not to include any helpful description of SAML.
Good luck, it's prolific throughout so many enterprises and academic institutions.
We might see some support start to materialise for OIDC as larger enterprises start to make use of e.g. AzureAD properly, but you're still going to see tons of universities especially running their own Shib instance. And larger enterprises are going to be difficult to shift off of SAML when they've had X years of experience with it and their supplier on-boarding processes are built around it.
The entire Jisc UK Access Management Federation is based on top of SAML, so that'll power the vast majority of academic resource authentication flows. You can argue that security is less of an issue for this use-case (particularly when e.g. Sci-Hub exists) but I think it's likely that once you have the infrastructure set-up for this it gets used for more and more services.
I'm sure a lot of governments have based regulations and standards on top of SAML... it seems exactly like the sort of thing they'd be drawn to - "mature", more complex than necessary, verbose.
Pretty sure Gov.UK Verify is SAML-based, though I get the impression the new single sign-on solution (now entirely in-house with no external IdPs?) will be OIDC based.
Peasy easy. Just start mounting SAML vulnerabilities attacks on all those organizations with friendly suggestions to fix their problems, on top of juicy demonstrations of dangers. After a few will be affected, a trend will appear - in all kinds of media, among other things - that one ought to move from SAML. Of course the teachings should be done remotely, from friendly jurisdictions.
/s
We have this problem with security for years, not, decades even. Vista (Microsoft OS) was universally hated - <appropriate xkcd picture with blinking Hitler inserted> - as it tried to address gaping holes in XP security; fortunately Win 7 was much more user-friendly, so security still had some wins. But not nearly "ultimate" wins, no.
Active Directory support both SAML and OIDC as of version 2016. Okta and similar support both indifferently.
SAML had a head start of a few years but OIDC caught up. It's helping that all social sign-on are exclusively OIDC.
Ignoring all other business constraints, the most secure architecture in my mind is integrated security w/ single-instance applications which are isolated to their own physical hosts.
Integrated security can include MFA as long as the additional factors are scoped exclusively to each applications' secure environment.
There is obviously a huge convenience speed bump if you leave it here, but there are small, incremental enhancements you can make to dramatically improve the UX without centralizing everything at the convenience of your attackers. For instance, you allow users to link accounts between 2 popular systems, and then rely on a simple pin code or similar for quickly jumping between them.
As for the problems caused by ambiguous representations, the most memorable one for me was the bitcoin malleability problem which boiled down to a thing people based a transaction ID on having multiple representations. Bitcoin used DER format precisely because it isn't supposed to allow multiple representations of the same thing mind you - that bug was created by openssl following John Postel's advice "be conservative in what you do, be liberal in what you accept".
The author constantly disclaims their approach. Why are they demonstrating in json, and why are they demonstrating in what they call 'pseudo-saml' when everything they need is right there?
They don't prove that SAML is insecure, because there is no vulnerable SAML in the post.
I wasn't sure what the point of that was either. I found that confusing and it seemed to detract from the strengths of the article.
I think OIDC is better than SAML (it avoids the particular problem with malleable signatures described in TFA). You can actually read the standard to understand how the protocol works, whereas SAML is a jumble of abstract nonsense. I still think OIDC is a bit too complex, and I wonder if it really needs all those different profiles (hybrid flow, authorization code flow, implicit flow).
Thankfully we've progressed from that era. s/XML/YAML
> The Central Authentication Service (CAS) was developed here at Yale University between 2000-2002. In 2003, a release of the server core code was refactored in collaboration with Rutgers University and in 2004 we collectively placed the code in the public domain under the oversight of Jasig (later Apereo).
v1 of the spec is a relatively short read, and relatively easy to implement correctly. I even wrote an implementation way back when with some homegrown changes (JSON for responses). It doesn't seem to get used much outside of academia but it's very minimal and understandable, and OAuth/SAML work in a roughly similar pattern (but different, CAS is centralized though there are proxies). It's now stewarded by a company called apereo[1].
Agreed with the other commenters though, SAML login is an "enterprise" feature, so it's going to be around for a very long time -- things that warrant upselling tend to stick around.
[0]: https://developers.yale.edu/cas-central-authentication-servi...
[1]: https://apereo.github.io/cas/5.1.x/protocol/CAS-Protocol-Spe...
It’s akin to saying “making a web browser is easy, just follow the spec”; but the spec is a sprawling set of standards, with many layers of revisions and many important details left up to implementations.
In both cases, implementations may not be nearly as cautious about the finer points of the spec. That’s a huge weight for either to carry, and probably the only reason we haven’t seen a similar proliferation of browser implementations with horrific vulnerabilities is that the spec is so huge and the incentives to build a new one are so low.
*clicks link knowing full well the travesty about to be witnessed*
So many businesses want me to support SAML. I always say no.
Assuming you now require that you deploy a solution which is free of SAML2, what do you deploy instead?
"Computed values" is almost as general as "unit". It tells me absolutely nothing.
One way to avoid this would be to require the input to be in some non-malleable normalized form: If normalization changes the document in any way then it fails the signature check. The advantages of this approach are that it doesn't require Base64-encoding the signed content and that it has no problem with embedded signature blocks.
This way you always accept as validated exactly the same input that is signed, and any normalization or processing is irrelevant¹. OTOH, once you do it like this, I'm not sure what's the benefit of allowing malleable data in the first place.
You may need to splice data back from multiple signed parts, but I'm not sure what else would be proper, if only parts of a document are signed.
-- Note: Only ever trust SignedData.
playWithData :: Dangerous -> Dangerous -> Dangerous
verifySignature
:: Dangerous
-> VerificationParameters
-> (SignatureInfo, Maybe SignedData)
¹ Normalization/processing will be exposed to attackers, so it still needs to be memory and computationally safe, and you should carefully evaluate do you really need preprocessing.SAML as far as I know doesn't specify how exactly an identity provider authenticates a user but only how, once a user is authenticated, the user has a specific "identity" in the context of the service provider that initiated the authorization/authentication process. Therefore the authentication mechanism on Google/Facebook's side can be OAuth or something else, but once completed, the mechanism to convey the identity of the user to the originating service is SAML.
All our enterprise customers on the Microsoft stack indicate SAML as the only viable option, whereas those on Google Workspace or on more custom IdAM setups in my experience don’t care if you as a vendor prefer SAML or OpenID Connect.
For OpenID Connect the developer has to sign up with Azure and have their app to the Gallery, you can not add a custom add yourself
Right now there are over 1100 Gallery apps using SAML, and only 500 using OpenID Connect
But yes, we don't support dynamic registration of apps for eg OIDC/oauth
I had thought I'd missed something when setting up some customers, and having to refactor to use SAML for SSO where we'd used OIDC for g-suite, but you're saying unless I throw it on the gallery (public app store?), It's only SAML?
Public it may be, searchable it does not seem to be.
Every search seems to direct me towards [1] , Which is about gallery apps, and then directs me to a Table of contents entry that doesn't seem to exist (the closest I found was a tutorials page which is about connecting to a bunch of pre-existing SaaS apps, not for a "custom-developed app")
...maybe this is something where I have to burn half a week playing with Azure AD on a trial account to figure out..
[1] https://docs.microsoft.com/en-us/azure/active-directory/mana...
this is the correct link if you want to develop something with the identity platform, the other link is more or less admin documentation...
Which is what I am talking about and ALOT of SaaS vendors do.
They have you go to AzureAD Go to Enterprise Applications, click "New Application" and then choose Application not found in Gallery.
When you do that, i can see no way to use anything other than SAML .
nothing in the Microsoft Docs, nor anything I can see on any portal gives the ability to for an Enterprise to Add their own Customer Open ID Connect application, you have to go through the process to add an App to the Gallery.
Will check out the docs and see what we can do, it's not good this was hard to learn.
basically the app registration process for openid connect and saml is basically the same one?!
When you do that, their are 3 SSO options for the new Applications, None of them are OpenID Connect
This is how 99% of all SaaS Vendors instruct people to add new Applications
My approach has been to use Keycloak as an identity broker. It's implementation is quite robust and supports a lot of flexibility in terms of mapping custom assertions and the like. But the actual application "only speaks OIDC" and relies on access tokens to be reissued by Keycloak.
The way it works in enterprise is that somebody wrote guidelines years ago that external software must support SSO with SAML. Then the guidelines were never updated and in some cases the company never realized they can support OIDC out of the box.
The exception is education, where Shibboleth is very entrenched with federations spanning all universities in some countries. Another exception for healthcare/defense that may not have updated any of their systems for 20 years, though they may not be customer if they have no internet connectivity and no SSO :D
Our desire to upgrade pretty much only comes from Architects telling us "thou shalt follow my shiny new standard" (and by the way, they read your docs; if you suggest SAML be dropped, they'll update their standards!). In that case we have to find time to upgrade, and of course we never just document how to do it once for our whole org, so all these engineers will be wasting time re-learning how to do the same upgrade. I'll bet you we'd save hundreds to thousands of hours per year by having really good migration docs. Even if your migration guide doesn't cover everything, they still take a significant chunk out of the time we need to figure it all out, and it lowers the mental barrier to the change.
I think the root cause of insecurity here is the near-universal attempt to use general-purpose XML libraries to build SAML. When the Go instability bugs were announced, I think there was a general take from the Go community that `encoding/xml` was not an appropriate foundation for SAML libraries, and, in this instance, I think the Go community was right: I think it should have been self-evident that you couldn't safely build SAML on `encoding/xml` (it doesn't even really handle namespaces!)
When I did my own SAML in Go at my last job, I wrote my own soup-to-nuts XML, including DSIG canonicalization. It was annoying, but you can't get SAML wrong; you need to be able to predict what every component in the system is going to do. What makes this worse is that most SAML systems defer the cryptographic stuff to libxml/xmlsec; for obvious reasons, some of them sane, nobody wants to implement DSIG themselves. But then they interpret the signed message in a general-purpose XML library, and now you have competing XMLs in a single system. It's bananas.
In the esoterica of DSIG, there are even worse problems. DSIG is a very flexible format; it has has pluggable canonicalizations, happily supports multiple signed subtrees under different keys in the same parent message, and supports both detached and embedded signatures. There's a famous, respawning DSIG bug where you can trick validators into verifying a signature on one subtree but passing on a different subtree to the calling application. A researcher at Duo Security tricked SAML implementations by embedding comments in text fields, which, after canonicalization, tricked parsers into returning altered text fields. Technically, SAML responses are supposed to be strictly schema-validated, so you can't sneak arbitrary subtrees into the middle of a SAML response --- but there's at least one extension field in a SAML message that has an any-typed free-form tag.
It's also worth remembering that SAML's problem domain is difficult even setting the XML part of this aside. It's pretty common to find very bad authorization vulnerabilities in multitenant SAML RPs, because you have to check mesaages carefully beyond their signatures. Many of the high-level OAuth2 protocol issues port over to SAML as well. It's tricky to audit!
People stick up for SAML because what it does (facilitating centralized SSO) is extremely valuable; it's probably more valuable than the bugs are harmful, as long as you're using well-known tools that people have already been incentivized to scrape for SAML vulnerabilities. If I was adding SAML support to something new, I'd consider beyond all the standard SAML checks also rejecting any message that doesn't have the same shape as what Okta, Onelogin, Google, or Shib generates.
But if you have the option, I'd also say avoid SAML.
Kerberos is limited to internal network and some very specific use cases (desktop auth). It's not competing.
If the company has fully integrated Active Directory/Kerberos. On any desktop computer, it's possible to get an OIDC/JWT token for the current user with a single API call. It's transparent, the user doesn't need to enter their password because they are already authenticated on the machine. That is to say, no application ever needs to support Kerberos in the current age.
We implemented our own SAML processing library, too: https://github.com/FusionAuth/fusionauth-samlv2
(We pay for valid security bugs.)
My company needed to implement SAML SP support in one of our products so we could get academic customers, particularly those that are part of the InCommon federation. We contracted with a company that specializes in SAML and Shibboleth to help us get it right. We decided to use Shibboleth SP running in a container; that container also has Apache httpd (as practically required by Shib SP) and a little Python shim app that generates a JWT and passes it back to our main app. Hopefully that's a good way of using the nearest thing to a canonical SAML SP implementation, without running our whole application through Apache httpd. In case anyone's interested, our Shib SP container setup, with the Python shim app, is here:
https://github.com/PneumaSolutions/shib-sp
It's probably still too specific to our application, but might be useful as a starting point for others.
Things get dicier when you go to languages like JavaScript, where there aren't really well maintained SAML implementations. But then that's true for everything.
Elixir for me. That's why I ended up running Shib SP with a Python-based Shim app in a container.
However, the “what does the future hold” of OIDC is not much brighter. OIDC is based on JSON Web Tokens (JWT), which manages to avoid some of these issues (e.g. signs the encoded value), but introduces new ones (JSON interpretation bugs, algorithm substitution bugs, etc). They’re similarly terrible by design [2].
However, what OIDC does relating to signing is far worse. In many OIDC deployments, the idea is you use something called “OIDC Discovery” [3] to discover the expected signing keys for the OIDC server. You fetch those regularly (e.g. daily), and do so over TLS. With SAML, you exchange certificates, and then rotate them every 2-3 years (with things blowing up on expiration), but with OIDC, you often end up using OIDC-Discovery, and thus can change keys daily.
This means that a single malicious TLS certificate can be used to MITM your OIDC Discovery exchange, and from there, impersonate any user from the identity provider to your system, the relying party.
I spend my days in the TLS trenches, working to improve the CA ecosystem, but I absolutely would not trust the security of all users to a TLS certificate. The reality is that BGP hijacks are still a regular thing, as are registrar hijacks. Even if you find out about a malicious certificate (via Certificate Transparency), and revoke it, virtually none of your tools doing the OIDC-Discovery fetch (like programming languages or curl) support revocation checking, and even if they did, it doesn’t work at Internet scale. To deal with this problem, some relying parties do a form of poor-man’s certificate pinning, but now they’re at risk of even greater operational failures than SAML expiration in the start.
In practice, it seems plenty of OIDC clients just shrug and go “yolo” - if they’re talking TLS to the IDP, that’s good enough, and no need to bother with signature validation of the assertion at all.
For all my hatred of XML DSig and SAML, I’ve seen few auth standards as bad as OIDC: because it looks good, but is hell to implement correctly. At least with SAML, you know it looks bad to begin with, so you’re hopefully on guard.
[1] https://www.nccgroup.com/globalassets/resources/us/presentat... [2] https://news.ycombinator.com/item?id=14292223 [3] https://openid.net/specs/openid-connect-discovery-1_0.html
I would bet a lot of money that a non-trivial number of people do exactly this in the real-world using SAML (Shibboleth: FileBackedHTTPMetadataProvider or DynamicHTTPMetadataProvider). It's not always manually managed.
https://www.usenix.org/system/files/conference/usenixsecurit... (from 2012)
We've implemented SAML for 10s of millions of users and devices. The spec is verbose, but the approach solves common business issues. My suggestion is to use SAML simply: federations and passing attributes between trusted parties, allowing to verify the payloads. SAML can do a lot, but keep it simple and use OOB services for more orchestration/metadata.
->
SaaS and other types of companies start solving this issue at scale
->
SAML becomes code that is mostly going through a handful of large companies that sell products in this space
->
since the protocol is now centralized, it can be updated to a better protocol
- that library is not supported as such. It was published under the DVV but that’s it.
- my understanding is the SAML is legacy they are trying to move away from. This is critical gov infrastructure so I doubt they take it lightly and I would presume it’s been incredibly heavily pen-tested.
- integration was time consuming and difficult to debug errors. Would not recommend.
Ho ho ho, they really haven't looked at a lot of XML.
SAML has been a game changer for us. We are a SAAS business and 80% of our help desk tickets were people who could not log in. SAML has largely fixed that for us.
Getting rid of SAML will be like getting rid of SMS 2FA.
It would therefore be comparable to OpenID Connect's authorization code flow.
Edit: it’s the audit assertions proposal for anyone interested. I consider it so high risk that I’m actively working to move projects in my purview off NPM in case it lands.
Signing the normalized content is of course a nightmare, given that there is no spec'd normalization.