The Difficulties of SAML Single Logout
wiki.shibboleth.net
wiki.shibboleth.net
In practice, logout in a federated environment seems to be
* A user logs into apps A, B, C (...)
* While using app C, the user clicks logout. App C trashes their cookies
* With luck(!) the IdP provides a low friction logout endpoint – think "GET /logout?redir=http://foo.bar/logged-out".
The app redirs to that endpoint, the IdP trashes cookies and directs back to the apps "you have been logged out" page.
Apps A and B? Go with god, and good luck. Keep your sessions short.
I asked about implementing logout cause I saw it was a thing in the SAML docs. The customer didn't care so our app trashed it cookies as usual and didn't report it to the IdP. It just wasn't worth the time.
I would log in to the main page where we can view government services. Then I would click through to the ATO (our IRS) to see some statements or whatever. I would then close the window and log out in MyGov then log back in with my wife's account. Click through to the ATO would show my details instead of hers. IIRC logging out wouldn't work either since that would mess up the session, I'd have to come back later.
Couldn't find anywhere to report this so I ignored it. I think they fixed it now.
https://www.smh.com.au/technology/revealed-serious-flaws-in-...
I was at Centrelink's IT department for just a couple months, and that was the only time in my decades-long consulting career that I had seen an adult man cry. Not for personal reasons. Work-related reasons. Several men, on multiple occassions, for different reasons.
That place crushed their spirits, and the tears just had to come out...
Imagine fixing an issue on a Model 204 mainframe, with a perpetually dwindling number of people you can rely on for help, knowing that if you don't fix it, people don't get their welfare cheque.
High stress, low reward. Unsung heroes in my books.
A "fun" anecdote I heard was that every 1 day of outage of Centrelink's mainframe would result in 5-7 children with broken bones.
You see, some alcoholic high-strung fathers in rural communities were spending their welfare cheques mostly at the local bottleshop. If the cheques stopped, their booze supply stopped, and their withdrawal symptoms would send them into mad rage. All too often, they'd take it out on their kids.
Apparently someone at Centrelink was tracking this kind of stuff by gathering paediatric admissions data from hospitals.
While I was there they had a 3-day outage that everyone just laughed off. You do the maths.
This one guy I knew told me that he had seen meetings dedicated to working out how they could spend money since they hadn't spent the whole budget yet. Services run over hundreds of servers when it could be done in one just to keep the whole team of system admins employed.
Government departments are filled with bloat to a level you wouldn't believe.
I think this is fairly common in government outside of Australia as well. With the budget it’s use it or lose it, so you better find ways to spend it, or you’ll have less margin next year when you might actually need it.
Almost all the higher ups are paid out of budgets based on how much you spend- so the supervising agency gets their 15 percent admin fee, your managers are part of the indirect cost pool etc.
If you spend 50% of budget it is game over upstairs. And you lose the money at year end, and future allocation is dropped. And all the budgets in long chain up get messed up. Almost nothing gets more attention and priority than this.
We've taken the approach of building a second API that integrates directly with the directory systems of record (Workday, Gusto, BambooHR, SCIM, etc.). These are often the actual "source of truth" and a more dependable source for the group membership changes that should trigger de-provisioning. But unfortunately even SCIM isn't well enough supported yet to be considered a base standard. (It's worse than SAML.)
Enterprise authentication and authorization is totally fvcked. Trying to fix it brick-by-brick.
But once I've given that information away, it's irreversible. You can't unsend data. You can ask nicely for it to be deleted, or revoke future access, but unless you have a time machine, federated SLO for all parties in a session ought to just be considered an impossibility, and planned for accordingly with short sessions.
In an enterprise environment all software has a basic level of trust of "this will try to do what it claims to do". That means the ID provider can directly ping the server of whatever service was issued the data or login cookies and say "this data should be deleted please".
And even if you do exclusively use SPs that support SLO, is the user supposed to wait while the IdP does all of that outreach in order to know that it worked, or 97% worked because one timed out? That depends on whether the user should even care about the outcome -- if they should care, then do they get an email when it achieves 100% hours/days later when that down-at-time-of-SLO SP is back up? Should they report an incident if they never see a success report?
So atypical that it borders on inconceivable. There's almost always an external party in there.
But if you somehow end up totally internal sure, try the back-channel binding. You get to toss SOAP messages around, and it might work. Of course now one of the invariants you built around no longer holds ("SAML works without having to let the IdP talk to the SPs" oh my sweet summer child...)
Just don't try SLO over the front channel, or else you'll have the joyous UX of a user clicking "log out" and all of a sudden they're bouncing logout messages back and forth between the IdP and two dozen internal applications. Hope the ninth app in line isn't down for maintenance or your <LogoutResponse> never makes it back to the IdP and the user wonders why they're staring at a server error page for an app they haven't used in hours and why half their environment is logged in and the other half isn't...
(Certainly that's not always possible, but for many systems it is).
It's not fun to administer.
The session ID needs to be cryptographically associated with the IdP so if you blow away the IdP session unilaterally you cannot decrypt or access any SPs session.
For instance but probably not sufficient as a solution, you could imagine IdPs running a JavaScript inside the SPs client session that provides half of a key pair and the SP providing the other half that combined form the session ID. Then once SLO is initiated the IdP script no longer provides its key.
I have not deeply thought further about how to completely design a flow like this but I strongly believe this kind of cryptographic session ID is the likely idea that will lead to a solution.
I made the switch out of IT to another field. Configuring a Shibboleth IdP was probably the hardest thing I ever had to do in my IT career, it really pushed my capabilities. SLO wasn't the only hard thing, the whole thing was immensely challenging and every time I'd restart Jetty I'd be holding my breath hoping it would come up again.
I love IT and programming, but I knew if I made a life long career of it, I'd grow to hate it.
I really like doing IT related things in my own way. I'd always want to do things to best practices and not just get the job done as I was told.
I implemented a SAML IdP [0] in MUCH less time than it took to configure Shibboleth. The specification for SAML is pretty easy to comprehend.
The implementation is really an experiment, but the configuration and usability is significantly better. Improving the implementation doesn't affect this. In some closed-source forks I've written a production version that's been in use for several years.
[0] https://github.com/rkeene/saml-idp/blob/master/lib/saml/saml...
Maybe your IdP expects SOAP over HTTP but your SP won't. Perhaps the SP insists on encrypting AuthnRequests. God help you if one side wants to do URL encoding and DEFLATE.
I've made my life easier by refusing to ask/answer questions around SSO and instead insisting on talking about "ADFS login". We still do SAML, but at least there's a baseline implementation that I can plan for.
This post indicates that 3p cookie issues only affect SSO providers without a hostname that lives within the primary websites domains: https://blogs.akamai.com/2020/01/cookies-single-sign-on-and-... (so SaaS offerings without domain masking). Is that your understanding?
Here's an interesting post on the samesite implications: https://www.troyhunt.com/promiscuous-cookies-and-their-impen...
We documented the login breaks from Samesite and 3p cookies, but logout was both not particularly a concern for folks (I've gotten about three inquiries in two years), and pointless to document tombstoning patterns for Samesite since ITP would break those anyway.
https://docs.microsoft.com/en-us/azure/active-directory/deve...
I did some research while back and found that shibboleth supports local storage for sessions [0], unfortunately the IdP+SP I'm using do not support such a thing.
[0] https://wiki.shibboleth.net/confluence/display/DEV/IdP+SameS...
Perhaps the safest way is for apps is to verify periodically if a SAML session is still valid, if there’s a mechanism to do so.
In general, that seems equivalent to the SP having a short session and the IdP having a long session. Suppose it's 1 hour and 8 hours. At 1h0m1s the next request to the SP causes a redirect to the IdP, and if the IdP hasn't heard about a sign out request, it just redirects back to the SP with a valid assertion that gives you another hour on the SP. No prompt to re-auth until the 8 hour mark.
Sharing of OS-level accounts remains an interesting challenge, so devices in those situations should involve identical account sharing at all application layers.
and I was looking around and ran across this post from the Shibboleth project.
Shameless plug for my startup. Hope that's still ok on HN :)
It's rare that a company will use multiple identity systems, but there exists lots of fragmentation across the large companies, which makes building a universal solution into your app very time consuming.
Active Directory: holds those accounts.
Okta/OneLogin: facilitates use of AD accounts to log in to webapps.
Here's our IdP docs: https://fusionauth.io/docs/v1/tech/samlv2/
Here's our SP docs: https://fusionauth.io/docs/v1/tech/identity-providers/samlv2...
I had independently concluded that the complexity probably wasn't worth it, but I hadn't considered shorter sessions as a mitigating factor.
Looks like the sands of time are the best solution yet again. I can easily spin 15 minute sessions as a superior alternative to SLO when talking to a customer. Determinism is the biggest point and I can play that against security & compliance very easily in my industry.
If you end up having to implement it, hedge with sessions as short as your compliance needs dictate. Expire your own session at the very beginning of the dance before you lose control.
And do not, under any circumstances, let them talk you into front-channel SLO. They probably won't unless they're totally clueless. But if they do it's the surest way to end up blamed for someone else's problems. Otherwise you'll end up with a support ticket some day that says "I clicked log out on bob1029's app and got a Datadog error page, what?!". And I'll smile.
This might require a trusted browse extension or browser feature to delete a "well-known" cookie across all sites.
It's fundamentally weird to ask for a feature "I want some website to end my sessions across all other websites on this device."