Keycloak with PostgreSQL on Kubernetes
blog.brakmic.com
blog.brakmic.com
At work, I ran 9 Keycloak clusters in production, handling tens of millions of sessions where the cost of losing sessions was high. The amount of time we wasted on getting it to work reliably with its default configuration of storing the sessions in its distributed, in-memory cache (Infinispan) is insane. It just isn't designed to handle such a work load reliably. Unless you're willing to spent months tuning it for every possible scenario, you WILL lose sessions.
If you are in this situation, shoot me an email. I have been through this pain and it took a lot of painstaking work to get to a highly reliable set up at scale.
Keycloak 17 made offline sessions lazily loaded by default.
[0] https://www.keycloak.org/docs/16.1/server_admin/#offline-ses...
Please! Post it, thanks
We were wondering about Redis in a similar IAM use-case (PingFederate) but it wasn't officially supported, so we decided to just go with persistent Postgres. I wonder if we saved ourselves a bunch of heartache.
As long as we wouldn't do any restarts, it would sort of work. Problems would pop up when due to high load, one or more nodes would become unresponsive and liveness probes would restart nodes. That would often cause the kind of cascading failure described above.
Most of these problems are also the result of running it in Kubernetes. We very quickly learned to remove the liveness probes and to massively increase the grace period. This helped, but only so much. We still had rather frequent failures similar to the one I just described.
Maybe if we wouldn't have run it in Kubernetes and we would be more knowledgeable about Infinispan, we could've gotten a stable set up. For us, as a small team without that specialized knowledge, we struggled to get a stable set up.
Whether it's process management or just say node having too little memory and spinning in GC too much.
Mixing app and DB (which is I guess happening here) also can be fun, as now app being overloaded can cause DB being overloaded. You'd probably be just fine if infinispan was used as a remote database instead of embedded one.
Newer keycloak versions (19 and up) have a configurable storage for the auth sessions (see storage-area-auth-session and storage-area-user-session). I haven't checked them but the documentation is promising.
For older session (last time I checked keycloak 15) you might want to use offline sessions but they don't allow SSO after the auth session was evicted from infinispan.
It’s a decent product but setting things up can be a bit dull.
One has to thread multiple Github repos and their incomplete scattered documentation together.
But once it works, it works.
> Features from the enterprise version are periodically moved to the open source version.
Instance also has a domain on top of that, but there are plans for a "simple" mode, assuming single org.
Since I somehow can’t resist complaining about absolute mana from heaven:
The only issue I see is this reliance on event sourcing — I get the reasoning but I much prefer regular state saving + audit log approaches.
Event sourcing seems like a complexity and performance liability — does anyone have any insight on the implementation/why I am wrong about my misgivings?
It also seems cockroachdb first, but I'm glad i can use postgres. One fewer database to deploy and manage, and for my use case (basically myself and occasional friends and family) that's perfectly fine.
Right, but it's like... why take that liability in the first place when you have a rock solid and extensible DB like Postgres under the hood.
Why not take the CQRS (good idea), but not go as far as full-on Event Sourcing, and just make sure you keep an audit table log or even executed operation log?
IMO in practice almost on one actually goes back in time with Event Sourcing. Also there are so many things can bite you it just seems unnecessary.
I did some digging through the code, and I really wish they'd made a big DB interface and then made the event store an implementation of that. It looks like they did it the other way -- the default interface being the event store, and PG/CockroachDB being the underlying. It's a subtle difference but means a huge deal for actual swappability of backends.
https://github.com/zitadel/zitadel/blob/main/internal/events...
I have to say, the code is also REALLY confusingly laid out. I just want to find the grpc/http handler that does like "create a user". I've been searching and clicking around for 10s of minutes -- maybe I don't read enough go.
> It also seems cockroachdb first, but I'm glad i can use postgres. One fewer database to deploy and manage, and for my use case (basically myself and occasional friends and family) that's perfectly fine.
I think of cockroachdb as basically postgres-with-stuff-bolted-on (albeit very good stuff, cockroach seems awesome), so I still consider it postgres-first! :)
There is an important improvement, though: the Postgres deployed here is not production ready (high availability, backups, monitoring, etc).
We run Keycloak on StackGres [1] which gives us production-ready Postgres setup (disclaimer: it's dogfooding). Happy to share the YAML manifests used to deploy Keycloak with StackGres. Maybe we will write a blog post as a follow-up to this one, for completeness.
[1]: https://stackgres.io
Another omission is that one could use a Keycloak operator instead of rolling custom YAML.
As a one-liner, though, for completeness: StackGres is fully open source (unlike Crunchy that needs a license for production); comes with a Web Console; 150+ Postgres extensions (including Timescale, Citus and many others); and many Day 2 operations fully automated.
Thanks for the hint.
I just disabled smooth scrolling. Sorry, not a sophisticated designer guy, using wordpress plus some UI themes.
Regards,
You will not be able to scale anything up anyway since it’s a single instance mounting the same data.
I don't think you're wrong. StatefulSets are simply on a "higher level" than raw Deployments and manually provided PersistentVolumes/Claims. I am thinking about writing another article that shows how to use StatefulSets in similar scenarios.
Regards,
Wow the learning curve was steep on that one. Not having ever touched OpenID or anything other than forms based authentication and not knowing ASP.Net very well.
But it's neat to get it all up and running. Still a few issues with getting Keycloak to redirect to HTTPS but we will get there.
That looks like the exact problem I'm facing. I'll try it out today!
Thanks again!
Everything else was going pretty smooth, although the authentication documentation for asp.net really sucks.
The documentation sucks for ASP.net and it's far worse for the Safe stack.
You have to understand the stack so you have to read up on the following.
ASP.net
Giraffe
Saturn
Fable remoting
Keycloak
OpenID
Once you have a good understanding of all of those you can start to understand the half a dozen blog posts that attempt something similar.Need to enable SaveTokens in session though, because you need the logout token for that.
If you have any issues, please ask and I will post some (short!) code snippets here :)
Ps: I also love f#, but I concluded that it’s not worth using it for asp.net. There are just too many f# specific things you need to figure out first. Just going with c# is the safer bet. But you can still use f# for your service layer and for tests!
I built my own auth, forms based with blowfish encryption is a few hours. Then I felt like I was doing it wrong. So I looked at OpenID. It's been two weeks and it just working. It'll take me. A few days to document it well enough to be happy I can keep it running.
I made a poor choice and introduced too much complexity.
How do people in the field handle configuration updates with code? For example, if I want to set it up as an identity broker to an idp, I would want that configuration backed by code, reviewed by my team. Is anybody using the keycloak terraform provider https://registry.terraform.io/providers/mrparkers/keycloak/l... in production?
Do people diff the realm json configuration as code and use that instead?
[0] - https://www.keycloak.org/docs/latest/server_development/#_au...
I am kinda curious though about the kind of personality type that enjoys this kind of stuff.
Of course, I have never heard of "Keycloak" before, so I checked their homepage:
"No need to deal with storing users or authenticating users."
Wait a second, is dealing with storing users and authenticaing them _so much pain_ that you rather inflict yourself with the pain of setting up and managing a k8s cluster?
I seriously don't get it.
What would you expect to see instead of a bunch of configuration files and cryptic commands?
The fact that you even need to ask ..
Like, it's so _obvious_ that any computer system can't possible be made to work properly without a bunch of cryptic configuration files and cryptic commands.
Indeed. The sad world of configuration files and cryptic commands.
Many thanks for the hint. It must have been a leftover setting from one of the other variants I am using locally.
Yes, postgres doesn't need to run as privileged. I changed it to "false" and updated the github repo.
Regards,