Okta Outage
status.okta.com
status.okta.com
why do i say this? At least as of Oct 2021 Okta didn't even have a complete compositional API to setup accounts -- thus requiring a 'josh-api', literally a guy named josh to manually provision new customer accounts by hand. the latency on those api requests was immense.
the company i was working at still depended on java 8, never budgeted time for refactor or maintenance, and had a plurality of other horrible dev practices they justified (calling themselves agile, devops, but not actually doing any of those things properly)
i was still in my probationary period when told me they were going to roll out Okta to all clients early in 2022 and charge for SSO so they could line their pockets, i gave my notice the next day (for many reasons, including okta). josh also gave notice and left on the same day.
given log4js, etc. this has probably been an extremely bad week there.
public String createAccount(String user, String pass){
var hashmap = readUserMapFromDisk(MYUSERS.PATH_TO_TXT_FILE);
hashmap.putIfAbsent(user, pass);
writeToDisk(hashmap, MYUSERS.PATH_TO_TXT_FILE);
return "All done " + LocalDateTime.now();
}EDIT: Just imagine the string args are in a more complex object that can be serialized to JSON
The control API (i.e. adding/removing roles, modifying policies, etc.) is available out of us-east-1. However, the bits of IAM that relate to distributing credentials to instances/tasks/lambdas and STS are all regionalized and isolated.
AWS is divided into multiple partitions. For the vast majority of users, there is one partition - the regular commercial - other partitions being China, GovCloud, etc.
Within each partition, there is a primary region that needs to be available for creation/mutation of credentials and policies. However, that data is replicated to other regions within the partition. That means the use of credentials that exist does NOT depend on the primary region being available. The replication is something that is closed monitored, and SLA breaches will result in pages.
https://auth0.com/availability-trust
And then read this tweet:
https://twitter.com/auth0/status/1471159935597793290
Edit: Ah, seems they picked us-west-1 and us-west-2 as the two regions..."In this case, we use two AWS regions: us-west-2 (our primary) and us-west-1 (our failover)."[1] So bit by a double-region failure.
[1] https://auth0.com/blog/auth0-architecture-running-in-multipl...
What else is everyone using?
Any thoughts on how to future proof this?
The notion of self-hosting is that you have to hire expensive operations staff and maybe you have sub-par experience due to lack of investment.
There is often an argument about "lack of core competency" too.
Facebook famously runs their own infra but that didn't work so well for them.
--
FWIW I'm actually of the other notion; I truly believe that you should minimise external dependencies. But that's because I'm a sysadmin (now: SRE) and it's my job to worry about reliability of systems. Less complexity and less external dependency ususally coincide with higher reliability.
A person could reasonably argue that it's in my interest to prefer companies run their own stuff, since it might be my job to maintain it, so it's self-serving. So I am not unbiased I suppose.
I saw many systems engineered for a lot of 9s be down hard for much longer than 9s promised, most often due to perfect storm of issues.
9s are great, and communicate pretty well what system was designed for, but they are in no way hard guarantee that the system will be up.
I used to work there and know the internals well. These aws outages must be causing massive chaos there.
MS doesn't do the things we need in a better way than other options, and it's almost always more expensive at product level and TCO level.
You hear people calling Microsoft expensive when they’re on some random mix of Gmail, Notes, CM9, or whatever.
Then MS seems expensive because it’s all or nothing. Dipping your toe in the water turns into a dive to the bottom of the pool.
A few people that really invest and enjoy a specific application does not make it great, especially when it turns our they are just doing more than they should be doing; i.e. when you have an InDesign professional that would be typesetting materials for publication but the person that writes the copy is also trying to 'typeset' the source in Word. It's great if you then feel like Word gets you cool typeset documents as a power user, but if 9999 people in a 10k company don't do that and just let the publication team do that properly in InDesign according to the media standards, it's no reason to keep it as a default available application.
A lot of the usage comes from "well, it was already there so I went and did it in that". Not because it was actually the standard, best choice or in scope of the task that was supposed to be done.
Same goes for things like notes and documentation:
- Code-level docs go in the repo (MD, RST mostly) - Org-level docs go in the wiki (Confluence) - Publications are delivered as copy to the publication team which then uses the DTP/typesetting thing of choice
Yet someone who would ignore that creates extra work by doing it in a different application first, then copying it around and converting it. That means that the person/process needs to be fixed, and doesn't mean we need Word as an expensive WordPad/Pages replacement.
Now, this might not apply to things like mini-orgs inside a bigger org, or very small companies and individuals. But I wasn't writing about those anyway ;-) At that level you don't really have the size and scope to make good choices anyway, and you're best off just sticking with one big vendor, not because they are the best, but because you won't be handling multi-vendor management anyway.
Well, you see... that's just it.
There is no direct substitute for Excel offered by any other vendor.
Similarly, PowerBI has no direct competition with even a tenth of the capabilities. Before then Analysis Services + Excel was absolutely the best, and nothing else could hold a candle to it.
Active Directory + Group Policy had no viable competition, and still doesn't.
For orgs that must be on-prem only, Microsoft Exchange only had Lotus Domino as a vaguely equivalent competitor. There are no open-source equivalents.
InTune + Windows + Azure AD + Hello 4 Business is hard to beat. You can assemble a mish-mash of vaguely compatible products, but it's a lot of work.
MS Teams is hard to beat for large enterprises because of the deep integrations.
Etc...
We have just as many hardcore Databricks users that wouldn't want to move, or hardcore MATLAB users and Mathematica users. We don't have anyone using PowerBI anymore, those all moved to Databricks and Tableau.
Active Directory and GP are a burning trashfire and only can't be replaced if you're stuck without modern MDM on a Windows Desktop construction. The only Windows we have left is VDI based on Citrix and Ivanti, everything else is "do whatever you want" where users have a BYOD choice with no internal access (and in reality they don't need it anyway) or VPN access with a choice of Kolide-based compliance or MDM-based compliance, and either are used to gate connections, in combination with DLP and standard anti malwares.
InTune sucks, Azure AD is nice, but when attempting to integrate with everything that already exists it sucks again. We have everything that is fully-managed or half-managed on JAMF and everything else is just isolated or internet-only (which was a requirement starting 2 years ago anyway). MS Teams never got a foothold here, sucks in so many ways people just rather have physical meetings or use email. Slack on the other hand works well and has been in constant use for over 5 years now. On-prem exchange was deleted and migrated to Google's thing a few years ago, works fine, does everything we need for quite a low price and quite high user happiness. It also ended up being the directory replacement. We no longer need Kerberos, except for some legacy Windows-desktop applications, but those are stuck inside VDI anyway until we can move on.
We used to have large VMware server farms and physical oracle boxes too. The first category was moved to AWS or replaced with microservices on Kubernetes, the second one migrated to Postgres RDS in AWS, but that is an ongoing process (roughly 10TB in table size done so far, and yes that did take tweaks like adding actual indices where Oracle would automagically do that on Exadata machines for you). Even with the human investment the cost is lower, productivity higher and we are less constrained by contracts. The best part is the elasticity we gained which was always problematic, even with the fake cloud (vmware-on-aws for example) or pretend-to-be-hybrid cloud solutions (azure) that never bear the fruit they advertise, or do but aren't actually better in the end.
Perhaps a big difference between the projects at this company and others is the vertical integration where things that are distinguishing to the company are pretty much created from scratch by internal development and maintenance teams. Things that don't matter or are shit no matter how you do it (looking at VDI) are delegated to MSPs but they have to run on our infra so we know the state of the infra no matter what the MSP is trying to tell/sell us. At the end of the day, this works great for us, everyone is happy, and a profit is made. This is in Western Europe if that makes a difference.
The main advantage is that the hardware token can be used in areas where mobile phones are prohibited, and of course immunity from a SIM swap attack.
Also, Yubikey would would not work, because like mobile phones, USB devices are restricted in some areas.
What kind of token would work, then? Something that only generates a TOTP, like those fobs some banks used to give out?
I don't know the actual answer to this question. My speculation is they don't want to have to support too many platforms, so they just flat out refuse to serve them.
We're in the middle of a migration from in-house auth (which we need to get rid off) to Okta and I think the people involved are finding Okta pretty confusing. But it's a big product and auth stuff is complicated, so I'm not sure how much it's Okta's fault.
My client uses it, it works mostly well. It does have its annoying limitations, though, such as no group inheritance and limited support for hardware tokens outside of Windows (no support on Safari/iOS, Safari/macOS, Firefox/Linux).
Okta does provide one at a reasonable cost, so it's easier to test with your new app deployments.
I haven't seen any sandbox feature either. The way we handle this is by creating a "test" app if it requires special rights, so we end up with SomeApp-test and SomeApp-prod. I don't know if there's any limit on the number of "apps" you can have.
By hardware tokens I mean U2F, in my case a Yubikey.
It's not 'u2f' exactly, but better (imho) for most people.
This is why I said support was "limited". Basically, only Chrome is supported cross-platform, which I don't use anywhere.
Should you? I don't know your situation and whether you can build an Okta-caliber level team internally. (My guess is that many smaller or non-tech focused orgs would have a hard time with that, but that's just a guess.) It's a hard question worth asking.
It's easy to think "we could have done better" when things are on fire, as opposed to all the times when the status chart is all green and you don't have to think about Okta (feel free to s/Okta/other service provider/) at all.
Disclosure: I work for FusionAuth, an auth provider that has both SaaS and self-hosted installation options.
Blaming Okta or any other group isn't the issue. Your customers don't care how you are down they only care if you are down.
Also, I got a bill from Amazon when I forgot to shut down a pagemaker instance and that cost me $700. I self host now buying a business internet package with a static ip. I also upgraded the machine but it wasn't necessary and in hindsight I shouldn't have done the upgrade but just fix the case.