Sentry.io outage
status.sentry.io
status.sentry.io
There's many opportunities for failure even if only a single zone goes away; most (if not nearly all) database solutions elect leaders for example, and "brown-outs" (as in, not total failures) can lead to the leader maintaining leadership status, or at least messing with quorum.
other situations can exist where the migration out of a zone leads to hardware becoming unavailable for consumption for other people, after all, the cloud is not magic and if peoples workloads auto shift to the surrounding (unaffected) zones then it will impact peoples ability to do the same migration as all the free hardware could be used up.
I can think of dozens of examples honestly where even if you had built everything multi-zonal you could be down due to a single zone; for instance if some unknown subsystem was zonal (like IAM?) or you use regionally available persistent disks and now they suddenly perform extremely bad with writes because they can't sync to the unavailable datacenter.
I believe multi-zone is less possible than we would like it to be, there are many cases where you can commit no error but still be completely at the mercy of a single zone going away.
There are many understandable ways to accidentally have a single point of failure. But if your conclusion after the outage is that there was no mistake, you have made two of them, and the second is much less understandable.
They've been pretty solid.
I've been really happy with Sentry. No muss, no fuss as a vendor.
I wish they would buy/bundle a status monitoring tool.
So, I agree, but it's not like they really had a choice :)
Anytime this debate comes up there are two camps of thought -- the AWS camp where "measuring service degradation" is one of the most insanely complex problems that we lesser devs will never understand -- and the other camp is StatusGator a cheap service that seems to be easily able to tell you when one of your services you use is having trouble - including AWS.
<3
As it stands if their reliability leaves something to be desired: run your own.
These things happen, they’re more common the more complicated or huge your setup.
Good luck to them in getting it resolved.
Is there a good name for this theory that self-hosting something is more reliable? I mean yeah, if you have an outage it only affects you and your customers so I guess it's an improvement in that regard, but thinking you're better at running software than a SaaS is hubris.
The inverse is “appeal to higher power”; (implication being that its almost religion since it’s based on faith).
However: if you think that running large systems is hard, then it follows that it’s better to deploy small systems.
There can be times to outsource for sure, and I’m being relatively glib: but this idea of rent vs own is quite pervasive in the industry and it leads to dangerous situations where entire companies grind to a halt because these trade offs were talked down.
Regardless of that; understanding the trade offs is important, the famous “your nines are not my nines” is good to reference here: https://rachelbythebay.com/w/2019/07/15/giant/
Disagreed on several points.
First off, if you're using something like Sentry, chances are that building and running software is your job so it should not be a problem. If self-hosting Sentry is a problem I would start doubting the skills of your tech team.
Second, running a service for your own use is very different to running a service that has to scale to the entire world. Your service won't need as many moving & distributed parts as the SaaS Sentry since it will only ever have to handle a fraction of the latter's traffic.
Finally, if you're in control, you can schedule risky maintenance operations during times where a potential accident (operator error, etc) won't affect your business. You can't do that with an SaaS.
Sentry is a pretty complex piece of software with a lot of moving parts. They have good orchestration around it, either Helm charts for Kube deployments or a giant bash script managing docker-compose for mono-machine, and it just works. However understanding either of them, and being capable of debugging aren't skills i expect your average developer to have. It's more for SRE/Infra/Platform/etc. folk. Like i don't expect the average developer to be capable of debugging Kubernetes on their own.
Of course, many of them absolutely can, but many more would prefer that to people who do that for a living. The type of developer using Heroku and other PaaSes precisely to avoid getting their hands too deep into infrastructure stuff.
> If self-hosting Sentry is a problem I would start doubting the skills of your
> tech team.
I can't think of any team I've worked at in the past year for whom reading the self-hosted Sentry docs and standing up an instance would be impossible.However, my experience is by and large at early-stage startups, and so I also can't think of any instance where I'd want to have my team working on that when they could be working on adding value to our product. If I can pay Sentry to handle setup, scaling, maintenance, etc (and I do!), then that's worth it when weighed against the dollar and opportunity cost of having my team handle it.
That's not to mention maintenance or other issues.
We as an industry build our businesses on proprietary solutions that are difficult (nearing on impossible) or extremely expensive to self-host.
I love that there are options for SaaS; I strongly dislike being locked into SaaS -- and it feels like a lot of the companies I join are locked into some SaaS offering with no possible migration plan, despite dubious uptimes, awful support, rising prices and overall dissatisfaction with the product. (atlassian springs to mind immediately)
It's tech debt, of a certain type; and some tech debt is acceptable, especially in a startup.
Having an option such as Sentries to self-host is basically the best you're going to get... Gitlab is the same: Pay them to host, or host it yourself, the product is the same, the only "debt" incurred is the setup and migration, which is much lower than, say, Jira -> YouTrack.
I am however a proponent of self-hosting open source things you could get as a managed service. I.e databases, elastic search, redis etc. Provided you have the expertise to do that and benifit from it. For example, I used to run a redis setup that would failover in <5s on crashes and with basically 0 downtime for upgrades etc. Now we use ElastiCache and an upgrade will cause serveral minutes of downtime. I haven't had redis crash ever, but I have had elasticache & rds instances dissapear. Sometimes failover works, sometimes not and you have little to no wayt ot find out what is going on.
I always assume that running an on-prem version of a SAAS offering is going to be a shit show, but I'm curious if anyone here uses it?
I would like to use their hosted version, or a more stable paid on-prem version. Pricing is not the issue for me, but I don't want to introduce another third-party service and have to include that in compliance documentation for our customers. That is the reason I still self-host many things to be honest.
I imagine 99% of installations (and certainly CI pipelines) would be fine with their wsgi webserver and sqlite instead of Clickhouse, Relay, memcached, nginx, Postgres, Kafka and a host of other services. I wanted to take a stab at this myself, but given the complexity of the system, uncertainty of being able to merge it back, and, last but not least, their license, I decided against it.
Nowhere near as nice as, say, GitLab's Ubuntu packages which you can simply apt-get install, unattended automatic (security) updates are available and work well, and it's been mostly a set-up-and-forget type of thing. For GitLab, backups and most ops procedures are well documented.
Sentry, on the other hand, requires you to run a script to upgrade with their docker-compose thingies. You must remember to upgrade.
Self-hosted Sentry runs on docker, which on Ubuntu is a bit of a security shitshow what with docker side-stepping UFW and all that.
Recently Sentry tripled (seems like) the number of containers they spin up, and it seems much heavier than a couple of years ago. You can still get away with 8G RAM, so it's not really that bad, to be honest.
Backing up Sentry was simple enough. However, a lot of questions are answered in the style "if you want to know more, buy the enterprise version". But, I've backed up and restored a Sentry system and it worked out OK, so I guess we're fine. Unless they change something, and then we're hosed without realizing it.
So, Sentry's self-hosted version is OK, but clearly not their number one priority. I have no idea what Sentry would look like with a self-hosted distributed/HA setup. Our data protection policies would absolutely prohibit us from using their cloud services, so maybe I'll have find out, but I'm not looking forward to that.
When it works, it's pretty good. I'm using it with sentry-native on the application side, which uses Crashpad to capture stack traces of native binaries (x86/ARM, Windows/Mac/Linux, whatever). It often doesn't deduplicate events properly, and the stack trace qualities vary dramatically by platform. Sometimes the stack traces it provides are total nonsense, but it does allow downloading the minidump files, so I can dig at them in Visual Studio and see what's really going on. I have discovered and solved many real bugs using it, so I've put up with the frustrating stability issues.
[0] https://gist.github.com/tycho/4279ce2ca47b293a85696695968263...
Upgrading is a little bit tricky: in my (limited) experience the upgrading procedure for the self-hosted version sometimes fails and requires to take the application offline to perform the upgrade (but it is very possible that I may be doing something wrong).
It is really an invaluable tool.
Ideally, things run as a single binary with bring-your-own-database but that's not the case here.
I wish I could use my already installed ClickHouse to save some memory but the version constraint and some very complicated docker compose setup on their end wasn't worth trying.
But I still love the product itself, so I'll let it chew my memory.
Upgrade has been fine as the manual says, you pull their git repo, run the installer and worked fine on all my past several occasions.
You get used to the quirks of running it and then it just works, but it did take some time polishing the health checks etc to get it HA and scalable.
None of the helm charts I tried actually worked or provided HA.
Is that just a scaling thing for your company? Or are their 15 discrete services under the hood to drive it?
That said, it runs untouched for long stretches of time no problem.
A Sentry data leak could include valid access tokens to, say, all data on a customer's Docusign account. We try to make sure data like that is scrubbed before sending to Sentry… but mistakes can happen. This is simply not a risk we are willing to take.
Thus, we self-host Sentry in our own secure infrastructure (we are a SaaS provider ourselves), and accept all the maintenance burden that it entails.
There is an open-sourced option called GlitchTip (https://glitchtip.com/) that is much simpler to self-host. I believe they forked the Sentry repo before the license change. It's not quite as feature rich but is a pretty great alternative.
But my team settled on Honeybadger.io which is fine, but I still prefer Sentry.
I'm using Sentry to monitor logs and it now makes sense. Ok, so I remove Sentry from Laravel's error handler but nothing changes. And it's weird because sometimes it works, sometimes not. I tweak some things on Cloudflare (turning on Under Attack mode etc).
I still have this feeling that maybe it's related to Sentry: I'm using it and HN says it's down. So I go and remove the composer package as well, just in case. And it worked.
It took like an hour... I didn't use Sentry anyways :)
What did I learn: do not depend on external services if you really don't need it.
You have a great history of availability and would like to say congrats.
Not sure if their POP ingest servers are on it though.
This is fine.
In theory things shouldn’t break but if your CI process involves signaling a third party, it can fail from that (hopefully you can temporarily disable it or something of the sort)
I prefer a 98% accurate error reporting better than a 100% one I can't push to production.
Or maybe I'm not understanding well the value Sentry offers, of course.
Edit due to other comment: this assumes CI/CD. If you were deploying only weekly then this might be good reason to delay
I think if your release cycle is longer, then you will prefer not to deploy without sourcemaps.