PagerDuty (YC S10) Raises $90M at a $1.3B Valuation
forbes.com
forbes.com
I've been using pagerduty for years, and I've been getting other admins to switch from private emails to it, so no more single point of failures. They just had to be shown how simple it was to configure and use.
I won't wake up in the middle of the night to a text, but a phone call, wakes me up. I don't even answer, I refuse the call then open the pagerduty app. Love it.
Now if only I could get it to not leave a message.
Having set systems like this up for numerous clients, the latter bit is usually a case of waiting a few minutes for the on call person to either accept the alert, or failing that move onto whoever the fallback is for that person.
But I agree for e-commerce or mission critical applications - you need redundancy and it's worth ponying up the extra money for an advanced system.
I also prefer not starting my day with a fire. I'd rather do it off-hours.
They could of course just make the core product better, but thats not cool enough i guess.
Also, I’m a pretty light user of PagerDuty, but what has been getting worse and worse over time? From my usage it’s been reliable software, that lets me set and modify on-call schedules, and notifies the right people when things go wrong. I’ve been at a company that’s been using it for the past 6 years, and the core functionality has been strong/reliable the whole time, as far as I can tell - though again, I’m personally a light user, just on-call a few weeks a year and doing minimal administration. What are the big regressions they’ve had?
Would love feedback from folks on here about it or what they’d like to get out of a “core pager duty” product.
Is PagerDuty exactly "niche"? The core product - providing a third-party system that handles the mechanics of paging people during a production issue - is something that virtually all software companies could use, in theory.
It's hard to think of a SaaS company that's less niche than Pagerduty (again, focusing on the market size of the core product, not the actual market penetration or the current state of the product).
I love pagerduty. (the company, being point, on-call... a bit less)
I think "virtually" is a bit strong. There are still a lot of software companies that build software that doesn't rely on company-controlled production systems.
Also, in practice companies that have large-scale production systems tend to have dedicated 24-hour human monitoring.
Thus, I do think PD is hitting a niche within a niche: small-medium sized companies that rely on always-up servers.
I think the huge valuation has more to do with how absurd startup valuations really are[1], and less to do with showing how big the market for a niche product is.
[1] Behind the gloss of absurdly high startup valuations https://42floors.com/blog/startups/absurdly-high-valuations
Discussed here => https://news.ycombinator.com/item?id=6874838
https://saasholic.com/the-rule-of-40-for-saas-and-subscripti...
20% growth will yield a 2-3x ratio. 40% growth 4-6x, and by 80% growth you’re looking at a valuation of 8-12x ARR.
What Pagerduty can become, is much more than just alerting for IT. That's what the $90mm is for.
They hint a little in the article at what some of that future can include: "PagerDuty has shifted some of its focus to doing more with the data it tracks for customers in the future, an area in which it plans to invest."
ServiceNow is a pretty generic business process and service management solution, with modules for ERP, software development, HR planning etc. And maybe also a module for alerting and incident management.
I haven't used PagerDuty, but it seems it's a pretty focused and specialized application.
Is there any indication that PagerDuty is trying to be the generic business workflow modeling software that ServiceNow is?
So no, I dont think pager duty is going after hr & payroll functions. But the value/market theyre after is much more than just event engagement.
It's a valid opinion, and your faux-Socratic response came off as needlessly condescending, that's all.
And if you read June's outage report, they are in multiple AZs... and it didn't help.
Generally not banks per se, but bondholders definitely do. That’s PE’s entire business model.
If Pagerduty makes their money from startup VC money, could a similar thing be happening here?
Just to illustrate, here's the subheadline from a random company that sells a lot to startups (guess who):
Modern products for sales, marketing and support to connect with customers and grow faster.
And now the Pagerduty subheadline: Resolve and Prevent Business—Impacting Incidents Quickly for Improved MTTA and MTTR
I mean, seriously, "Prevent Business"? Why would you want to prevent business? This stuff is the insanity of million dollar contracts.(also they wear suits)
Also, the website seems pretty straightforward: if something bad happens on the servers, PagerDuty helps notify the right person to resolve it.
Still, the difference is enormous. Your explanation is so much better. Eg you used words like "servers" and "something bad" and "notify the right person". I still maintain that their landing page scares startup founders away and that we can safely deduct they don't make the majority of their money from VC funded companies.
I wish there were more creative pricing models to get us in the door somehow.
The problem is that our support structure, being a smaller newer company is that everyone grabs a paddle and rows. So for a 50 person company 30 are support on-call staff all with small windows of operation.
They will have a long-tail of free, and small clients - then a few whales and it seems completely doable...
Pagerduty is a great service.
Well, they're good at alerting. They're not great at incident revie/. The interface makes it pretty difficult to look at, say, "what are all the incidents I got paged for related to this schedule".
In general, the interface is pretty painful to use for large organizations. This is a fairly common theme, though....
How did you come up with that metric? Pretty much every company at that stage has negative income. If you mean revenue, that isn't accurate either. Even the largest tech giants are valued at ~8x annual revenue or more, and that number is higher for the SaaS sector. It is common for a smaller, fast-growing company to be valued at 15-20x ARR.
I'm not just referring to a lack of new features... it’s a lack of improving their core product. It's buggy and painful to use. I can’t tell you how many times acknowledging an issue (during the crisis around the issue) failed and caused cascading problems with other engineers being paged. That’s the one thing this product should do without fail.
It truly feels like 6 months were invested into a fairly basic CRUD rails application with a Twilio gem installed and then forgotten.
Good for them to have generated so much revenue with such a weak product. We paid them for years even after we removed everyone from rotation. I am sure lots of companies do this because it is the only thing out there.
What exactly does this mean? A "lack of innovation" doesn't tell me anything. It does what it says on the tin. That's usually what I'm looking for from a product like this.
Because the things we monitor are heavily interrelated, and an outage or degradation in one service will often affect others, it's not uncommon for us to see multiple alerts at once.
The typical scenario goes like this:
1. Pile of alerts goes out.
2. I acknowledge all of them in the course of acknowledging the first.
3. Some of the alerts will decide, despite having state-changed to yellow in the app, that they somehow haven't been acknowledged, so the app will issue the notification tone — for, again, previously acknowledged alerts.
4. Some more alerts will decide they still haven't been acknowledged, and will play the notification tone again, to remind me that I have outstanding, "unacknowledged" alerts.
All of this while I'm trying to fix the broken thing and am being interrupted to attend to the feels my alerting tool has about having alerts.
That is to say: PagerDuty is getting in the way of my fixing the outage. No, I can't just ignore them. We have a global team, and whoever's secondary is probably asleep, because we've taken care in crafting our on-call rotations such that people shouldn't get alerted at 3am. If they sleep through the alerts, they will escalate up my management chain. Then I'm having to explain to my director or VP why there are "unacknowledged" alerts — while I should be fixing the fucking fire.
It's an abhorrently crap user-experience, in a way that is antithetical to the tool's very purpose, and I'm not the only person on my team who has either wanted to throw, or actually has thrown, their phone through the nearest wall because of it.
To me, keeping the product stable and making sure that I never miss a page when I'm having downtime is by far the most important thing. I commend them for having designed and developed a system that appears to be one of the most robust services you could use. It means I'll continue to use them in the future without hesitation or lengthy debates on cost.
I help manage a fairly large on-call rotation and we've actively been finding gaps in their product where we think the product should be of assistance. Through support tickets or even hopping on a call with them they didn't seem all that interested in hearing about it or finding a solution for it. I mostly just got a sales pitch about upgrading our plan for new reporting features.
Some of the problem we've frequently encountered:
1. Anyone who is offboarded from the company will just get removed during the next LDAP sync. This just moves everyone up a day for their next shift, no notifications so the visibility is hard. If you try to remove someone manually there is a pre-requisite that they get removed from any schedules they belong to manually.
2. Overrides do not coincide with a specific shift, but a point in time. While I understand why overrides work the way they do, when combined with the problem outlined above, can be a real pain to fix the rotation if someone is taken out since there is no Audit Log.
Here is a scenario that happens:
- Person A is on-call on 9/29
- Person B is on-call on 9/30
- Person C is on-call on 10/1
- Person D has an override for Person B since they couldn't make their shift. Person B now does not have a shift until it goes through the rotation again.
- Person A is offboarded after the override was configured.
- Person B's next on-call shift was moved up one day to 9/29. Person B now is on-call again, with no notification.
- Person C's next is now on-call on 9/30, which has an override from Person D and now is not on-call.
As you can see, it would be beneficial if there was at minimum a notification of a gap in the rotation through automated means, or at least allowing some overrides be tied to a specific persons shift rather than point in time.
We built a tool internally that every few minutes keeps a copy of the rotation and the order of its participants and detects any changes. If it does, it will open a GitHub issue with the user that was removed and the position in the rotation they were at before being offboarded. I often put myself in that spot to preserve the rotation from any breaking changes until the rotation goes through at which time I remove myself.
3. Assistance to find someone to take your shift. Say you get scheduled for 9/30 but find out you can't make it, with one click of the button PagerDuty could email 4-5 people about your shift time and ask if they can take it. Someone can accept it and it'd notify the person their shift has been covered, and others that the request was sent to it has been taken care of. It could factor in any fatigue or length of time someone has been on-call before it includes them in the pool of engineers to cover the shift.
4. Per schedule notification policies. Anyone can change their notification settings or how they're notified. For one particular rotation, we'd like to enforce certain minimums to make sure the push notification is sent immediately, and the engineer is called if not acknowledged in 5 minutes. Currently we cannot enforce that.
5. Audit Log. Who added X to the rotation? Who removed X from the rotation? Who changed the length of shifts? I could add more. For Enterprise level software an Audit Log would be great. They mention they have one internally but don't have plans to expose it for customers.
While their API allows you to build all these tools, having it a first-class part of their product would be wonderful.
Fraction of the price, reliable, and not packaged with monitoring.
Here's what I'm looking for as a user:
- primary / secondary escalations pulled from a single roster without overlaps - add / remove people from roster without fucking up who's on call next
But it's pretty obvious they're more focused on listening to fence-sitters than paying customers. Which may be a viable strategy, considering their new valuation.
I think this is good news!
Edit: formatting, added more precision
It's an on-call management system that receives events from multiple sources and routes to the appropriate people.
It can then phone/text them when it gets an alert, and escalate to the next person in the list of it doesn’t get an answer.
It’s not super-hard to replace with some twilio scripts/etc but it’s cheap and usually works as expected.
It’s nice to pay them to run alerting scripts, so they don’t run on the same infrastructure as production code.
Deleted comment