3,735 karma · joined March 31, 2010
Building incidenthub.cloud
I run an outage detection service - and some of these issues, like parsing hundreds of - sometimes undocumented - status APIs, make for an interesting engineering problem.
Now if I don't see it on a page I check the page source - some blogs don't advertise the feed but it's there.
- New ideas - Easier way for new users to try out the product. It's currently a few steps, but I want to optimize it for specific user personas.
Thanks for explaining #2 - but the text "Auto" might be a bit non-intuitive for a user.
I understand this is an initial launch, so not trying to nitpick, but these are what felt like would add immediate value:
- Adding a social login like Google would allow faster signups.
- Maybe I missed it - but I could not find any options for automatic monitoring which would update the status page. There are manual options for adding an incident.
- Post-signup, it takes me to a login page. Why ask to enter the company name instead of user/pass directly?
- If I could click on "Components" on my dashboard (/admin) and go directly to the Components page, that would be nice.
I like it that you took the time to write a privacy policy that is both readable and precise, especially as to the list of third-party services.
My biggest technical challenge remains dealing with the immense number of different APIs (and not-APIs) in the different status pages out there. Marketing remains my biggest overall challenge as my background is engineering, but I've learnt quite a bit since I launched this.
Irrespective of his skills or how broken the system is, dishonesty can never be condoned.
It was born out of a personal need in past roles and teams. I launched it last year.
It was born out of a personal need in past roles and teams. I launched it last year.
We did face similar situations - but we fixed them after the cost went up on prod. I guess this has more to do with how much and how fast an "undetected" cost in pre-prod can explode in production. We used to keep an eye on the prod cost numbers after a deployment, and then tackle each one, because the increase was not that quick.
I'm not sure either about a "standard" way, so I'm just thinking aloud here, and I've not tried this myself:
For application changes, measure the difference in cost in pre-prod, in terms of percentage increase, between the previous deployment and the current one, and use that to estimate the possible prod increase. I suspect this will become messy very fast as the other factors to include would be num requests, CPU/memory usage, and so on.
Each application team should be able to view the total cost of running their service - and thus be held accountable to reduce costs when necessary.
Without data you are running blind. Cost optimization cannot be solved by a standalone team - it has to be owned by everyone.
Source: Personal experience reducing cloud costs in a slightly smaller team.
Not having MFA opens it up to potential data breaches causing havoc.
E.g. Getting alerted when my public cloud has incidents, my monitoring SaaS is unable to send alerts, GitHub has intermittent issues, Slack has trouble delivering messages, Stripe has payment failures for an edge case, and so on.
It was important to catch these issues before they impacted my apps. For paging services, I wanted to avoid potential disasters because my on-call teams were not receiving alerts (true story - AWS had an outage and PagerDuty was not sending alerts). For non-infra services like Slack, I wanted to know if there were message drops or delays.
There are RSS feeds for these services but I wanted a single place to monitor the ones I am using and be notified, rather than manually check 10-15 or more RSS feeds.
It's primarily meant for teams that want a simple way to setup such alerts.
I'm the solo dev on this project. I've been in backend development/ops most of my career, so my frontend skills are not great yet, which is evident in the UI :)
Would love to hear feedback.