Show HN: Diahook – Webhooks as a Service
diahook.com
diahook.com
Diahook makes it easy for developers to send webhooks. Developers make one API call and we take care of deliverability, retries, and offer a great developer experience for their users. Essentially, we make it possible for everyone to offer a Stripe-like webhooks experience.
At my previous company, our users were constantly asking us for webhooks, both for consuming in their own services and for integrations with no-code solutions like Zapier. However, we kept on deferring building them because we weren't willing to commit the engineering time, resources and ongoing maintenance required of a webhook delivery system.
There are a variety of challenges when it comes to sending webhooks. For example customer endpoints fail or hang much more often than you would think, so you need to implement retries, but also make sure that such failures don't slow down or block your send queue or the rest of your system. Additionally, because of how webhooks work, anyone can send fake webhooks to your customers, so you need to make sure to cryptographically sign the payload, and make it easy for your users to verify it. You also want to avoid overloading your users' endpoints, so you want to automatically rate-limit webhook sending, as well as disabling failing ones, and notifying your users when you do.
I love webhooks, and I think everyone should be offering them! Our goal with Diahook is to make it faster for developers to add webhooks to their service, take care of the above challenges (and more), relieve them of having to worry about maintenance and scaling, and offer their users a UI for inspecting, debugging and replaying of past webhooks out of the box.
I'd love to hear about your experience building (or using) webhooks systems. What's a must have? Any war stories to share? Got any questions? Suggestions? Please let me know!
Docs: https://www.diahook.com/docs/
API viewer (and OpenAPI specs): https://api.diahook.com/docs/
Don't developers need to handle deliverability and retries to Diahook?
User endpoints on the other hand, fail all the time, and often require a few retries.
With that being said, we plan on having client side (in our libraries) redundancy (try another endpoint if one fails), and ways for you to gracefully handle errors locally for users who need this.
Edit: I realise this comment was a bit sloppy, added some context in another comment https://news.ycombinator.com/item?id=26400970
My suggestion would be not only a library, but a local queue that can be managed with the library. Also, having some experience with working at a no code provider previously, ensure that it's trivial to query your service for what webhooks you can confirm were received by your infra and then delivered, perhaps by UUID.
Last, make sure it's dreadfully difficult to make changes to your systems where data loss could occur (messages dropped while returning 200 to counterparty systems, for example).
I agree about all the issues you outlined though. One thing we plan on doing in terms of internet being down and cloud provider not having routes: have a lot of endpoints spread geographically in every region, and maybe in the same/nearby data centers.
Thanks again for the feedback!
We just send the same requests over and over until they're successful. It's relatively simple to use.
Edit: ahh, I already mentioned it, it's buried in my comment.
That doesn't sound realistic to me. Sendgrid, Twilio, and all other APIs have downtimes and other problems, just check their status page: https://status.sendgrid.com/history.
I always had to implement some retry logic in API clients at some point or another. That doesn't mean your service doesn't bring value but it's not enough to just expect it to work because it is focused on uptime. When implementing a service that requires external dependencies I would always expect them to fail and try to design with that fact in mind.
We already have users that have their own queue system for sending us messages because that's how they deal with similar APIs too.
So I definitely see the problem, what I meant in my comment is that it's similar to other API companies, and it's not unique to us!
I work in healthcare tech, and in the last couple years I've worked with legacy medical systems that use FTP to implement "pseudo-webhooks". For each integration with a third-party technology provider, there is a dedicated FTP server that reads/writes files to an FTP server owned by the third party. When this service is interrupted there is no retry logic - someone has to spend time manually checking if files made it across.
If someone could make this legacy process easier on developers and HIPAA compliant they would be rewarded handsomely, as these organizations (hospitals, medical device makers, health-tech companies, etc) are comfortable paying enterprise-level prices for any technology they use.
Is this what it would look like? https://docs.mulesoft.com/ftp-connector/1.5/
If there’s any links or vendors you could share so I could learn about this I’d greatly appreciate it.
I will say there are EMRs like Athena Health which actually build for the developer experience. Integrating with Athena is relatively easy, and you can scale up the integration across customers easily. Epic (the market leader) is much harder and more expensive to do this for, since every distribution usually has some customization.
Here's the only post I wrote for it that focuses on security, which is pretty critical for a webhook system https://www.easywebhooks.com/how-to-secure-a-webhooks-api. You might want to consider adding protection against Replay Attacks and support for Challenge-Response Checks if you do not already!
Thanks for the link, will read it tomorrow!
Basically you periodically send a GET request to the target API with a token, and have them respond with the token encrypted with the same secret they'd use to decrpyt your webhook signature.
You could also consider sending dummy ping messages that may or may not have a valid signature (of course make sure this behavior is documented) that you would expect the target API to return a 4xx error if the signature is incorrect.
These extra steps are definitely not table stakes for a webhooks system, but could be enough to make sure the webhook event providers are being the best possible stewards of their user's data that they can =D. A lot of this complexity can also be wrapped by a client library you provide, which is a big win for everyone on its own.
Twitter's webhook API has an example of CRC btw! https://developer.twitter.com/en/docs/twitter-api/enterprise....
My employer (Segment) has something similar to Diahook (as I understand it) at its core: https://segment.com/blog/introducing-centrifuge/
Thanks for sharing this though, a cool read for tomorrow!
It turns out my interview question is now a SaaS!
Kudos on the launch OP. This is a much needed service and it’s one of those things that once you see exist you go d’oh, why didn’t we have something like this all this while.
Thanks a lot, yeah, there are so many moving parts and pitfalls. I think that what makes it a good question, also makes it a good service (I hope).
It ended up taking us one afternoon (this afternoon in fact!) to integrate Diahook. Was super straightforward and Tom was highly responsive.
There are other features/reason that would make me use a "webhook service" and I've contemplated building one before but mainly for async message passing between systems. Anything sync and you almost always need to take care of all the failure states.
This could be useful as a library or side-car. I can't help but think of this as a plugin for something like Envoy even though Envoy probably does a lot of this already.
I think the same can be said on every other API company, including Stripe, Sendgrid, Twilio and etc. If this is a concern for you, you need to take care of that, if it's not a concern for you, then Diahook is not any different.
Anyhow, we take of other things other than just deliverability. We sign the webhooks for you to prevent SSRF, we implemented a retry logic, we have monitoring on all of these, and all of the other things I mentioned elsewhere.
I mentioned it elsewhere in this thread, I don't see how this can be a library. It's a standalone service that needs to be run and monitored, needs to scale with your usage, and you need to make sure it doesn't hang, the queue doesn't get too long, and etc. I actually say it in the post, webhooks aren't as simple as they seem.
Really? I thought this would be covered under most SLAs.
> It seems to me that people saying your service is also potentially unreliable are missing the point a bit.
The reason I (and _I guess_ other people) are thinking of this is that I read not having to invest significant engineering resources into building a robust failure resistant system as one of the main, if not the main selling point and I don't see how that is possible if I have to build it for Diahook in the first place. I understand Diahook has a great SLA and promises to be up as much as possible but things will always go wrong and there are so many points of failure outside of Diahooks control. So as a service maintainer, I still need to invest the same amount of engineering in the component that calls the webhooks.
SLA is a business thing, not a technical one. When I evaluate multiple services and compare their SLAs, I see the probability of one service causing less disruptions to my business than the other. I don't see that as a reason to not write fault tolerant code. Engineers cannot rely on a higher SLA as a reason to throw fault tolerance out of the window. That's my main issue with using a service like this. However, there can be other features that could still make me sign up. Not having to write fault tolerant code just isn't one and anyone buying into that is shooting themselves in the foot.
Nice. These are very useful feature and I can see myself subscribing for these. May I suggest you highlight these as the main features and demote, remove or re-word "not having to invest significant engineering resources in making a reliable webhook calling system" bit (not exact words). Just a suggestion, don't claim to know more about your business than you but as a possible customer, this claim being at the top would throw me off TBH.
> I mentioned it elsewhere in this thread, I don't see how this can be a library. It's a standalone service that needs to be run and monitored, needs to scale with your usage, and you need to make sure it doesn't hang, the queue doesn't get too long, and etc. I actually say it in the post, webhooks aren't as simple as they seem.
They definitely aren't and I know that from building some quite non-trivial ones and I did build some of those as libraries used across projects in the same company. Also, wouldn't my scale be the same irrespective of whether I call Diahook or anything else directly? I still need to make same amount of calls, have same queue, same retries etc.
> I think the same can be said on every other API company, including Stripe, Sendgrid, Twilio and etc. If this is a concern for you, you need to take care of that, if it's not a concern for you, then Diahook is not any different.
This should definitely be a concern for everyone unless missing outgoing webhooks is fine for a service and I agree Diahook wouldn't be any different which was my original comment. Sorry if I came across as dismissive. I was just pointing out how using something like this cannot be a reason to eschew fault tolerance in outgoing webhook code.
> Nice. These are very useful feature and I can see myself subscribing for these. May I suggest you highlight these as the main features and demote, remove or re-word "not having to invest significant engineering resources in making a reliable webhook calling system" bit (not exact words). Just a suggestion, don't claim to know more about your business than you but as a possible customer, this claim being at the top would throw me off TBH.
The landing page needs improvements. I agree that this can come across a bit wrong.
> They definitely aren't and I know that from building some quite non-trivial ones and I did build some of those as libraries used across projects in the same company. Also, wouldn't my scale be the same irrespective of whether I call Diahook or anything else directly? I still need to make same amount of calls, have same queue, same retries etc.
The webhook system is another system you need to scale (including monitoring), it's not the same as your main system. You need to make sure that your queue and workers can handle the load, monitor backlog, and etc. I don't think it's quite the same.
> This should definitely be a concern for everyone unless missing outgoing webhooks is fine for a service and I agree Diahook wouldn't be any different which was my original comment. Sorry if I came across as dismissive. I was just pointing out how using something like this cannot be a reason to eschew fault tolerance in outgoing webhook code.
I was definitely too sloppy in my original comment, I didn't try to eschew fault tolerance, I know how important it is! What I was trying to say is that for people who care about this high level of fault tolerance already have systems in place for the rest of the APIs they use, and the people who don't, don't. I don't think it's substantially different to other critical APIs in that sense.
1. How do you deal with endpoints that are down or 500'ing? What kind of retry policy or backoff occurs? Related, how do you notify clients when their endpoints are having trouble? (I ask because that's PII that I then have to share with you).
2. Is there support for message signing (e.g., HMAC) to let clients verify that the webhook really came from us? How does it work?
3. Any kind of deduplication support? One use case we had was that certain webhooks only required the latest delivery. e.g., product inventory. If we previously failed to deliver a webhook, but have a newer version pending, our next attempt should try to send the newest version.
4. Is it possible to delay hydrating data in the payload? It seems the expected usage is that I send a blob of data and you take care of the last mile and send it. Is it possible to send you an id and then you call back to my service and fetch the latest version of that object just before delivery?
5. How are webhook subscriptions actually managed? e.g., my app lets users register webhook URLs? How do I get those URLs to your service?
6. The big question: What do I do if your service is down?
1. Exponential backoff. We don't notify clients, but we notify you with a webhook. We actually don't have it implemented yet, but will have in the next week or two.
2. Yes, and libraries for some languages (more coming) to easily verify it. See the docs for how it works, but standard stuff. We also plan on offering end-to-end encryption in the very near future so that we can't even access your payloads.
3. No, and no plans for supporting it for now. Does anyone do it? I think sending all of them is actually a feature.
4. Could you elaborate? You mean changing the webhook content in each attempt? No plan at the moment. Similar to the previous point. We consider webhook payloads as immutable messages.
5. You can either use our API to add them, or redirect users to a management UI we built with a one-time password. We will also offer JS libraries to make it easy to build your own UI.
6. I answered it a few times in this thread. Here's my original sloppy comment (with a link to a comment with more context): https://news.ycombinator.com/item?id=26400510
Build that and people will flock to you.
There are many industries where push APIs haven't caught up and are much needed.
Let me know if I can help akshay@terminal49.com
Here's an example query that will poll an npm package and send a payload whenever the latest version changes.
https://www.onegraph.com/graphiql?shortenedId=FYH1SY
subscription PollNpmPackage {
poll(
webhookUrl: "https://example.com"
onlyTriggerWhenPayloadChanged: true
schedule: { every: { minutes: 60 } }
) {
query {
npm {
package(name: "graphql") {
distTags {
latest {
versionString
}
}
}
}
}
}
}
This is just an illustrative example. I used NPM because it doesn't require any authorization. We also support proper subscriptions for NPM that listen to their couchdb change feed.There are several reasons I believe we weren't successful with it:
1) There is a perception that webhooks are "easy", and so developers might not even think to look for a solution. After all, anyone can send an HTTP request. The devil is in the details of course, and a good webhooks system is actually a lot of work.
2) I believe most devs looking for a "push" solution are usually looking for something like WebSockets (which has a more common perception of being hard). And so the people finding us or landing on our page had a specific goal in mind, and our additional offering of webhooks wasn't enticing.
3) Our webhooks feature simply wasn't that good. It handled retries, fan-out, rate limiting, ordered delivery, and full payload customization. That's a start, but a complete solution also needs inspection, response feedback, test calls, and UI widget.
With the right product and GTM, maybe it can work. Good luck!
Thanks!
What happens when the requests to your API fail? Do I need to retry? Will there be an SDK that can help with this?
It's a good question, and I answered it in a sibling comment: https://news.ycombinator.com/item?id=26400510
Essentially I think it's a risk with any external API you use. What happens when Stripe goes down? Twilio? Sendgrid (when you use magic links login)? Our whole focus is uptime, that's what we do. :)
Will definitely be using this on my projects.
Congrats!
And yeah, I'm familiar with hookdeck, your co-founder already commented in this thread. :)
Some notes: 1) I think you are to cheap. 2) signatures of sent webhooks is pretty important for the receiving party.
1. Ah, maybe. We are still new and trying to figure it out. Thanks for the feedback.
2. We already offer them, including libraries for some languages to easily verify them!
Good luck!
PS. On a side note, it would be very nice to have a standard for webhooks similarly to OAuth. We're designing a solution for accepting various webhooks and the variety of negotiation schemes required by various cloud services is absolutely counter-productive.
Both for the sending part (it's really just a REST API under the hood for now), and the consuming (verifying signatures).
At NewsCatcher [1] we've been asked for webhooks to our data a few times. We still did not start the implementation. Do you think Diahook would be a good solution to consider?
Also, 1$ per 1,000 hooks. What exactly does it mean? Like any push through your hook?
[1] newscatcherapi.com
It's $1 per 1,000 messages. Retries, or if there's no one to send to are free. So it's just for actual webhooks sent. I'll clarify it now.
Anyhow, I'd love to chat and help you get started, please email me at tom @ the domain!
Doesn't mean I'm right.
Congrats on the launch and good luck!
On the other hand: now we have 2 Services that could crash and our app still need a retry logic.
Or we could use it as a fallback, if our services is down, we sent to Diahook.
With that being said, this is something we plan on helping with too. We will have retries to different endpoints (in different geographies), and ways to easily implement it locally. We already offer idempotentcy so that helps.
At some point you're going to have to quantify that and provide enough evidence (of one sort or another) for this claim to be credible. After all, every service under the sun will claim things like 'scalability', 'performance', 'uptime', 'resiliency', whatever. Devil is in the details.
This is something we focus on though, and as part of this focus, we will focus on earning this trust and making it as transparent as we can.
I don't want to nitpick but ... are you really 'focused' on this? You're a startup with limited resources trying to get a product up and running with commercially-viable feature-set. You're not setup for high availability. You just aren't. Proper HA is very hard and as you climb the '9s' in uptime, you introduce huge amount of complexity, cost, planning and manpower - which is something you haven't done and can't do as a startup and, honestly, it isn't worth to do for a Webhook SaaS.
And that's OK! The nice thing about Webhooks is that nobody runs mission-critical workflows with them because there is an implicit expectation of unreliability. This means that for HA/uptime you just need something reasonable. Your service may not survive an AWS outage (or whatever cloud service or datacenter you're deployed in) or a directed DDOS attack, but if your infrastructure can handle traffic spikes and occasional node going down - you're golden.
P.S, what you are doing with Fly.io is super cool!
We happen to be pretty good at infra, so a library / self hosted setup is usually preferable. This is partially because we prefer to monitor things ourselves, and partially because we don't want to send things like auth tokens through a third party service. We might be unique, though.
I'll reach out about Diahook, as we can also have a self-hosted offering, and I think it could still make sense for you. As you can still benefit from the rest of the offering, including client side libraries and management UI (will be fully customisable soon).
I'm actually very privacy conscious myself, which is why we plan on offering optional end-to-end encryption in the very near future. This way only you and your customers can access the data, not us. :)
Our libraries will deal with the encryption/decryption for you.
I wonder how this one stacks up given is centered for webhooks specifically