Give me /events, not webhooks
blog.syncinc.so
blog.syncinc.so
Many large customers eventually had some issue with webhooks that required intervention. Stripe retries webhooks that fail for up to 3 days: I remember $large_customer coming back from a 3 day weekend and discovering that they had pushed bad code and failed to process some webhooks. We'd often get requests to retry all failed webhooks in a time period. The best customers would have infrastructure to do this themselves off of /v1/events, though this was unfortunately rare.
The biggest challenges with webhooks:
- Delivery: some customer timing out connections for 30s causing the queues to get backed up (Stripe was much smaller back then).
- Versioning: synchronous API requests can use a version specified in the request, but webhooks, by virtue of rendering the object and showing its changed values (there was a `previous_attributes` hash), need to be rendered to a specific version. This made upgrading API versions hard for customers.
There was constant discussion about building some non-webhook pathway for events, but they all have challenges and webhooks + /v1/events were both simple enough for smaller customers and workable for larger customers.
Watch me get downvoted like crazy by all Nodejs developers. Even though they could accomplish exactly what they want with much less code and far less complex systems to maintain.
It worked very well.
EDIT: To add some context, my team had come off building a webmail platform, and so we'd done lots of interesting stuff to qmail and knew it inside out. We then launched the .name tld and built a model registrar platform that on registration would bring up web and mail forwarding for users that wanted it. We used SMTP to handle the provisioning of those while keeping the registration part decoupled from the servers handling the forwarding. We also used it to live-update a custom DNS server I wrote.
We had similar-ish constraints. SLA was internal, not imposed (the .name registry had externally imposed SLA's, but the registrar platform did not), but the zones were very simple - either NS records pointing elsewhere, or identical CNAME/MX records, so we needed only a short string per address.
I don't remember if we used CDB files or if we stored individual records directly in ReiserFS filesystems (our mail platform had relied heavily on the ability of ReiserFS to handle vast quantities of tiny files, so were comfortable with that), but it was definitively something simple.
Similar for the web forwarding, which just required a url to redirect to.
If a node should ever need to be replaced, all we'd need to do would be to start a queue on a new box but not process it, then rsync over the dataset from another server, and start processing the queue, and add it into rotation when up to date. If we'd needed stricter consistency guarantees it'd have been a different consideration.
For many types of workloads I'd pick another queuing system today, but the amount of readily available tooling for e-mail, especially once you need federation, reflection/amplification etc. does make it an interesting choice for some things.
It also made debugging the message flow trivial: just add a real mailbox to the cc:....
I didn't downvote you but I bet they come from this part. People don't like this kind of negativity.
> But it's not "cool" technology or "web-based" so developers won't consider it. Watch me get downvoted like crazy by all Nodejs developers.
I say this from experience, as someone who's used a few stupid technologies over time.
I find node to be surprisingly well rounded.
SMTP itself is interesting, although it comes with fun new footguns like STARTTLS.
I'm sure absolutely nothing bad will come from that last bit. Oh look:
"And yes, STARTTLS is definitely less secure. Not only can it failback to plaintext without notification, but because it's subject to man-in-the middle attacks. Since the connection starts out in the clear, a MitM can strip out the STARTTLS command, and prevent the encryption from ever occurring."
I find:
> If the client is configured to require TLS, the two approaches are more-or-less equally safe. But there are some subtleties about how STARTTLS must be used to make it safe, and it's a bit harder for the STARTTLS implementation to get those details right.
I previously thought that was the default, good to know it isn't / might not be
Thanks everyone :-)
You might have some valid criticism about the cryptography; I would not be able to judge that (except when you are basing it on wildly outdated information). I’m not an expert on the details; you could most assuredly argue circles around me when it comes to the cryptography, and possibly about the DNSSEC protocol details as well. But, from my perspective, your continuous claim that “nobody uses” DNSSEC is simply false. DNSSEC works, usage of DNSSEC is steadily increasing, and new protocols (like DANE) are starting to make use of DNSSEC for its features. Conversely, I only relatively rarely hear anything about MTA-STS.
#!/bin/sh
while read domain
do
ds=$(dig ds $domain +short)
echo "$domain $ds"
done
... and note that virtually none of the domains, in any sane list of top domains, are signed. That was true several years ago and remains true today, despite the supposed "increase in usage" of DNSSEC.What's actually changed is that registrars, especially in Europe, now apparently auto-sign domain names. That creates a constant stream of new, more-or-less ephemeral signed zones that gives the appearance of increasing DNSSEC adoption. Of course, this is also security theater (the owners of the zones don't own their keys!). The real figure of merit for DNSSEC adoption is adoption by sites of significance, and that has been static, and practically nonexistent, for a decade.
It is no surprise to me that people working on the DNS talk quite a bit about DNSSEC. People who worked on SNMP talked quite a bit about SNMPv3, and IPSEC people probably really believed there would be Internet-wide IKE. None of those things happened, because what matters in the real world is what the market decides. Most especially at the companies with serious security teams, DNSSEC is a dead letter standard.
In fact, the new CDS and CDNSKEY DNS records allow it to work the other way around; DNS server operators can auto-sign domains, and the registrars need not be involved at all.
> The real figure of merit for DNSSEC adoption is adoption by sites of significance
People said the same about IPv6. Or maybe you do, too?
> People who worked on SNMP talked quite a bit about SNMPv3
I seem to recall you mentioning quite often how WHOIS was dead and would be replaced by RDAP. That didn’t happen either.
> IPSEC people probably really believed there would be Internet-wide IKE
Interestingly, that problem could in theory be solved by DNSSEC. We’ll see what happens.
STARTTLS exists for two reasons (https://www.fastmail.com/help/technical/ssltlsstarttls.html):
1. Wanting to accept mail insecurely.
2. Not wanting to use two different TCP port numbers to send and transfer mail.
To solve these problems they created STARTTLS. But obviously, STARTTLS isn't actually secure (even though that was the point of supporting TLS). So to make it secure, it's suggested to use DANE - a standard built on a different procotol, requiring a feature that is controversial, potentially dangerous, and not widely implemented. So you can use a kludge (STARTTLS) with a kludge (DANE) to send and transfer mail securely. But should you?
Since 2018, RFC8314 says that e-mail submission should use implicit TLS, not STARTTLS (https://datatracker.ietf.org/doc/html/rfc8314#section-3). Therefore the use of STARTTLS, and the use of DANE to make it secure, are deprecated. So while you shouldn't use DANE for anything seriously, you really shouldn't use it for SMTP.
DANE is necessary as long as there are still some agents using backwards-compatible behavior; i.e. falling back to unencrypted communication if TLS is in some way blocked.
It is only when we allow backwards compatibility that something is needed to differentiate to the clients whether the server is new enough to allow TLS or not.
As a developer who used Salesforce for nearly a year once upon a time, I can confirm that exposure to stupid decisions in a platform can affect the developer.
Node, though? Could you expand on the stupid decisions in Node? And does Deno address those?
Some of the stupid in node just comes from the fact that there's still a lot of reinventing the wheel, and doing a less good job of it. Like, we've got all these backend frameworks, but still nothing at all that compares to eg Spring. Can you even find a nodejs lib that does HATEOAS properly and completely? How often do you find yourself doing string parsing, or handling a JSON object, when you know it would be more efficient to be handling a stream, or that really the kind of work you're doing ought to be handled by your framework but isn't?
As for nodejs itself, it's much better in 2021 than it was in the past. But it's still a massive runtime. And I have mixed feelings about eg Worker threads. As for node_modules, I get the sense that we're just replaying the history of Microsoft's dll story, needing to relearn all the lessons that should have been learned already.
As for Deno, I think it comes with great ideas. In many respects, I like it better than Nodejs. Most of its good ideas, Nodejs is flexible enough to accomodate. One of Deno's main advantages is that it doesn't have any legacy to support, so it can embrace things like ECMAScript modules more easily. Its library system is closer to Go, although I think the end result is that a lot of folks end up doing one-off systems that end up looking a lot like the nodejs module resolution system in the end. Deno's main disadvantage is that it is not compatible with nodejs libraries. That's also an advantage insomuch as you have a clear module import spec from the get-go.
In short, the stuff that Deno can do, Nodejs can do, and I'm not sure that it's cleaner system can overcome the fact that the same is accomplishable in Nodejs. I'd be more than willing to use Deno in a greenfield project because I like all the technology choices it makes, but fundamentally, the technology choice you're making is whether or not to use V8, and adopting Deno is almost just a way of pressing the reset button the ecosystem, which may or may not be a good thing depending on your needs.
"Here's a crazy idea ... This are the properties for why it works ..." Would have sounded much better.
Stripe needs to be usable by both the developers building intense, scalable, reliable systems and the people teaching themselves to code in a limited context on a limited platform.
I think it depends on the developer. There's developers hammering out boring business logic as fast as possible and there's developers with a deep understanding of machine internals, protocols, and infrastructure. For the former, SMTP is black magic they'd probably never think of and involves engaging the one infra person that's always busy
It also means standing up and managing "infrastructure"
If I ever heard a dev at work say "No I won't use that new tech, it's too untested/I'll have to spend more time figuring out how to make it work well", I would shit my pants. Whereas if it's old tech, "it's not modern/I'll have to spend more time figuring out how to make it work well". It's practically software ageism...
I tell you that as someone who fiddled with sendmail.cf’s and other .conf’s way too much long before nodejs became a thing. Now it’s a relief.
Purely anecdotal of course, but I follow a number of the latter, and they're either sparsely employed or often employed in a capacity where it doesn't matter There was a comment in this thread where one person had such an idea, and it was rejected for what were essentially business reasons.
At this point in my career (10 years in the game), let me simply defend node as the tool that got me here. Using it then to bootstrap my career was just as practical as using SMTP as you describe now.
Lesson being, that sometimes there are unexpected reasons why a specific piece of technology shouldn't be used.
Either way, the auditors and the infrastructure might not want to handle an order of magnitude more traffic (API usage is really in a different league than occasional human email). Expect all emails to have to be stored for around a decade.
It sounds like one not allowed to use http for restful APIs without calling it a website. (And that org require website to be audit to fulfill accessibility measurement for physical disability)
I’ve been building for the web for 15 years and it shows how far I can hyper focus on certain communications implementations that I’m not looking at pre-existing options that really meet a large number of use cases. I suppose it also means making sure your data consumers are comfortable working with the protocol but it’s a really top notch idea.
But... Why?
The HTTP protocol is so much easier to manage, load balance, use, etc.
> there are risks when you go down.
Solved by SMTP on protocol level. With HTTP, must be solved on both client and servers application level.
> webhooks are ephemeral. They are too easy to mishandle or lose.
SMTP has this baked into its heart. Loosing messages is possible, certainly, but rather hard to do. With HTTP, it's really simple.
> In the lost art of long-polling, the client makes a standard HTTP request. If there is nothing new for the server to deliver to the client, the server holds the request open until there is new information to deliver.
SMTP is push, not polling. So all those issues are solved for you.
Yes. And if you want to poll, POP is polling, and IMAP has both polling and immediate notifications.
And if that doesn't succeed for some reason, it reliably queues and retries.
That's a push.
But here's another way you could (ab)use the mail system for delivery, provide a mailbox for the client and just allow IMAP or POP access and throw the messages into that. The client can log in to access and process them (which they would likely be automating on their own mailboxes anyway). It does mean it's housed at the provider, but it's also pretty easy to scale. There's lots of info on how to set up load balanced dovecot clusters out there, and even specialized middleware modes (dovecot director) to make it work better so you can scale it to very very large systems.
If the distinction is too hard to make: think of it as using the 'Simple Event Transfer Protocol' that just happens to use exactly the same protocol as SMTP.
Yeah, but it's a protocol for transferring email. As I noted with "you have to be very careful or have a lot of control over the receiving server to ensure you actually get delivery", you can amstract most of the mail system out as long as you ensure you are running the server they deliver to, but you would also need to rely on them making sure their outgoing server is good for this, which probably means dedicating it to this and not running any real mail through it (so you avoid outgoing company email filters, etc). At that point, both sides are running specific bespoke mail servers, which cuts down on the usefulness of the solution because of how much setup and administration it requires.
It used to be nobody ran incoming and outgoing filtering on email, so it was a robust channel for communication with retries, and notifications for failed delivery, etc. These days it's not exactly that because of all the spam mitigations and company compliance and risk mitigations that might be in place, etc. In fact, just setting up a new mail server and attempting to send to microsoft (live/hotmail), yahoo or gmail is extremely hard, because they have a high bar for acceptance, and large swaths of the easily obtained IP space have already been blacklisted from prior spam use so you start at a bad reputation and have to work to get it to a level you've even be allowed to talk to other by working with all the third party (and first party) blacklist maintaining entities.
Notably, in choosing to use an atom/rss feed, you need to determine what the webserver serving it is, how to implement authentication on top of it (is it a token/oauth, HTTP auth, param auth, etc), what is the underlying data store (SQL/NoSQL, some message system), how to scale that system if you expect it to be large and span multiple servers and/or datastores (mail systems right now deal with hundreds of thousands of users and gigabyte plus mailboxes of millions of messages).
Choosing IMAP to deliver this info means there there are well worn solutions for all the decisions you need to make (including howtos to implement oauth at the server level), as well as client level libraries in almost every language. Basically, you could decide to use it and not have to worry about forging a new path on that system basically ever, because there's plenty of people that have already implemented it at a larger level and with the same features (even if you would be using them to slightly different effect), and they've contributed the info on how to do it and what the performance ramifications are to the public domain.
I'm not seriously advocating for it, but that's more because clients will look at you funny than for any technical reason. Technically, it actually has a lot going for it. Unfortunately as an industry we fetishize the new and bespoke because obviously our own unicorn projects are so new and special and will serve so many people that some off the shelf solution could never be as good....
I think the bigger issue is that consumption isn't particularly friendly. Also, you still haven't solved the versioning issues.
Why you need to insult a whole body of people, rather than just make a claim about the technology, I don’t know.
In an enterprise setting it becomes more complex if a 365 subscription is required, or active directory authentication is needed to receive emails. Does someone need to monitor the inbox to confirm it's working etc.
But after you mentioned it, I do wish that this was an alternative to webhooks that more service providers offered.
So the suggestion is to use email? That's not how others are interpreting it. [1] And it doesn't make sense to me either. Emails as they are "baked into the whole internet already" are unencrypted with tons of middlemen, and even their transport isn't guaranteed to be encrypted. Email is also munged and messed with in weird ways, with fun stuff like each middleman tacking on their own headers and filtering it out based on unknown rules. It also introduces a ton of latency and severely prioritizes "eventually reaching the destination" over timeliness. And more downsides I can't think of off the top of my head. That seems like a really poor choice for an event delivery mechanism.
First, you don't have to use the rest of the internet's e-mail system. Stripe can run their own mail servers that deliver straight to clients on non-standard ports using implicit TLS, ensuring security and no middle-men. This also ensures delivery is as timely as possible (sub-second typically, as mail software has to be fast to handle its volume).
Let's say you want to poll (ex. "/events"). The client uses IMAP to poll the Stripe server with a particular username/password. Check a folder, read a message, delete it on connection close. There are of course ready-made solutions for this, but you can also write simple IMAP clients really easily using libraries.
Let's say you want pushes (ex. webhooks). The client sets up the alternative to the webhook-server they'd have to set up anyway: an SMTP server. Use a custom domain, one that has nothing to do with the customer's main business, so nobody ever gets confused. Configure it to only accept mail from a "secret mail sender" (aka webhook secret). Part of the "SMTP webhook URI" would be what mailbox to deliver the webhooks to. The client then configures an MDA on their mail server to immediately deliver new messages to some business logic code. If the MDA or business-logic code has a bug, the messages will stay in the client's mailbox until they are "delivered" successfully. If the client's SMTP server is down, Stripe keeps retrying for at least 3 days, more if Stripe wants.
Stripe could actually implement both by keeping messages in an IMAP folder on Stripe's servers, and deleting the messages once the SMTP server confirms delivery to the client. Of course all messages already have unique IDs so removing dupes is easy.
You could implement all of this in a week, write almost no code, and still handle all the weird edge cases. Virtually all of that time is just reading manual pages and editing config files. The end result is a battle-hardened fault-tolerant event-driven standards-based distributed message processing system. The maintenance cost will be "apt-get update && apt-get upgrade -y", and anyone who can configure Postfix and Maildrop can fix it.
The API is entirely HTTP, and tries to meet users where they are by providing tools that they are most comfortable with. Frequently, these users are familiar websites or mobile apps. As such, webhooks are implemented over HTTP.
If there was an alternate way to integrate with events, it'd be something that's either:
1. accessible to novice users, or
2. delivers on high throughput/latency needs of the largest users, or
3. resolves a storage/latency/compute cost incurred behind the scenes
Thinking about these:
For #1, websockets would score better than SMTP
For #2, kafka (or a managed queue like SQS) would score better (many support dead letter queues and avoid the latency at the mail layer)
For #3, it isn't clear that SMTP reduces the latency, compute, or storage costs
SMTP might be familiar -- and it's possible for you to build your own webhook → SMTP bridge if you wanted it -- but doesn't score well enough on any of these metrics to be built in-house.
[Disclaimer: I work at Stripe, and these opinions are about how I'd approach this decision. They're not the opinions of my employer.]
But I'm not convinced on the latency/compute/storage comparison with Kafka or other solutions. I think a POC would need to be built and perf tested, and then tweaked for higher performance and lower cost, like most software. Considering the volume of traffic that mail software is designed for, I can't see how even a large provider like Stripe would have difficulty scaling a mail system to match Kafka. It's not like mail software is written in Java or something ;-)
> Any push notification like system is going to be just as error prone.
E-mail servers are built with retry logic and queuing logic already. The point is if you need queuing anyway, it offers a tried and tested mechanism with a multi-decade history and a vast number of interoperable software options. While there is now a relatively decent number of queuing middleware options, none of them have as many server and client options as SMTP.
SMTP isn't the best choice for everything, but it works (I've used it that way), it's reliable, and it scales with relative ease.
> And what happens when the send fails?
It gets retried. Retries are built in to mail servers. That's part of the point.
> And what about multiple message handlers?
What about it? Most SMTP servers provides mechanism for plugging in message delivery agents rather than delivering to a mailbox, or you let it deliver to a mailbox and pick it up from there. Or you plug in whatever routing mechanism you want to distribute the messages further. The sheer amount of ready-built options here is massive.
> Did someone write the code to check the inbox for them and handle them? When a send fails multiple times, is that logged and is there a system for clients to check that log?
Pretty much every e-mail server ever written provides a mechanism for handling persistent failures, and many of them offers heavily configurable ways of doing it. But yes, you'd need to decide on what to do about persistent failures. But you need to do that whatever queuing system you use.
> Message transfer isnt the hard problem in this domain.
The point of the article is exactly that reliable message transfer is the hard problem in this domain.
Developers won't be able to use the existing email systems of the company, too critical and managed by another team. They will never be able to reconfigure it and get API access to read emails. Note that it may or may not be reliable at all (depends on the company and the IT who manages it).
Developers won't be able to setup new email servers for that use case. Security will never open the firewall for email ports. If they do, the servers will be hammered by vulnerability scanners and spam as soon as it's running. Note that large companies like banks run port scanners and they will detect your rogue email servers and shut it down (speaking from experience).
As for "being hammered", rejection of invalid recipients before even getting to the DATA verb is cheap.
Having actually run both an e-mail service and SMTP used as messaging middleware, I have dealt with these issues.
You could work around it but should you? You're exposing the company to fines and risking your job.
Better think of another way to integrate with the vendor, or find another vendor.
P.S. SMTP is easy to identify on ANY port, it's replying a distinctive line of text when TCP connection is opened.
If they freak out over an SMTP server but don't freak out over a web server, then the are indeed absolutely utterly incompetent fools that should never work in this space.
In both cases code written by the company developers will eventually process untrusted textual input, and you need to deal with that with the same level of caution, and the protocol does nothing to change that.
> You could work around it but should you? You're exposing the company to fines and commiting a fireable offense. Better find another product that's easier to deploy.
I would not work around it - I would make the case that there's no difference in exposing a carefully chosen SMTP server than exposing a web server, and if the security team fail to understand that, I'd resign, because it'd be a massive red flag, and I've been successful enough to be in a position to not need to work for companies like that.
For that matter, in 25 years in this business I've yet to run into your hypothetical scenario, including at large companies, so I'm not at all convinced it'd be a genuine problem. Yes, I've been at companies where I'd need to provide a justification for getting a port opened. But never once had an issue getting it approved - including SMTP.
> P.S. SMTP is trivially identifiable on ANY port, it's giving a line of text when the TCP connection is opened.
I was responding to "Security will never open the firewall for email ports.". Point being that if they care about the specific port numbers, it doesn't matter.
[And I'll again point out I've actually run infrastructure like this].
I'm speaking from real experience too. It takes a while to open firewall in some environments, if you ever can.
One bank was the worst. There was a super stringent process to expose things externally. Opening the firewall port was just the beginning and that'd take 2-4 weeks if all goes well.
You'd struggle like hell to expose a SMTP server though because it would be immediately be rejected and flagged based on the port. Banks have to store, monitor and ensure the origin of all emails, they don't allow shadow email servers. And it's plain text so more reasons to ban (also a problem with HTTP, you should do HTTPS if anything).
Defense was simpler, mainly because there was no external connectivity in many cases. You don't need to worry about how to open a firewall when there's none :D
I might be missing a point or two here, but I don't see how SMTP can work for this case at all. You would require every API consumer to setup a SMTP server (which is another piece of infrastructure to maintain), and then somehow have a layer of authentication so the recipient can control who post messages on that server (overhead for the publisher per new customer). Then we still haven't resolved the issues on the customer side (a bad code that could pop all messages and now we might require the publisher to replay them again).
I haven't even started to think about security and network hardening challenges yet. Again, I might be missing the point but this is not a case of cool tech overuse to me.
As for "setting up an SMTP server", the point is that compared to the current requirement of a webhook, you're going to need a queuing mechanism or a pull mechanism or both anyway. So you can build a custom solution, or you can pick an existing queuing mechanism that people have spent literally decades providing a vast array of software options for.
And yes, you're right, you can always ending up needing a way to trigger a replay because no matter what you do the customer might do something stupid. Nothing you do will get you away from that. So either you require them to always pull, or you provide an option to push and an API to trigger redelivery for when they've done something stupid. If you opt for push, SMTP is an option worth considering, because no other queuing mechanism has as many available ready-made and battle hardened queuing options.
There are many cases where it'd not be suitable, but in the situations where SMTP is a bad choice, webhooks are likely to be an absolutely awful choice.
I speak from actually having run messaging on SMTP both as an e-mail provider with a couple of million users and having used it as messaging middleware in production.
You will probably point me to SMTP success messages, but a removed mailbox might only be known by a backend server.
Also mail infrastructure will potentially include heavy spam filters etc. making it quite inconvenient. Not even mentioning security aspects with limited availability of transport layer encryption with proper signatures.
Same goes for security - your objection is true for public e-mail delivery without additional requirements on the servers or clients, but that is not relevant for a private infrastructure.
Edit: not sure why I'm being downvoted. I work at stripe and this is literally how it works.
Years ago there wasn't even a Kafka portion, that's newer.
One thing I really admire is how Stripe makes it transparent which events were fired both in general through the Developer area, and on specific objects like customers, subscriptions, etc..
Our main focus is on handling Trello notifications, and the Trellinator library I wrote is built in, our objective is to create more API wrappers over time to make it as simple as possible to deal with as many APIs as possible. You can see some example code here:
https://trello.com/b/IoHmhz5c/benkobot-community-board
You currently require a Trello account API key/token to sign up, but you can use it as you described to be a generic endpoint, transform however you want with JS then post the data onto another endpoint.
Like many others, I now pattern my own APIs after Stripe's.
(I worked on the same team as bkraus, non-concurrently).
For teams that are building webhooks into your API, I'd recommend including UI to view webhook attempts and resend them individually or in bulk by date range. Your customers are guaranteed to have a bad deploy at some point.
The only event API I ever want is notifications there's new data, and then an interface by which I can query all new data which has arrived by some sort of index marker - because this is fundamentally reliable. It means whatever happens to my system, I can reliably recover missed events, skipped events, or rebuild from previous events.
And this is in fact exactly how something like Kafka actually works! Complete with first-class support for compacting queues to produce valid "summarized" starting points.
Any streaming system essentially should never start as a streaming system - it should start as a slow-path pull-based system, and have a fast-path push system added on top of it if needed - because then you've built your recovery path already, rather then what happens way too often which is just "oh yeah, we'll develop that when it breaks".
Extra points for being able to set something like 1s between pings (now you see why I like the option ID for a range).
I think this is a quite interesting and important point. When we talk about "doing the simple thing first" too often we end building something that is technically simple but fickle. The trick to making the simple thing reliable is to figure out which part is the slow-path (or failure mode), and then only building that. Unfortunately, it often means out result ends up technically "boring" since all the interesting optimizations are what we cut out, but I think that's worth it if the end result is a more useful product.
It's something I've been working with and thinking about for a while. I think it applies to a way broader scope than this discussion.
Yes, this is pretty much the right thing to do. It can be a bit more work for the API consumer, partly because they need to track state of their last-read ID, and there's more moving parts.
If you're building a webhhook+events system like Stripe's, you might consider adding an option for a mostly-empty webhook body, which can speed things up in this use-case, but still allows "the easy way" of just processing the event from within the webhook body.
(For readers thinking of implementing this, note that "query for new data" means hitting a dedicated /events api, not individual tables, which might have unpleasant load/performance consequences).
The problem is that all of those services are internal to our network, and aren't accessible from the outside world. We cannot set up a webhook to Jenkins because Jenkins does not have a publicly accessible URL. We cannot set up a webhook to Gitlab, or to Prometheus, or to Sentry, or anything else, because those are all internal services.
The only option there would be to create a new, public-facing server, set it up with a domain name and SSL certificate, expose it to the world, and then give it access to those services - which defeats the point of having those services internal and secure if we just create a non-internal system and give it access to them.
Alternately, we have that new, public-facing server buffer those requests and have other services poll them, somehow, so that it cannot connect in, but now we're getting into the same situation as described in the article.
If there were an API, I could easily create a small daemon that would watch for events and dispatch them accordingly, and then respond to them as needed; instead, my only option is to build some kind of Frankenstein - or to give up entirely, which is the more reasonable solution.
Then again, this is Microsoft Teams, where creating an application requires an Azure account and jumping through a ton of hoops, so they're no stranger to stupid ideas that no one wants to deal with.
You might look into Cloudflare Tunnel (formerly Argo). It is free and allows you to poke a hole in your firewall to a specific service. If that meets your security requirements.
https://blog.cloudflare.com/tunnel-for-everyone/
Also in my personal testing I didn't need to pay.
The beauty of this is that you also can compose with other nodes and for a distributed service by calling the local service as a proxy and routing the requests to the other nodes of the same api.
It took more time than i've predicted because its also expected to deliver UI and most of the 'HTML5' api to native applications (instead of Javascript), which is a massive platform by now (and the #1 reason why newcomers to browser technology cant compete, giving the feature creep tax imposed to them).
The idea is also to distribute over a DHT so you can just serve your application over torrent without needing to register anything..
The only way to get there is by empowering users and developers and taking some of the control from the cloud platform giants.
In my point of view the only way to break the browser monopoly now is to create a new path forward, a branch.. its not the time to follow the rules, its time to break them or else the future doesn't look so bright in my opinion..
This is a problem with receiving any inbound data from a third party. At least with HTTP, it's pretty trivial to set up a robust reverse proxy with nginx.
My company's internal apps use a mix of VPNs and IP fenced load balancers. We are migrating to app proxy.
No inbound connections + access based on Azure AD identity with conditional access (restrict apps to Intune enabled corporate devices) and MFA is an absolute killer.
My only complain is that connectors are not very DevOps friendly. Cloudflare Tunnel is much better in this area.
That's not the only solution -- you could also develop a bot that will do those specific things.
In the days of yore I know of at least three companies that were using IRC bots to similar effect long before webhooks ever existed.
Because of that prior experience, this is how I currently manage a similar set of problems, albeit not on Teams in my current role.
The downside was that the event api required a huge amount of scope, so if you weren’t careful and were compromised, someone could use that token to scrape all messages in the system.
Slack recently added socket mode for precisely this reason: https://api.slack.com/apis/connections/socket
* https://zulip.com/api/real-time-events
* https://zulip.com/api/register-queue
* https://zulip.readthedocs.io/en/latest/subsystems/events-sys...
We use this same long-polling based /events API interface for all official clients (web, mobile, terminal), our interactive bots ecosystem (https://zulip.com/api/running-bots), and many integrations (E.g. bridges with IRC/Matrix/etc.).
We also offer webhooks, because some platforms like Heroku or AWS Lambda make it much easier to accept incoming HTTP requests than do longpolling, but the events system has always felt like a nicer programming model.
(Zulip's events system was inspired by separate ~2012 conversations I had with the Meteor and Quora founders about the best way to do live updates in a web application).
A good platform should offer both of these and more (for example Slack does webhooks, REST endpoint, websocket-based streaming and bulk exports), and let the client pick what they want based on their use case.
A way to fix this is to use an application-level keepalive (TCP keepalives are generally useless), but then that increases the load on the server and adds a scaling burden.
Meanwhile, unless the event stream is stateful (more overhead!), the client has lost all events since the connection has dropped, and the client can't even be sure when the connection actually dropped.
With webhooks, assuming the callback sending service has a generous retry policy, and the customer's receiving service does not return 200 unless the webhook has been completely processed, or persisted to storage, you won't lose events.
I've been at Twilio for the past 10 years. We recently started offering an event stream service (that customers had been requesting for some time), but it's complicated to get right (on both the server and client side) and difficult to scale, and, frankly, webhooks have worked fine for most customers for a very long time.
Exactly why mqtt has the ping packet for the client.
are you sure? specifically, are you sure a persistent connection has _more_ of a cost than repeatedly re-establishing a connection & TLS, etc.?
in terms of energy costs alone, DNS resolution, establishing routes, generating cryptographic session keys, etc. it's definitely not as cheap.
in terms of today's computation power, the "memory" costs of maintaining a connection are minuscule, and the performance "penalties" are negligible.
example: lets say you have 50k event subscribers. if nothing happens, then, aside from a few TCP keepalives (which are not strictly speaking required, and can happen very infrequently), no traffic moves. if instead you have polling once every second, then that's nearly ~13-14 connections a second, each one with at least 4 round trips of traffic. that's a measurable amount of load.
Straight long-polling should be avoided, but intermittent polling is a good solution for performance when you don't want to use all your socket bandwidth.
You should use normal short polling and Server-Sent Events.
Also it makes no sense to say long polling is more reliable than SSE, because SSE is essentially a non-hacky implementation of long polling.
With webhooks, as in the article, you only get state changes; you need some separate mechanism to achieve (or recover) the initial state.
I'm thinking of simpler situations like my source host's CI spinner that seems to get stuck all the time due to missing the ping back from Jenkins about build statuses. In that case it really would be fine to always just say "I think the state is X, please answer me now or in the future whenever the state is other than X." I don't care about anything other than an up to date sync.
Kafka solves exactly the issue that the author is complaining about. This is a safeguard to ensure that data isn't dropped in the event of an issue, and provides mechanisms to replay events.
The tradeoff between pushing and polling have been argued since forever.
In other news, mechanics who work with bolts often do so with ratchets. This is a cumbersome compromise, just give me Torx fasteners!
(And, of course, I don't want Kafka. I want Google PubSub. No, wait, I mean SQS. No, wait, I mean I want zeroMQ. No, I mean....)
Certainly the event producer is in a better position to maintain a queue without missing events, but it also means they need to buffer more data in their queue system to accommodate for your receiver's downtime
Conceptually, the important thing is each stage waits to "ACK" the message until it's durably persisted. And when the message is sent to the next stage, the previous stage _waits for an ACK_ before assuming the handoff was successful.
In the case that your application code is down, the other party should detect that ("Oh, my webhook request returned a 502") and handle it appropriately -- e.g. by pausing their webhook queue and retrying the message until it succeeds, or putting it on a dead-letter queue, etc. Your app will be "out of sync" until it comes back online and the retries succeed, but it will eventually end up "in sync."
Of course, the issue with this approach is most webhook providers... don't do that (IME). It seems like webhooks are often viewed as a "best-effort" thing, where they send the HTTP request and if it doesn't work, then whatever. I'd be inclined to agree that kind of "throw it over the fence" webhook is not great and risks permanent desync. But there are situations where an async messaging flow is the right decision and believe it or not, it can work! :)
I know the goal for most systems is just to be 'up to date' Not to get the entire history. So in most cases you don't need to stash all the messages, you just need to be able to retreive the latest state of stuff...
For example, you rolled out code on the receiver side that did the wrong thing with each message. Now there's no way to replay the old webhooks events in order to reinstate the right behaviour; there's no way to ask the producer to send them again.
The only way around this is to store a record of every received message on the receiver side, too, which the article author thinks is an unnecessary burden compared to polling.
Personally, I think push is an antipattern in situations where data needs to be kept in sync. The state about where the consumer is in the stream should be kept at the consumer side precisely so it can go back and forth.
Embedded systems don't do that for webhooks because they can't (very little RAM or non-volatile storage) but customers clamor for webhooks anyway because it's what their web developers know how to use. So inevitably they're going to lose data but they're only getting what they asked for.
Of course, then you need a way for the receiver to retrigger or view the webhook if one gets missed, which starts to look like you have to have a polling endpoint anyways, though.
Also, swapping out one messaging system for another is trivial. Pick the one best suited to the environment you're working in, and if that environment changes, changing messaging queues is going to be one the easiest transitions you'll make.
> If the sender's queue starts to experience back-pressure, webhook events will be delayed, and it may be very difficult for you to know that this slippage is occurring
I've never before seen anyone try to argue that properly dealing with backpressure is a bad thing. The author's proposed model makes this situation even worse. With kafka, consumers can continue processing the event stream and you can continue to serve reads from your primary datastore. With the author's model the event stream lives in your primary datastore, so if that starts to lock up the blast radius is much larger.
Now, if that premise isn't based in reality, or if it's already been solved some other way, discredit it without giving it too much air time.
A one liner about kafka being cumbersome and then building your own solution, warts and all, doesn't need to exist in the same thought if you've made the reader mentally disregard it as a possible solution.
HTTP Endpoint -> Push to message queue (kafka, SQS, etc) -> Acknowledge receipt
That's a pretty straight-forward design that's widely used, robust, and easy to put together. I've probably done that same workflow 100s of times without issue.
As long as you guarantee the message was pushed to the queue before acknowledging, that will be fabulously reliable. You need to make contingencies for duplicate messages, but that's not usually difficult.
I think that a combo of webhooks / events is nice, but "what scope do we cut?" is an important question. Unfortunately, it feels like the events part is cut, when I'd argue that events is significantly more important.
Webhooks are flashier from a PM perspective because they are perceived as more real-time, but polling is just as good in practice.
Polling is also completely in your control, you will get an event within X seconds of it going live. That isn't true for webhooks, where a vendor may have delays on their outbound pipeline.
I see the value in this, but I actually disagree with the article in terms of that being the best solution. Long-polling is significantly different than polling with a cursor offset and returning data, so you wouldn't shoe-horn that into an existing endpoint.
Doesn't that only work in the case where the server treats each webhook delivery as ephemeral? If you're keeping a queue to allow reliable / repeatable delivery, that's definitely not "zero resource usage", right?
1) /events for the source of truth (I.e. cursor-based logs) 2) websockets for "nice to have" real-time updates as a way to hint the clients to refetch what's new
One thing that makes sense: if you go down use polling so you can work at your own pace. But this isn't really at the same time. When/why does it make sense to do both simultaneously?
1. Fast system that delivers messages very quickly but is not always partition-tolerant or available 2. Slower, partition tolerant system with high availability but also higher latency (i.e. a database)
The author goes through this in the very first section. Webhook events will eventually start getting lost often enough for the developer to think about a backup mechanism.
Long-polling works if you have a lot of memory on your database frontend. Most shared databases want none of your long-running requests to occupy their memory which is better used for caches.
Even if your message bus has the ability to store and re-deliver events, you might want to limit this ability (by assigning a low TTL). Consider that the consumer microservice enters and recovers from an outage. In the meantime, the producer's events will accummulate in the message service. At the same time, the consumer often doesn't need to consume each individual event but rather some "end state" of some entity or a document. If all lost events were to get re-delivered, the consumers wouldn't be able to handle them, and would enter an outage again. This is where deliberately decreasing the reliability of the message bus and rely on polling would automatically recover the service.
There are other reasons, of course. The author is absolutely correct in their statement, though: whenever a system is implemented using hooks / messages, its developers always end up supplementing it with polling.
The underlying problem is that HTTP is a state transfer protocol, not a state synchronization protocol. HTTP knows how to transfer state between client and server once, but doesn't know how to update the client when the state changes.
When you add a /events resource, or a webhooks system, you're trying to bolt state synchronization onto a state transfer protocol, and you get a network-layer mismatch. You end up with the equivalent of HTTP request/response objects inside of existing HTTP request/responses, like you see in /events! You end up sending "DELETE" messages within a GET to an /events resource. This breaks REST.
A much better approach is to just fix HTTP, and teach it how to synchronize! We're doing that in the Braid project (https://braid.org) and I encourage anyone in this space to consider this approach. It ends up being much simpler to implement, more general, and more powerful.
Here's a talk that explains the relationship between synchronization and HTTP in more detail: https://youtu.be/L3eYmVKTmWM?t=235
Edit: from the network point of view, it’s either call-back or a persistent call-wait/socket, or polling. The exact protocol is irrelevant, because it’s networking limits and efficiency that prevent everyone from having a persistent connection to everyone. A persistent connection can’t be much better than any other persistent connection in that regard, and what happens inside is unrelated story. Or am I missing something?
Oh yes, changing HTTP is so easy.
You can add it to your own website with a few simple headers, and a response that stays open (like SSE) to send multiple updates when state changes: https://datatracker.ietf.org/doc/html/draft-toomim-httpbis-b...
...and you can get these features for free using off-the-shelf polyfill libraries. If you're in Javascript, try braidify: https://www.npmjs.com/package/braidify
https://developer.mozilla.org/en-US/docs/Web/API/Server-sent...
Obviously long polling will work in more places but SSE is great if the server side needs to do the pushing.
Also for syncing stripe in particular you might want to check out the work done by Supabase (still experimental):
It’s great that we’ve went full circle. But make no mistake that this only means one thing: that servers are cheaper than ever. We can now afford to entertain previously extravagant ideas.
At this point, its honestly just easier to just have a websocket + events endpoint with a cusor both.
Websockets or server sent events at least signal to intermediaries that a longer term connection will be open.
We've not come full circle. This is just one blog saying "how about long-polling". While also ignoring that since then we've gained web sockets, HTTP/2 and HTTP/3, each of which make long-polling pointless in three different ways.
None of those have strong support across language frameworks for using them as a client from a server context. Especially http/3.
I'm sure many people reading this have many idle GitHub repos set up with webhooks into some kind of build server. The repo might see no more than one commit a week.
It makes absolutely no sense for GitHub to long-poll (or websocket) all of these build servers.
(Now what would make sense is for /events to support a way to flip over to a webhook when it's idle. IE, long-poll for a minute, then the next request sends a URL for a 1-time webhook called on the next event.)
ISPs have very different views on how long a connection is allowed to be kept open, and will absolutely kill long-polling connections without remorse. This extends beyond ISPs, too. Practically all network infrastructure between API consumer and API will have to accommodate long TCP connections, which is unfortunately not as trivial as it sounds.
The next cool feature of long-polling is that the server might not know that the connection is broken until it actually tries writing to it, so if the back-end relies on having or not having a connection, this will make for some interesting edge-cases.
Load-testing products using long-polling also has interesting implications, such as hitting the "65k" problem on test clients and network infrastructure.
Add another layer on top (or below, like Kubernetes, OpenShift, whatever) and you should be getting a prescription for anti-depressants from the very beginning, because you WILL need them.
I guess this is the current zeitgeist's version of "stock up on $PAINKILLER" or even "keep a bottle of $LIQUOR hidden in the bottom drawer of your desk".
But if the producer retries and the consumer does not respond with 200 until it has processed the message, no consumer side message queue is needed, the consumer can rely on the producer to reach at-least-once delivery.
In both cases (webhooks and /events) storage is needed producer side, so nothing consequential has changed, only with /events you need long lived TCP connections which ties up (e.g. you can't do this on FaaS endpoints)
Functions as a service are absolutely ideal for low frequency webhook receivers. SO SO SO cheap.
Part of the point of the article was that you may deploy bad code which returns 200, but doesn't actually take the correct action with the events, and then you have lost all that data and have no way to get it back, which is why you have a consumer-side message bus to hold the webhook history, so that you can replay the webhooks if you made a mistake. Your comment does not address this at all.
If the service exposes a /events page, and especially one that supports long polling (or SSE), then you no longer need a consumer-side message bus, and you might not even need webhooks at all.
I definitely think webhooks should be offered, but I agree with the article that webhooks shouldn't be the only thing.
Having an ephemeral messaging system and a ledger to reconcile against is a nice, simple way to provide immediacy and eventual consistency (where eventual could be days). It's a pattern we're using all throughout our infra.
There was nothing fancy about it, you could just listen to an endpoint and it was a stream of the append only log of events occuring in the DB to the point that you could literally use it to feed a replicated master or slave (or backup).
I imagine your use-case is a bit more nuanced, but I sure do love that model.
If you need strict consistency guarantees, sure. Otherwise, don't piss off your API consumers, webhooks work just fine.
That really depends on the setup you're using. There are plenty of server platforms out there where setting up a cron is no more complicated than making a GET route handler.
In Python, it's either a cron script or Celery beat. In PHP, usually a cron script.
That all means more processes running outside of your web framework, more stuff to deal with, more complexity. Now you have to manage e.g. running a celery process, managing crontabs wherever you deploy...
Even in languages like Elixir where long-running scheduled processes are cheap and native, you still have to write a GenServer in comparison to just using your web framework.
I'm fine with the approach in the article. I actually like that they're giving the option for multiple ways of getting those events. Just please, please don't make /events the only option...
I typically use a language with an event loop (async/await), so there is always something like `setInterval(poll, 500)`. All that code needs is a connection to the internet. If the server is down for 24 hours I can start it up and it'll read the missed events. I can set up 10 dev environments with just an API key difference in the config. I can batch apply events in a single database transaction, ensuring consistency.
But with a webhook, I need to ensure that my server can accept incoming connections, the external API knows that location. Each dev env needs to replicate this public HTTP server set up. I need to monitor the uptime of the server closely as missing events or erroring on a subset of events could leave my database in an inconsistent state.
Or way simpler if you're not building a web application in the first place.
Polling is chatty when updates are infrequent.
Ahh, but all problems are queues. What happens if your webhook destination is down and your source system sending queue backs up? What happens if your destination endpoint is up, providing 200s to requests, but throwing away the data quietly due to a mistake during a rollout? Or your webhook source quickly ramps to a volume that is effectively a DoS attack? (These are all problems I encountered in a role at a low code/no code product).
I strongly endorse others in this thread who indicate long polling events is a suitable pattern. You want the best worlds of data durability and consistency through polling functionality, but also as close to real time event firing as possible.
Of course, some APIs you have no control over, and are stuck building robust, chatty polling infra to support because your (paying) users demand access to those APIs. Such is the schlep.
For consumers, I agree with most here that Kafka is certainly overkill. We've gotten away with a very simple architecture to have reliable event consumption. We point all webhooks to an (AWS) API Gateway backed by Lambdas. The Lambdas push the events to an SQS queue (FIFO-queue, if it needs some sort of sequence), and we take our time consuming the events through a very generic poll.
TBF they could do something similar with `/events`, instead of pushing events to a webhooks-sending queue just push them to the events buffer, which could even be a circular buffer just to point out that the essay is completely wrong. TFA is not asking for /events, they're asking for a very specific kind of /events with a large non-drained buffer. Something which would only ever work for low number of events: $dayjob's github integration takes in several events per second.
A proper event stream would be nice though, github's webhooks delivery system is not exactly reliable.
They can't offload webhooks at their own pace if the two parties want reliable delivery. The server providing the webhook might be experiencing a prolonged outage, in which case the sender needs the ability to buffer the events anyway.
A hanging GET is where you get chunked encoding (which is the only option in HTTP/2 anyways) and possibly never-ending stream. I've implemented a (proprietary, for now) "tail -f" over HTTP that does this when Range: bytes=0- (i.e., end offset not specified), completing the transfer (i.e., final, empty chunk sent) only whenever the file is removed or renamed away.
You'll want to add a heartbeat to any hanging GET /events.
Quite the opposite. HTTP servers and clients are essentially a solved problem. Massive scale-out, load balancing, retries, authentication, authorization, rolling deploys etc. can all be done out of the box by a hundred different providers. Anything to do with maintaining a large number of open TCP connections is still a massive pain on the server side.
With long-polling, now you're managing a custom daemon, basically. That's a big step down in reliability-by-default, and a bunch more work to do it right.
In either case, you'll be looking at more work if you want to check any kind of log on the other end for missed messages, but that looks pretty similar for either, and not all systems need that level of accuracy (and if they do, they probably need even more and this whole thing is Doing It Wrong)
One big issue with the pull-based model though is with concurrency. If you have multiple workers polling an API endpoint for new data, you need to synchronise the ‘last seen’ ID or timestamp across all workers. Otherwise, worker A and worker B might pull the same data and you could end up with duplicates. There’s no silver bullet here, either model requires work to harden against edge cases.
We're in the business of selling things, we're not in the business of building HTTP services. Our ERP-like thing is down for maintenance from 10pm to 11pm, so we can't use it as a platform for responding to webhooks.
I'll hack something together in a pinch when it's the only way to get orders from a marketplace service, but then eventually I'll have to explain to my colleagues how to linux, how to HTTPS, how to python, how to WSGI, and all other stuff our company typically doesn't do, but has to do now, because this particular marketplace wants to POST orders to us.
On the one hand, it seems like something too simple to expect people to pay for.
On the other, it's so simple it wouldn't be a huge loss to try it out and see if they will.
Can this be a 3rd party service? It certainly can be, but it's hard to make a generic one for any kind of webhook. Some marketplaces expect a dynamic response, like replying the order number we assigned internally (I typically just echo back the number they gave us with some prefix, but it's still more smarts than just replying with empty 200 OK).
And I've seen services which aggregate popular local marketplaces into API which is easier to work with, but they require to concede some other parts of the business we'd rather keep in-house, like assortment and inventory management.
Also, any SLA on uptime?
I recently had to work with an api to integrate a card reader terminal. The system was intended to work with an on premise pos so it was oriented around events sent back and forth over a websocket. Only one was allowed per restaurant.
In the end, I had to build a distributed cluster that managed a pool of processes that each spun up a websocket connection and relayed it to our pubsub system. Luckily I built out system in elixir so it was pretty easy to spin up hoarde but this would have been extremely difficult in any other language.
By contrast, a webhook is easy to scale from ANY language. Just make a http endpoint. Those are way easier to scale than a stateful process. Not being able to replay a webhook event is an issue with the publisher of the webhook. not a consumption issue.
That being said, I've also lobbied for a similar endpoint so I don't get support tickets for "missing data".
Both? Both are good.
One issue is that webhooks and HTTP clients can be pinned to a version, but the events listed at `/events` are whatever the Stripe account default version was at the time of the event creation.
So all of your code clients polling `/events` needs to ensure it can handle many different versions.
Another issue is that child lists are limited to 10 items, which means that you need to do a direct download to get list items > 10. This means the event list is lossy as items > 10 are never contained in the event stream.
Stripe feature request:
- An option to include all list items.
- `/events` that can be version-pinned / contains events with the same API version.
Gerrit stream-events doesn't solve the issue if the connection is dropped and events occur while disconnected.
I personally (for my hobbyist use case) would prefer the article's /event system over webhooks as webhooks require you have a system that is available on the Internet to receive the webhook. Where having this /event system would not require that.
1. webhook with retry up to 7 days. 2. a REST api to fetch all data
Sending webhook out properly require lots of effort, especially idempotent key concept to avoid duplicate data. And control concurency to avoid swarming the webhook endpoint.
So at the end of day, both are require same amount of resources, either on the sender side or the receiver side.
I wonder if the two approaches could be combined to simplify things for consumer apps at the cost of slightly more complexity on the producer side? Instead of POSTing the actual event data to webhook, the producer just uses consumer's webhook to "poke" it - to tell the consumer app "hey, you have new events waiting for you". On receiving the poke the consumer endpoint handler/PHP script can just turn around and do a GET to "/event" with anything > last downloaded event id query. That way you don't have to support long polling on the producer's servers and it's not a big problem if consumer misses couple of webhook "pokes". The next time it does receive a webhook "poke" successfully, it will download all the events and be all caught up. If real time notifications are not strictly required then producer side can even run the webhook dispatching code on a scheduled basis to coalesce multiple events in a single "poke" to a consumer to be more efficient, if desired.
For a great number of applications it's not that big of a deal to miss a webhook, and the extreme simplicity that it gives the developer is worth a great deal. With how enormously complex a lot of systems have gotten, I really favor simplicity whenever possible.
You do not need storage in the producer AND in the consumer, you just need a queue in the producer. Yes, even that is annoying, but the suggested architecture will still lose data if the long poller is down unless there is storage in the producer... so nothing significant has really been solved
Well the best advantage of Webhooks over polling is that you receive events straightaway no matter the volume of events. If you already know the volume and the volume is high, of course polling is going to be better for everyone.
It requires a sequence though, so that you know that the client can know that packet it just received isn't the one it expected, but you could build that with a chain of IDs like you're proposing.
To be correct with such a system you have to be prepared to queue the incoming webhook events, do the catchup query, then replay the queued events.
edit: Also, idempotent is not the right term here. Idempotent just would mean the event could be handled by the receiver multiple times w/o changing the meaning. If you need the events to be applicable out of order, then you need them to be commutative. This is a much more difficult property to ensure and in practice, I am guessing, almost non-existent in deployed webhook APIs.
Sure, there is no queue. There is an append only log. Kafka is not a queue but an append only log.
I do not like the proposed solution. I do not like it because it assumes that I have to maintain the infrastructure to do all the distributed logging on my end. As in, most likely I have to maintain a Kafka cluster.
There’s also one thing glossed over in this article. What if your consumer went past certain messages but it mishandled them? You can go to the past either.
If your consumer needs ordered delivery, web hooks might not be the best solution, indeed. But it might cost you more because I need additional infra to provide you with that.
If what you want is guaranteed state synchronization, pub/sub alone can't give it to you.
> In our integration with Stripe, it would be neat if we could request /events with a parameter indicating we wanted to long-poll. Given the cursor we send, if there were new events Stripe would return those immediately. But if there wasn't, Stripe could hold the request open until new events were created. When the request completes, we simply re-open it and repeat the cycle. This would not only mean we could get events as fast as possible, but would also reduce overall network traffic.
What they offer is a data plane and it makes sense. Although, it doesn't contradict the idea of webhooks, but rather complements it. A consumer can get notified via a webhook when new data is available. Whether the data itself comes with the webhook, or is available via an additional API request, is a matter of design. Personally, I like the idea of separating the control plane from the data plane. However, in some cases it can be an overkill.
Make the even ingestion job idempotent, so it can handle multiple receives, then hook up a pull / poll every 5 minutes or hour or whatever to start at the last pulled date (very important, don’t start at 5 minutes ago or you’ll miss downtimes). Then optionally set up webhooks or a push model.
EventSource is also nice because you can add cursors, so if the connection drops the source naturally catches up when the connection is re-established. While the interface is browser oriented, there's no reason not to use it in other contexts.
It says "This mode protects a table against concurrent data changes" but it does not elaborate how. Is it similar in consequences to what MariaDB describes?
In our case - active data export process, will slow down a bit main application while reading new events.
Best solution would be using external tool to handle streams, like Event Store, Kafka or new RabbitMq Streams. But we prefer to stick with Postgres.
Why not simply offer e.g. a public STOMP endpoint?
This is a huge helper for me personally.
The article talks about issues with webhooks such as not being reliable if the service goes down and messages are lost. It also talks about developers daisy-chaining multiple services together to put forward a solution which is not robust.
That's why you need a broker that does event distribution, supports multi-protocols (REST, AMQP, MQTT, WebSockets...) natively without any proxies and supports Webhooks. You can push messages to your REST clients and if they disconnect, the messages will pile up in a queue, ready to be consumed when the client reconnects.
Solace PubSub+ Broker does all of this. Disclaimer: I work at Solace.
we are a data and cpu intensive API - and moving from an client-request flow - to pushing to an API to get a pending response, then hitting their queue/webhook once the data is ready. Eventually i think only-once semantics found in kafka, other stream processing will become the norm and better alternative to webhooks, with the right security constructs.
giving write access to partners to a queue or even database connection to directly push data, may sound suspect today, but be common practice in a few years.
[0] hookdoo.com [1] hookdeck.com
Did anyone here understand what their value proposition even is?
Webhooks are perfectly fine for what they're intended, which is inconsistent push-based notifications to loosely coupled web apps. If you require "consistency", you supplement with polling and queues and other junk. If you require real consistency, you must use a distributed consensus algorithm.