WebSockets vs. Server-Sent-Events vs. Long-Polling vs. WebRTC vs. WebTransport
rxdb.info
rxdb.info
edit: From the article: To workaround the limitation you have to use HTTP/2 or HTTP/3 with which the browser will only open a single connection per domain and then use multiplexing to run all data through a single connection.
I think the article calls this out. There is still a limit on the number of logical connections, but it's an order of magnitude larger.
https://github.com/w3c/ServiceWorker/issues/980#issuecomment...
Also unfortunately Chrome doesn't keep SharedWorker alive after a navigation (Firefox and Safari do):
https://issues.chromium.org/issues/40284712
Hopefully Chrome will fix this eventually, it really makes it hard to build performant MPAs.
Edit: of course you could use: https://caniuse.com/sharedworkers but android does not support it. We migrated to the lib because safari took its time… so mobile was/is not a thing for us
For HTTP 1, simply shard the domain.
Websockets get really complex to scale past a certain level of use.
Any day now: https://www.google.com/intl/en/ipv6/statistics.html
This might be the best thing about Elixir/Phoenix LiveView. I haven't actually had to care in quite some time :-) (though to be fair, I keep things over the websocket pretty light)
AWS you can use NAT Gateways for 6to4 and do v6 only subnets
I am not the one making a claim, I expected some arguments with it. "IPv4 is not web scale" is not good enough https://youtu.be/b2F-DItXtZs
I've been using IPv6 more recently, and one nice thing as a developer is being able to use the same IP address for local connections and internet connections. Simplifies managing TLS certs for example, since the IP address used by Let's Encrypt is the same one I'm connecting to while developing.
I guess maybe what GP is getting at is that with vhosts on IPv4 you need to have some sort of load balancer in order to share the IP, but with IPv6 you can flatten this out and give every host it's own IP?
> With ipv6 they can now be fully scaled easily but they are absolutely awesome, much easier to scale because you can give your client a simple list of sse services and its essentially stateless if done right.
If you don't understand it either please stop saying random stuff about IPv6 that we already know and has nothing to do with this thread.
A client gets an SSE endpoint (hostname). That endpoint maps to an IP. A server at that IP receives the connection. Which part is better with v6?
Are we talking about the few cents it would cost you to give each server an IPv4? Are we thinking about a distant future where that cost is not negligible compared to the cost of compute? Something else?
I don't use AWS so ok thanks. And if I did, I would use their ingress/gateway solutions, which totally circumvent this problem (while being quite expensive anyway).
What's your preferred provider?
But yes it's the superior (simplest, most robust, most performant and scalable) way to do real-time for eternity.
The browser is dead, but SSE will keep on doing work for native apps.
I wonder why they didn't just a multipart streamed response.
Supports my metadata, very commonly implemented format
Damn, that’s a huge downside
So depending on how much the phone needs to utilizes the radio, the higher the power level is?
That’s just my theory though.
I use K9-Mail app for email working 24h a day, it has multiple accounts on different IMAP4 servers. You know, IMAP requires one keep-alive socket per subscribed folder and I have no problem with battery usage.
I receive emails instantly. There is polling option in settings, I've disabled it.
WebSockets lack flow control (backpressure) and multiplexing, so if you need them you either roll your own or use something similar to RSocket.
Also SSE can't send binary data directly. You have to base64 encode it or similar.
WebTransport addresses these and also solves head of line blocking. But I'm concerned that we might run into a similar problem as we had with going from Python2 to Python3 and IPv6. Too easy for people to keep using the old version, and too little (perceived) benefit to upgrading.
As long as browsers still work with TCP, some networks will continue to block UDP (and thus HTTP3/WebTransport) outright.
Yes, head of line blocking is an issue, but TCP provides flow control, and if you're not using that, you're going over HTTP3.
> WebTransport addresses these and also solves head of line blocking. But I'm concerned that we might run into a similar problem as we had with going from Python2 to Python3 and IPv6. Too easy for people to keep using the old version, and too little (perceived) benefit to upgrading.
At one time or another, one could have said the same thing about TLS transport, HTTP3, or XHR itself. Because of the comparatively huge domination of a few key browser engines, it's much easier to roll out new browser capabilities & protocols.
> As long as browsers still work with TCP, some networks will continue to block UDP (and thus HTTP3/WebTransport) outright.
By that logic, as long as browsers still work with HTTP 1.1 without TLS, some networks will continue to block HTTP 2 and TLS. While that's not entirely incorrect, the broad adoption of HTTP2 and TLS in particular suggests it's less of a problem than you think.
Unfortunately, the way browsers implement WebSocket it undermines TCP's flow control. It's trivial to crash a browser tab by opening a (larger than RAM) file and trying to stream it to a server using a tight loop sending on a WebSocket. WebSocket.bufferedAmount exists, but as of 2019 I failed to use it to solve this problem and had to implement application-level backpressure.
> At one time or another, one could have said the same thing about TLS transport, HTTP3, or XHR itself. Because of the comparatively huge domination of a few key browser engines, it's much easier to roll out new browser capabilities & protocols.
> By that logic, as long as browsers still work with HTTP 1.1 without TLS, some networks will continue to block HTTP 2 and TLS. While that's not entirely incorrect, the broad adoption of HTTP2 and TLS in particular suggests it's less of a problem than you think.
HTTP3 actually falls under my concern. There are still networks that block HTTP3, because it has really nice fallback to HTTP2/1.1, so there's no obvious impact on users.
So I guess the real question is will QUIC be an HTTP/2 or an IPv6, or something in between? Was HTTP/2 ever actively blocked the way UDP is? If so that certainly gives us hope.
The reason I care is that I'm currently developing a protocol that WebTransport is an excellent fit for. But I can't assume WebTransport will work because UDP might be blocked, so I'm having to implement WebSocket support as well, which is a lot more work.
It wasn't particularly interesting to block to begin with. UDP if blocked, then it's blocked in a name of Security.
Before you were saying, "WebSockets lack flow control (backpressure) and multiplexing, so if you need them you either roll your own or use something similar to RSocket.", and now you're saying you can't roll your own? ;-)
> HTTP3 actually falls under my concern. There are still networks that block HTTP3, because it has really nice fallback to HTTP2/1.1, so there's no obvious impact on users.
Erm, HTTP2 & HTTP 1.1 have their own problems, some of which you yourself have identified. We actually rolled back from HTTP2 to HTTP 1.1 because of problems with HTTP2, particularly with mobile performance.
Our migration to HTTP3 has been all win so far. While UDP might be blocked for security reasons, there are security reasons to move to HTTP3.
That said, there are cases where only HTTP/1.0 is supported.
> The reason I care is that I'm currently developing a protocol that WebTransport is an excellent fit for. But I can't assume WebTransport will work because UDP might be blocked, so I'm having to implement WebSocket support as well, which is a lot more work.
I feel your pain. I've been using WebRTC, and the vast majority of the time UDP doesn't seem to be blocked anymore. That said, adoption of HTTP3 seems to be about half that of HTTP2 right now. Not bad for a brand new protocol, but I'd say we still have a significant amount of time to go before HTTP3 is the dominant protocol. I think the path forward is going to require some toil by the likes of you to support both, but no reason you can't support HTTP3 better! ;-)
You can, but you need to do it at the application level, which is a lot more involved. It would be nice if it were as simple as checking if bufferAmount > some threshold and then waiting on a promise before attempting to send again. That's essentially what you get with ReadableStream and WritableStream, which are provided by WebTransport.
> Erm, HTTP2 & HTTP 1.1 have their own problems, some of which you yourself have identified. We actually rolled back from HTTP2 to HTTP 1.1 because of problems with HTTP2, particularly with mobile performance.
Not sure if we're actually disagreeing here. In any case, we can both agree HTTP2 is not a panacea, and actually worse than HTTP/1.1 in some cases.
> That said, adoption of HTTP3 seems to be about half that of HTTP2 right now. Not bad for a brand new protocol, but I'd say we still have a significant amount of time to go before HTTP3 is the dominant protocol.
Here here. Overall I'm bullish on HTTP3 in the long run. I really just hope random enterprise networks don't decide to block it. In any case, it's going to be a big win for the places where it works.
I keep hearing it, but I've never actually seen such a network. There are many things that run on UDP. I can see it being closed in some tiny offices (but those usually lack brain power to accomplish it) or some dystopian corporate offices you can only see in a movie.
I really don't see how the fact that some networks might ban UDP has anything to do with it. Some networks ban google.com and wikipedia.com, you don't see them failing.
EDIT: According to this[1] issue, Berkeley guest wifi allows port 443, which would solve HTTP3/WebTransport. That's certainly hopeful, and I hope all networks are like this in the future. It's just not a forgone conclusion yet.
[0]: https://tailscale.com/blog/how-nat-traversal-works#have-you-...
If it works literally everywhere else, and the product is valuable, they'll change.
HTTP2 should still work in that scenario then you don't need to worry about multiplexing.
If you connect to a WebSocket over an HTTP2 connection then you don't need to worry about multiplexing since you can rely on the browser doing it for you - HTTP2 connections support over 200 concurrent streams.
I'm referring to opening multiple data streams on a single connection, which is very useful in a number of contexts. This is supported by HTTP/2 and HTTP/3, but it's only exposed directly to the browser runtime through WebTransport.
HTTP2 and HTTP3 connections support the multiplexing of hundreds of streams across a single TCP or QUIC connection. Each stream can be a request or WebSocket. So now you can just open a WebSocket per logical data stream and leave the multiplexing to the browser.
So now you will have a single transport level network connection whether you use multiple WebSockets or implement multiplexing yourself atop a a single WebSocket.
Note that WebTransport handles this for you, because the server can initiate new streams.
If you're sending lots of data, you also need per-stream backpressure.
I also think calling the WebTransport API complex is overblown. If you don't want the more advanced things, you can ignore them. If you want to use it like a WebSocket, just open one bidirectional stream and you're basically done. If you want to avoid head-of-line blocking, just open a stream for every message. It's a little more complex, but it's not the kind of thing you need a library for. Github Copilot will probably write the code for you. It's true there aren't as many server libraries out there yet, since WebTransport is still maturing. And we're waiting for Safari to add support.
Huh. The signalling server is implemented in websocket, typically.
It cant be implemented in webrtc unless you propose a existing decentralization of existing clients to boostrap'
devices switch off network or slow down etc,... for battery conservation, or when you don't explicitly do the I/O using a dedicated API for it.
new connection setup is a costly operation, the server has to store the state somewhere and when this stateful layer faces any issue, clients keep retrying and timing out. forever stuck on performing this costly operation. it's not like there is an easy way to control the throughput and slowly put the load on database
reliability wise long polling is the best one IME, if event based flow is really important, even then its better to have a 2 layer backend, where frontend does long polling on the 1st layer which then subscribes to websockets to the 2nd layer backend. much better control in terms of reliability
In my experience it even works great when the poll interval is long (for example 20 seconds) but when you also include the message list in each response. That way the client will be up to date when it interacts with the server: user presses a button -> the client sends a request to the server -> the server reponds with data and also a list of the latest messages.
I may be wrong here and the spec suggests a good way to do it, but i've seen so many different approaches that at this point might as well say there's none.
If you call to other domains, then this problem is no different to what we had with CORS years ago.
They're probably comparing it to the fetch and XHR APIs, which both allow custom headers.
Doesn't the initial request get to send a full set of standard HTTP headers, cookies and all?
What about intercepting the request with a service worker?
The even more irritating thing is that there is nothing preventing this, and every server I've tried supports it. It's only the browser WebSocket API that was designed without this. Cookies are the only thing browsers will deign to send in the initial request.
Beyond handling the custom headers aspect, it also supports any request method (POST, PATCH..), allows you to include a request body, allows subscribing to any named event (the EventSource `onmessage` vs `on('named event')` is very confusing), as well as setting an initial last event ID (which can be helpful when restoring state after a reload or similar). And you can use it as an async iterator.
I love the simplicity of Server-Sent Events, but the `EventSource` API seem to me like a rushed implementation that just kinda stuck around.
Another problem we've never worked out the solution to, is how to send a termination - signalling "there are no more events coming". We always end up having to roll our own, though it felt like something that should've been handled at the protocol layer.
Clients will reconnect if the connection is closed; a client can be told to stop reconnecting using the HTTP 204 No Content response code.
(This all assumes you only care about maintaining a connection when the tab is in the foreground.)
I’m wondering what problems people have run into when they tried this.
Yes, all the benefits of http/2 (or 3) are great, but we should also be aware of what we can take advantage of in http 1.1, especially since it's effectively universally supported.
Nobody reads the specs anymore, and to a certain degree I can't blame them, as the protocols/standards have become quite complicated.
Comet/sse/chunked transfer needs xhr to work. x-mixed-replace was würd back in the days and still is.
Edit: maybe you could also use an iframe/frame which holds a chunked connection but that will only give you text.
You can use a frame/iframes, but you can also just have content that is updated with multi-part MIME that doesn't cause the page layout to be redone.
with x-mixed-replace (as the name implies x-: experimentell) you can stream the page over and over again and the browser would change to the new version. (chrome still supports that for images, cheap webcams)
Tbf without frames neither mechanism made much sense, (even back than) because it would be horrible to use with form fields.
I started to play with the web in the early 2000s where xhr/long-polling/comet(via iframes, later it used xhr onreadystatechange with chuncked encoding, without sse, which basically was created because of that) started to gain traction and x-mixed-replace was extremely niche even back than because of the limitations it had on the page
Edit: comet (streaming script tags, inside an iframe) of course worked, back then. But I never heard of an implementation before 2006 (maybe a few years earlier like 2-3)or so, would’ve been worth a Wikipedia change if you would have old entries. Also http/1.1 was 97
Oh Server Push has all kinds of issues with it. There are lots of good reasons to prefer the newer protocols for a lot of use cases.
Frames were the common way to deal with that, but I even did stuff with having a separate "named" browser window.
> Also http/1.1 was 97
IIRC there was support for it in Netscape before it was really a standard.
That is probably true, since x-mixed-replace started the whole chunked transfer encoding.
It’s probably also what really triggered the invention of comet and later xhr. Netscape was way ahead and Microsoft just pushed it out with money, integration, activex? and of course unfair advantage.
To the OP, you can still build APIs with long polling. They are uncommon because push patterns are difficult to design well, regardless of protocol (whether long-polling, SSE, websockets, etc).
Whiteboarding a push API is a good exercise. There is a lot of nuance that gets overlooked in discussions whenever these patterns come up.
I still use it all the time. There are plenty of applications where the request overhead is reasonable in exchange for keeping everything within the context of an existing HTTP API.
The networking that makes Second Life go uses long polling HTTPS for an "event channel", over which the server can send event messages to the clients. Most messages go over UDP, but a few that need encryption or are large go over the HTTPS/TCP event channel.
At the client end, C++ clients use "libcurl". Its default timeout settings are not compatible with long polling. Libcurl will break connections and make another request. This can result in lost or duplicated messages.
At the server end, Apache front-ends the actual simulation servers, to filter out irrelevant connection attempts (Random HTTP attacks that try any open port, probably). Apache has its own timeouts, and will abort connections, forcing the client to retry.
There's a message serial number to try to prevent this mechanism from losing messages. The Second Life servers ignore the serial number the client sends back as a check. Some supposedly compatible servers from Open Simulator skip sequential numbers.
The end result is an HTTPS based system which can both lose and duplicate what were supposed to be reliable messages. Some of those messages, if lost, will stall out the user's activity in the game. The people who designed this are long gone. The current staff was unaware of how bad the mess is. Outside users had to find the problem and document it. The company staff has been trying to fix this for months. It seems to be difficult enough to fix that the current action is to defer work on the problem.
So, no, long polling is not "stupidly simple".
The right way to do this is probably to send a keep-alive message frequently enough that the TCP and HTTPS levels never time out. This keeps Apache and libcurl on their "happy paths", which work.
Really, anytime there is any form of push (whether SSE, long polling, etc) then you need another way to re-hydrate to the full state. In which case you are nearly at the point of doing plain old polling to sidestep the complexity of server-driven incremental updates and all the state coordination problems that entails.
Of course with polling, you lose responsiveness. For latency-sensitive applications (like an interactive mmorpg!) then HTTP is probably not the correct protocol to use.
It does sound like Second Life has its own special blend of weirdness on top of all that. Condolences to the engineers maintaining their systems.
Full Refresh; yes please, in the protocol, with a user button, with local client state cached client code and reloaded state on reconnect. Maybe even a configurable polling period; some services might offer shorter poll as a reason to pay for a higher tier account.
If the user ever has to push a "retry" button, the networking levels are very badly designed. Just because some crappy web sites work that way does not mean it's OK.
It can also be very helpful for out of band issues, like ISP hiccups, random hardware failures, bitflips, etc.
No, that's the whole point of long polling. The server delays the reply until it has something to say. Then it sends it immediately.
The trouble here is middleware which does not comprehend what's going on and introduces extraneous retry logic.
I mean, if you're not respecting long polling, of course long polling doesn't work. That's like complaining that http doesn't work because your networking stack doesn't look at port number and distributes packets randomly to any process.
What's your hit rate for the fallback?
[0] https://developer.mozilla.org/en-US/docs/Web/API/WebSocket
You should in most cases just use websockets with a keep-alive ping every 30 seconds or so. It's not common anymore to block websockets on firewalls, so fallback solutions like Faye/Socket.io are typically not needed anymore.
WebTransport can have lower latency. If you're sending voice data (outside of regular webrtc), or have a realtime game its something to consider.
WebTransport is a bit more work than other ones, like SSE, but the flexibility and performance make it work it IMO.
That means if you build something that requires web sockets, prepare to have a deluge of support/refund requests from the most valuable clients who think your site is broken.
I suggest just having a once-per-second polling fallback, perhaps with an info bar saying 'the network you are connected to is degrading your experience'.
We also see in these cases that streamed HTTP can also be broken by the firewall - for example a chunked response can be held back by the firewall and only forwarded to the client when the request ends, as a fixed-length response. Obviously that breaks SSE and means you can't just use streamed comet as a fallback when websockets don't work.
What I'd like to add is that Centrifugo also supports HTTP-streaming – not mentioned by the OP – but this is a transport which has advantages over Eventsource - like possibility to send POST body on initial request from web browser (with SSE you can not), it supports binary, and with Readable Streams browser API it's widely supported by modern browsers.
Another thing I'd like to mention about Centrifugo - it supports bidirectional WebSocket fallbacks with EventSource and HTTP-streaming, and does this without sticky sessions requirement in distributed scenario. I guess nobody else have this at this point. See https://centrifugal.dev/blog/2022/07/19/centrifugo-v4-releas.... Which solves one more practical concern. Sticky sessions is an optimization in Centrifugo case, not a requirement.
If you are interested in topic, we also have a post about WebSocket scalability - https://centrifugal.dev/blog/2020/11/12/scaling-websocket - it covers some design decisions made in Centrifugo.
You can also use it in a client/server setup. Check out 'WebRTC SFU'
I wrote a little bit about the different topologies in [0]
[0] https://webrtcforthecurious.com/docs/08-applied-webrtc/#webr...
I've come across https://github.com/soketi/soketi and https://centrifugal.dev/ but not sure if there are more battle-tested solutions.
They are building on top of Cloudflare, and getting started is a breeze.
That being said they are also fairly new, but based on everything I have seen, I am a fan
Disclosure: Pushpin lead dev.
In addition to the free and open source server, we also provide a cloud offering and on-premises versions that support clustering using Redis Streams, Kafka, Pulsar or Postgres LISTEN/NOTIFY as backends.
The solution is used by many big actors in production for years:
How Raven Controls uses Mercure to power big events such as Cop 21 and Euro 2020: https://api-platform.com/con/2022/conferences/real-time-and-...
Pushing 8 million Mercure notifications per day to run mail.tm: https://les-tilleuls.coop/en/blog/mail-tm-mercure-rocks-and-...
100,000 simultaneous Mercure users to power iGraal: https://speakerdeck.com/dunglas/mercure-real-time-for-php-ma...
Open Swoole is very easy to setup and there's lots of tutorials online. Got my ass kicked a little bit trying to making my websocket secure (wss) but I'm the end it worked fine.
- Is this a full or partial state update?
- What if the client misses an update?
- What if the client loses connectivity?
- How can the server detect and clean up clients that have disappeared?
How those are answered in turn raise more questions.
SSE or long polling or even WebSockets is a relatively unimportant implementation detail. IMO the bigger consideration should probably be ease-of-use and tooling interoperability. For that, I would say that long polling (or even just polling) is the clear winner.
It's a message. Mapping protocol units to messages was always your business.
> What if the client misses an update?
Sequence numbers on updates combined with a "fill in" mechanism through a separate request.
> What if the client loses connectivity?
Then more important things won't work either.
> How can the server detect and clean up clients that have disappeared?
The SSE client will restart dropped connections. You can have the server opportunistically close connections that haven't received messages recently. The browser will automatically reconnect if the object is still alive on the client side.
> For that, I would say that long polling (or even just polling) is the clear winner.
Coordinating polling intervals while simultaneously avoiding strong bursting behavior is genuinely not fun.
None of this favors SSE, or anything else for that matter. You still have to think about bursting / thundering herd behavior after server restarts, which would affect any protocol. The browser auto reconnecting doesn't mean you can assume the connection remains alive indefinitely if a reconnect hasn't happened; you may want periodic notifications as a keepalive to enable more immediate recovery. Long-lived server state introduces its own distributed coordination and cleanup problems, which strict request-reply polling sidesteps entirely. Etc.
There is no silver bullet, only tradeoffs in the design space that need to be matched to your application's requirements.
Then why is long polling the "clear winner?"
Sure you can find random libraries to help but they do need to be specifically built for SSE. Plain (long-)polling does not need any of that, which makes it more attractive from an API interoperability standpoint.
That makes no sense. Long-polling scales linearly like all the other ones as well.
- other approaches give you locality without sticky load-balancing: let's say your application server needs to subscribe to a topic once a connection is established, with long polling you need to setup and teardown that subscription every time, other approaches let you keep that HTTP stream alive and periodically send some stuff on it resulting in mostly just memory overhead.
- each returned payload will result in at least 1 extra packet (the initial request headers, assuming the response headers and the payload fit into a packet) and at least 1/2 RTT delay.
That said, I'd argue SSE is the browser equivalent of mobile push notifications.
Web Push is a distinct API for browsers to emulate native push notifications via service workers. https://developer.mozilla.org/en-US/docs/Web/API/Push_API
However, since the discussed techs all have major problems with mobile connections, I still think it should have been discussed in more depth.
In WebSockets, you can use plug&play extensions which can modify the payload on both the client and the server, which make them also ideal for tunneling and peer-to-peer applications.
I've written an article from an implementer's perspective a while ago, in case you are interested [1]
[1] https://cookie.engineer/weblog/articles/implementers-guide-t...
https://developer.chrome.com/blog/shared-dictionary-compress...