The hidden complexity of scaling WebSockets
composehq.com
composehq.com
ActionCable is Rails' WebSockets wrapper library, and it addresses basically every pain point in the post. However, it does so in a way that all Rails developers are using the same battle-tested solution. There's no need for every project to hack together its own proprietary approach.
Thundering herds, heartbeat monitoring are both covered.
If you need a messaging schema, I strongly recommend that you check out CableReady. It's a powerful library for triggering outcomes on the client. It ships with a large set of operations, but adding custom operations is trivial.
https://cableready.stimulusreflex.com/hello-world/
While both ActionCable and CableReady are Rails libraries, other frameworks would score huge wins if they adopted their client libraries.
send(payload).then(reply => ...)But client and server use it to maintain a record of events.
That solves none of the issues outlined in the post or the comments.
At its core, JSON-RPC boils down to "use `id` and `method` and work the rest out", which is acceptably minimal but does leave you with a lot of other issues to deal with.
What people seem to be often missing for some reason is that those two map naturally to existing semantics of the programming language they're already using.
What it means in practice is that you are exposing and consuming functions (ie. on classes) – just like you do in ordinary libraries.
In js/ts context it usually means async functions on classes annotated with decorators (to register method as rpc and perform runtime assertions) that are event emitters – all concepts already familiar to developers.
To summarize you have system that is easy to inspect, reason about and easy to use - almost like any other package in your dependency.
Introducing backward compatible/incompatible changes also becomes straight forward for everybody ie. following semver on api surface just like in any ordinary package you depend on.
Those straight forward facts are often missed and largely underappriciated.
ps. in our systems we're introducing two deviations – error code can also be strings, not just numbers (trivial); and we support async generators (emitting individual objects for array results) – which helps with head of line blocking issues for large resultsets (still compatible with jsonrpc at protocol level, although it would be nice if they supported it upstream as dedicated semantic in jsonrpc 2.1 or something). They could also specify registering and unregistering notification listeners at the spec level so everybody is using the same scheme.
The terminology is not ideal, I grant, but a JSON-RPC "notification" (a request with no id) is just a request where the client cannot, and does not, expect any response, not even a confirmation that the request was received and understood by the server. It's like UDP versus TCP.
> emitting individual objects for array results
This is interesting! How does this change the protocol? I assume it's more than just returning multiple responses for the same request?
Our implementation emits notifications for entries and rpc returns done payload (which is largely irrelevant just the fact of completion is relevant).
As I said it would be nice if they’d support generator functions at the protocol level.
Honestly, I'm a little surprised and more than a bit depressed how we effectively reinvent the OSI stack so often...
JSON-RPC is request-response only; the server cannot send unsolicited messages. gRPC supports bidirectional streaming, but I understand that setting it up is more complex than WebSockets.
I will concede that horizontal scaling of RPC is easier because there's no connection overhead.
Ultimately, it really depends on what you're trying to build. I also don't underestimate the cultural aspect; fair or not, JSON-RPC feels very "enterprise microservices" to me. If you think in schemas, RPC might be a good fit.
As an example, with a goroutine you have to be careful to handle all errors, because a panic would take down the whole service. In Elixir a websocket handler can crash anywhere without impacting the application. This comes at a cost, because to make this safe Elixir has to isolate the processes so they don't share memory, so each process has its own individual heap, and data gets copied around more often than in Go.
Unless you're the default `net/http` library and simply recover from the panic: https://github.com/golang/go/blob/master/src/net/http/server...
I found that scalability isn't a problem (it rarely is these days). The real problem is crappy network equipment all over the world that will sometimes break websockets in strange and mysterious ways. I guess not all network equipment vendors test with long-lived HTTP websocket connections with plenty of data going over them.
At a certain scale, this results in support requests, and frustratingly, I can't do anything about the problems my customers encounter.
The other problems are smaller, but still annoying, for example it isn't easy to compress content transmitted through websockets.
Everything works great in local/qa/test, and then once we move to production we inevitably have customers with super weird network security arrangements. Users in branch offices on WiFi hardware installed in 2007. That kind of thing.
When you are building software for other businesses to use, you need to keep it simple or the customer will make your life absolutely miserable.
It's so much easier to reason about than websockets, and a naive server side implementation is very simple.
A caveat is to only use them with HTTP 2 and/or client side logic to only have one connection open to the server, because of browser limits on simultaneous requests to the same origin.
[0] https://developer.mozilla.org/en-US/docs/Web/API/Server-sent... [1] https://developer.mozilla.org/en-US/docs/Web/API/EventSource
I've also discovered similar networking issues in my own application while traveling. For example, in Vietnam right now, I was facing recurring issues like long connection establishment times and loss of responsiveness mid-operation. I thought I was losing my mind - I even configured Caddy to not use HTTP3/QUIC (some networks don't like UDP).
I moved some chunkier messages in my app to HTTP requests, and it has become much more stable (though still iffy at times).
* Read, write and track all application-level state in a persistent data store.
* Identify sessions with a session token so that application-level sessions can span multiple WebSocket connections.
It's a lot easier to do this if your application-level protocol consists of a single discrete request and response (a la RPC). But you can also handle unidirectional/bidirectional streaming, as long as the stream states are tracked in your data store and on the client side.
Then again, the frontend and backend are a distributed system, so not that weird one comes to similar conclusions.
[1]: https://news.ycombinator.com/item?id=42813049 Every System is a Log: Avoiding coordination in distributed applications
I think part of the problem is that early systems wanted to eagerly process requests while they are still coming in. But in a system getting 100s of requests per second you get better concurrency if you wait for entire payloads before you waste cache lines on attempting to make forward progress on incomplete data. Which means you can divorce the concept of a payload entirely from how you acquired it.
At what point should one scale up & switch to chips with embedded DRAMs ("L4 cache")?
But you don’t get credit for having three tasks halfway finished instead of one task done and two in flight. Any failover will have to start over with no forward progress having been made.
ETA: while the chip generation used for EC2 m7i instances can have L4 cache, I can’t find a straight answer about whether they do or not.
What I can say is that for most of the services I benchmarked at my last gig, M7i came out to be as expensive per request as the m6’s on our workload (AMD’s was more expensive). So if it has L4 it ain’t helping. Especially at those price points.
Eg “The 1M challenge”
That said, even with great framework-level support, it's much, much harder to build a streaming functionality compared to plain request/response if you've got some notion of a "session".
[1]: https://www.canva.dev/blog/engineering/enabling-real-time-co...
This touches something that I think is starting to become understood- the concept of a "session backend" to address this kind of use case.
See the complexity of disaggregation a live session backend on AWS versus CloudFlare: https://digest.browsertech.com/archive/browsertech-digest-cl...
I wrote about session backends as distinct from durable execution: https://crabmusket.net/2024/durable-execution-versus-session...
I imagine the way it might go is that the client would first send an HTTP request to an endpoint that returns routing instructions, and then use that in the custom headers it sends when initiating the WebSocket connection.
Haven't tried this myself though.
Of course you do want to make sure the client has exponential backoff and jitter when reconnecting, as to avoid thundering herd problems.
The relevant state will need to be available to all servers as well. Anything that's only known to the original server will be lost as it drops connections. On a modern deployment with a database and probably redis available this isn't too big of an ask.
I have found for safety you need to allow an arbitrary delay of 100ms before killing sockets to ensure message completion which is likely why the protocol imposes a round trip of control frame opcode 8 before closing the connection the right way.
What? How would public network even know you’re running a websocket if you’re using TLS? I dont think it’s really possible in general case
> Since SSE is HTTP-based, it's much less likely to be blocked, providing a reliable alternative in restricted environments.
And websockets are not http-based?
What article describes as challenges seems like very pedestrian things that any rpc-based backend needs to solve.
The real reason websockets are hard to scale is because they pin state to a particular backend replica so if the whole bunch of them disconnect at scale the system might run out of resources trying to re-load all that state
I remember this one in particular making me upset, simply because of another extra buffer pass for security reasons that I believe are only to prevent proxies doing shit they never should have done in the first place?
It's a good idea for short-lived HTTP requests, but will cause problems for a persistent connection.
i don't hate them, they're great for what they are, but they're for realtime push of small messages only. trying to use them for the rest of your API as well just throws out all the great things about http - like caching and load balancing, and just normal request/response architecture. while you can use websockets for that it's only going to cause you headaches that are already solved by simply using a normal http api for the vast majority of your api.
been running a phoenix app in prod for 5 years. 1000+ paying customers. heavy use of websockets. never had an issue with the channels systems. it does what it says on the tin and works great right out of the box
Why would I use websockets over SSE?
For a bit of background, in order to tackle scalability, the initial approach was to explore serverless architecture. While there are both advantages and disadvantages to serverless, a notable issue with WebSockets on AWS* is that every time a message is received, it invokes a function. Similarly, sending a message to a WebSocket requires invoking an HTTP call to their gateway with the websocket / channel id.
The upside of this approach is that you get out-of-the-box scalability by dividing your code into functions and building things in a distributed fashion. The downside is latency, due to all the extra network hops.
This is where vramework comes in. It allows you to define a few functions (e.g., onConnect, onDisconnect, onCertainMessage) and provides the flexibility to run them locally using libraries like uws, ws, or socket.io, or deploy them in the cloud via AWS or Cloudflare (currently supported).
When running locally, the event bus operates locally as well, eliminating latency issues. If you apply the same framework to serverless, latency increases, but you gain scalability for free.
Additionally, vramework provides the following features:
- Standard Tooling
Each message is validated against its typescript signature at runtime. Any errors are caught and sent to the client. (Note: The error-handling mechanism has not yet been given much thought into as an API). Rate limiting is also incorporated as part of the permissioning system (each message can have permissions checked, one of them could rate limiting)
- Per-Message Authentication
It guards against abuse by ensuring that each message is valid for the user before processing it. For example, you can configure the framework to allow unauthenticated messages for certain actions like authentication or ping/pong, while requiring authentication for others.
- User Sessions
Another key feature is the ability to associate each message with a user session. This is essential not only for authentication but also for the actual functionality of the application. This is done by doing a call to a cache (optionally) which returns the user session associated with the websocket. This session can be updated during the websocket lifetime if needed (if your protocol deals with auth as part of it's messages and not on connection)
Some doc links:
https://vramework.dev/docs/channels/channel-intro
A post that explains vramework.dev a bit more in depth (linked directly to a code example for websockets):
https://presentation.vramework.dev/#/33/0/5
And one last thing, it also produces a fully typed websocket client, so if using routes (where a property in your message indicates which function to use, the approach AWS uses serverless).
Would love to get thoughts and feedback on this!
edit: *and potentially Cloudflare, though I’m not entirely sure of its internal workings, just the Hibernation server and optimising for cost saving
const onConnect: ChannelConnection<'hello!'> = async (services, channel) => {
// On connection (like onOpen)
channel.send('hello') // This is checked against the input type
}
const onDisconnect: ChannelDisconnection = async (services, channel) => {
// On close
// This can't send anything since channel closed
}
const onMessage: ChannelMessage<'hello!' | { name: string }, 'hey'> = async (services, channel) => {
channel.send('hey')
}
export const subscribeToLikes: ChannelMessage<
{ talkId: string; action: 'subscribeToLikes' },
{ action: string; likes: number }
> = async (services, channel, { action, talkId }) => {
const channelName = services.talks.getChannelName(talkId)
// This is a service that implements a pubsub/eventhub interface
await services.eventHub.subscribe(channelName, channel.channelId)
// we return the action since the frontend can use it to route to specific listeners as well (this could be absorbed by vrameworks runtime in future)
return { action, likes: await services.talks.getLikes(talkId) }
}
addChannel({
name: 'talks',
route: '/',
auth: true,
onConnect,
onDisconnect,
// Default message handler
onMessage,
// This will route the message to the correct function if a property action exists with the value subscribeToLikes (or otherwise)
onMessageRoute: {
action: {
subscribeToLikes: {
func: subscribeToLikes,
permissions: {
isTalkMember: [isTalkMember, isNotPresenter],
isAdmin
},
},
},
},
})
A code example.Worth noting you can share functions across websockets as well, which allows you to compose logic across different ones if needed
some of the complexity is self-inflected by ignoring KISS principle