Millions of active WebSockets with Node.js
unetworkingab.medium.com
unetworkingab.medium.com
Thats 65k per CLIENT IP address (i.e one per client IP:port pair). There should be no reason to use multiple IP addresses on the server.
If (for some reason??) you need each client to have more than 65k connections, you can add a port instead of an IP
> The client side has similar settings but does not need to set up multiple IP addresses, obviously.
You got this backwards and are now having to use a pool of server IPs to connect to instead of a single one...
How can this be fixed without more IP-addresses?
Simply put: if you had a webserver answering on port 443 then the backlog will be a number of prepared slots that all eventually will have 443 in them as the local port for the connection. You can have as many of those as you want. But only one process gets to listen to port 443 on any given IP.
---
Yes, it is listening on a single port, but accepted connections bind to a separate socket.
Here's how it works under the hood:
> The accept() call creates a new socket descriptor with the same properties as socket and returns it to the caller. [...] The new socket descriptor cannot be used to accept new connections. The original socket, socket, remains available to accept more connection requests.
Pulled this from random IBM zOS docs, but it's in compliance with the Unix standard: https://www.ibm.com/docs/en/zos/2.4.0?topic=functions-accept...
To further complicate the issue, any proxies or NAT in between may reduce the number of “clients” seen. A naughty ISP may only allow you to talk to 65 k sessions at once, or per data center.
And then there’s load balancers, which increase the number of ports in use if you only have one IP.
This article covers Node.js for me, I guess.
I don't see why there would be anything stopping Go from being similarly capable as it also has a good reputation for concurrency and what I hear does preemptive scheduling.
Java can probably do anything except be fun and lightweight so assuming you want to figure out the hoops to jump through. I assume it could..
Elixir can do it with the ergonomics and expressiveness of Python/Ruby. If you enjoy that level of abstraction I recommend it.
Personally I think just following the official guide [1] will give you all you need to get a taste of the language and the platform and decide if you like it or not.
If you were talking about websockets in particular I guess realistically most people use Phoenix Channels [2] that give you websockets in ten lines of code.
[0] https://elixir-lang.org/learning.html
[1] https://elixir-lang.org/getting-started/introduction.html
This talk by the same author is also a good introduction in video format: https://www.youtube.com/watch?v=JvBT4XBdoUE
Also versatile.
https://github.com/centrifugal/centrifugo (Server/Admin)
https://github.com/centrifugal/centrifuge (Server core)
https://github.com/centrifugal/centrifuge-js (Library)
It's a complete solution, including server, admin panel and client library.
As a generalization (again, really depends what you're going to be doing), I'd expect people to get a lot further with a Go or Java based implementations. Specifically, if those connections are interacting with each other in any meaningful way, I think shared data is still too useful to pass up.
I've written a websocket server implementation in Zig(1) and Elixir(2)
(1) https://github.com/karlseguin/websocket.zig (2) https://github.com/karlseguin/exws
What does this mean? What are some scenarios where connections interact with each other? I work with dotnet. To me, every request is standalone and doesn’t need to know any other request exists. At the most, I can see doing some kind of caching where if someone does a GET /person/12345 and someone else does the same, I maybe able to do some caching. However, I don’t think this is what you meant by shared data.
Did you mean like if someone does a PUT /person/12345/email hikingfan@gmail.com instead of the next get request reaching to the database, you keep it in the application memory and just use it?
Or am I completely missing the point and you’re talking about near real-time stuff like calls and screen sharing?
It was being done 7-8 years ago. If you search you should find a few articles on this.
Anyway - there's a lot of "non-standard" stuff in ASP's code there.
Also IMHO it's better to have a strong typed language behind your project, if it will be big, dynamic languages and big projects tend to be a nightmare for me.
* As a build/compile-time concern, using Node doesn't preclude strong-typing, so maintainability is also not a strong argument against the runtime itself, given you can use e.g. TypeScript.
> In fact, Pingora crashes are so rare we usually find unrelated issues when we do encounter one. Recently we discovered a kernel bug soon after our service started crashing. We've also discovered hardware issues on a few machines, in the past ruling out rare memory bugs caused by our software even after significant debugging was nearly impossible.
For sure not everyone will be able to achieve that in their first try or when getting started, but for sure is possible, but with Node I'm not confident enough to say that, for sure if works to hack something quickly and put in online, with Rust it takes longer and there are not too many platforms yet where you can easily deploy your app.
[0] https://blog.cloudflare.com/how-we-built-pingora-the-proxy-t...
Check out https://github.com/panjf2000/gnet, it also has some links at the end.
Here's me looking at the websocket traffic (I think): https://youtu.be/4rlffwHUchk?t=1857
I might not fully understand the technicals of it but I got it up and running and use it almost every day! :D Maybe I'll someday understand it.
Most of the work of accepting and holding concurrent TCP sockets done by the OS, not the language runtime. One can easily tune Linux kernel to 1M concurrent sockets.
The real issues: memory usage per concurrent socket (idle or active), and ability to do something useful with all these active connections, e.g. send pings every 30s, or broadcast a message to all of them.
I'm not sure NodeJS/C++ based system from this post will allow sending pings every 30s to 1M websockets, let alone to do something useful with them, beside some low-traffic or infrequent notifications (Of course one always need to perform realistic loadtests in order to answer these kinds of questions).
Erlang/Elixir/BEAM have a relatively large memory usage per active socket, but it allows doing something useful with them under an easy to use programming model (read: no callback hell).
I might be mistaken, since I'm not up-to-date with the multithreading/concurrency/isolates in nodejs/V8.
IMO async/await is error-prone and isn't an ideal programming model.
That doesn't mean anything. V8 is single-threaded but Node.js I/O is non-blocking. The reason Node became popular in the first place is that companies started adopting it to fill the gaps in their existing infrastructure (e.g. Java) to offer "realtime" (i.e. web sockets or its experimental equivalents) communication.
> IMO async/await is error-prone and isn't an ideal programming model.
What's the point of unsubstantiated statements like that (btw "x is error-prone" is an empirical claim, so prefixing it with "IMO" just means "I can't back this up and don't care if it's true") other than stirring up pointless language rivalries?
Just acknowledge that your off-hand comment about "callback hell" was anachronistic and don't try to come up with excuses to justify your preferences. I think Elixir is neat and hope it can see sufficient adoption for me to justify getting invested in it but that doesn't justify poopooing other languages, especially ones you admit not to have up-to-date knowledge about.
That's the problem with Software Engineering in general.
> "IMO" just means "I can't back this up and don't care if it's true")
The research study that will either back or refute my claim would be prohibitively expensive and therefore impractical.
E.g. "in my experience async/await can easily result in bugs that can be hard to detect" or "when teaching beginners, I've found that they have a harder time wrapping their head around async/await than when learning about agents" or "async/await still requires writing imperative code, requiring the programmer to pay attention to behavior that agents can abstract away through declarative code". Now, I don't know if any of those statements are true or if they reflect your experience but these are examples for what you could have said, assuming you didn't just want to say "I don't like JavaScript and I prefer Elixir or Erlang".
And yes, people just dumping strong opinions with little more substance than gut feelings is very much a problem in Software Engineering. That doesn't mean we can't work on that and practice a little more hygiene and respect for each other.
EDIT: To be clear, saying "I don't like X an I prefer Y or Z" is perfectly fine too as long as you are honest about this being your own preference rather than some grand truth about the universe. The problem comes from insisting that everyone else is wrong for not feeling the same way.
The Readme says:
> µWebSockets.js is a web server bypass for Node.js that reimplements eventing, networking, encryption, web protocols, routing and pub/sub in highly optimized C++
You do actually see the same kinda thing happening in other ecosystems, especially Python – something runs slower than you want it to (because Python), so you re-write it in Rust/C/C++ and hook into the extension system of the lang for a seamless DX. Off the top of my head things like `asyncpg`, `orson`, `uvloop` and esp relevant in this case `websockets` are all async libs that do this.
Mildly tangential but as someone who has spent a lot of time making the Python backend for my startup run as fast as possible, I would recommend avoiding Python if performance is your primary concern.
The actual library, as already pointed out, is written is C/C++ and bypasses Node for everything, other than providing an usage API layer.
Its the same as claiming that Tensorflow is written in Python ;)
Any developer out there who have been fascinated by and have used socket.io should already be using this.
Maybe someone with more swift experience can chime in! Whats the state of swift on the server these days?
What's the driving factor for that limitation?
> The client side has similar settings but does not need to set up multiple IP addresses, obviously
Also as I know, Linux will only use one IP for outbound connections, unless forcefully bind to another IP address in the code.
The server only uses only one port. 9001 according to the code. https://github.com/uNetworking/uWebSockets.js/blob/875f16e1f...
This blog post does not make any sense.
There's two ways to get more connections: use more client IPs or use more host IPs. In this post the OP has decided to add more host IPs.
In real world there will be more client IPs and the host will be listening on 1 IPv4 on 1 port.
so 2 * 16 == 65k