Maintaining 65k open connections in a single Ruby process
wjwh.eu
wjwh.eu
That's not right. Connections are identified by the 4-tuple of (source ip, source port, destination ip, destination port). The server ip and port are fixed in this example, but there's still 2^48 combinations of client ip/port.
> It is amazing that a problem which took huge engineering efforts to solve back in the early 2000s can now be solved easily by anyone in just a handful lines of code.
But you didn't solve the problem! The whole point of C10K was the need for scalable ways to figure out which connections could be operated on (i.e. had something to receive, or could be written to after having previously been blocked for writes). Keeping an array of sockets and just repeatedly iterating over it with the assumption that all those sockets are active is the opposite of scalable.
You are definitely right that this is not a "full" solution for the c10k problem, but it is an interesting start IMO. Maybe for next month I'll implement basic epoll functionality to make an echo server or something :)
{Protocol, PeerA, PeerB, PortA, PortB} ex:
{tcp, 192.0.2.1, 192.0.2.2, 80, 12345} is distinct from {tcp, 192.0.2.1, 192.0.2.3, 80, 12345}
With roughly 2^32 ipv4 addresses and 2^16 client ports for each ip, you should be able to establish 2^48 connections to your server IP. Although, it usually requires a bit more than standard effort to actually use quite that many client ports per IP, by default you have to fight against high vs low ports, and anyway not a lot of OSes have a very sane allocation policy for ports.
Note also: if you're testing this on localhost, IANA has assigned you 2^24 ipv4 addresses for localhost, 127.0.0.0/8, so if i assume you're going to listen on all addresses, and connect to all addresses from all addresses, that gives you 2^24 x ( 2^24 x 2^16 ) = 2^64 connections. It might take you some time to iterate through an array that large :)
32b IP * 16 bit port?
> You are definitely right that this is not a "full" solution for the c10k problem, but it is an interesting start IMO.
It's not a solution to the C10K at all, because the C10K problem is outdated and was trivialised a decade back. Whatsapp was doing 2 million connections on a single box (with resources to spare) back in 2012: https://blog.whatsapp.com/196/1-million-is-so-2011 and folks were working on C10M: http://highscalability.com/blog/2013/5/13/the-secret-to-10-m...
> Maybe for next month I'll implement basic epoll functionality to make an echo server or something :)
Epoll is bad, why would you want to do that?
Source please! I'm genuine - what is bad about it?
Google comparisons with kqueue from BSD and iocp on Windows.
Edit: comment below clarifies servers.
So 4 source ips. Guessing they were just bumping into a (configurable) max open file setting.
Total computer power - Cost of maintaining sockets = Available computer power to use sockets.
If total computer power - cost of maintaining sockets is close to 0, there's nothing one can really do with those sockets. If the system has 65K sockets open is it's barely breaking a sweat, you can still do plenty of work!
Edit: As masklin referenced in reply to me[1] that jsnell pointed out, that's the ports for a single IP. So yeah, it's not all the ports, but it is all the ports for a single source IP, unless I'm missing something. You can of course assign multiple source IPs to extend this, but it is a limit to be dealt with.
1: And then deleted it, at least as of now, even though I think it was a good point.
http://pubs.opengroup.org/onlinepubs/9699919799/functions/ac...
http://pubs.opengroup.org/onlinepubs/9699919799/functions/li...
Beej's guide is a popular intro, but I'd bet any text on network programming covers sockets.
It makes much more sense, and in fact I'd been pondering that since I commented as how that could work in practice with servers that actually keep connections open but can handle a lot of requests wasn't obvious in my prior mental model.
> Beej's guide is a popular intro, but I'd bet any text on network programming covers sockets.
I highly suspect I knew this at some point, but forgot it over the past 15-20 years since it didn't have direct relevance to most of my projects since that time. :/
While a best effort has been made to mimic the Linux semantics, there are some semantics that are too peculiar or ill-conceived to merit accommodation. In particular, the Linux epoll facility will -- by design -- continue to generate events for closed file descriptors where/when the underlying file description remains open. For example, if one were to fork(2) and subsequently close an actively epoll'd file descriptor in the parent, any events generated in the child on the implicitly duplicated file descriptor will continue to be delivered to the parent -- despite the fact that the parent itself no longer has any notion of the file description! This epoll facility refuses to honor these semantics; closing the EPOLL_CTL_ADD'd file descriptor will always result in no further events being generated for that event description.
Brian Cantrill on epoll https://youtu.be/l6XQUciI-Sc?t=3424
With MRI Ruby you'll still effectively be limited to one core due to the GIL, so if you need to do any compute intensive work for each socket instead of just sending a string you'll rapidly run into that barrier. As I state in the post, memory usage is not all that great either.
You mean kqueue, and that was in 2000. Epoll came later, and is bad.