Ruby - Handling 1 Million Concurrent Connections
github.com
github.com
This reminds me another post from a guy who measured and graphed JVM performance in adding (or multiplying, can't remember) numbers. And he concluded, that, well, JVM is fast at doing dumb arithmetic in a loop.)
C - Handling 1 Million Concurrent Connections
Java - Handling 1 Million Concurrent Connections
Javascript - Handling 1 Million Concurrent Connections
Go - Handling 1 Million Concurrent Connections
If it was 10^10 connections, now we are talking about some clever hacks to get that to work.continue with Ruby as backend for audio/video chats...
consider Ruby for streaming podcasts...
etc. etc.
http://blog.whatsapp.com/index.php/2012/01/1-million-is-so-2...
Here is the same in Erlang for reference (from a few years ago, I would be interested to see if there is a more efficient way now):
http://www.metabrew.com/article/a-million-user-comet-applica...
Re-establishing a million connections at once is going to be hard on the network - the million were built up over a period of time previously yet now they're being re-established Big Bang style.
It's more the impact side of the risk equation i'm thinking of than the probability.
EDIT: typo
The positive in the one big machine scenario is that you have potential to take strong efforts to keep it reliable. The advantage in the lots of machines scenario is that there is a better chance you have well tested failover solutions.
It is the combination of impact and risk that I am discussing.
A lookup. Get any reasonable big publicly available simple data-set and import it into any persistent storage (a sorted file is OK) and let a client perform a lookup for a row (preferably utf8 text ,)) then render a simple html table in response. As stupid as MVC.
or
Perform any simple lookup, but only for authenticated user (against some passwd-like file, to make it easy).
Then we could see how cool any JVM stuff or over-engineered OO ruby frameworks really are, with all that shinny graphs and smooth curves.
You can also open 1 million connections using one linux box a) increase local port range 1024 to 65535 b) setup 17 ip address. c) open from each ip address 58824 connections
The issue here is concurrency on the service software. If you have to launch a million instances to listen on a million different ports, you are doing it wrong.
So if you are on 192.168.1.1 and want to connect to a specific port on 192.168.1.2 there aren't enough free port numbers to get 1 million connections. Thus the extending of the "ephemeral port range" (the local port number the kernel is allowed to assign) and addition of more local IPs.
Note: This scenario is only valid for two computers talking to each other. As gilgoomesh said, if you have multiple clients you have virtually unlimited valid connection tuples (src addr, dst addr, src prt, dst prt).
I like what he has done but don't know why it made HN frontpage. I'm guessing it's because it has Ruby, 1 million, and concurrent in the title.
Not bashing the author, it's good work in its own right. I'm trying to understand how stuff like this makes frontpage HN. Where's the value of this article? That Ruby can handle 1 mil connections? Am I missing something (even if it's obvious?)
I think someone mentioned better real world test cases. I agree that would be a place to start. Perhaps he could send more meaningful data to clients. Maybe market data for some stocks or something. '\0' is not very useful after all!
Benchmarking experiments of this nature are actually remarkable because they are so rarely done. Most scaling and performance principles still arise from conjecture. It's not trivial to set up a test like this.
Also, _this_ is the place to start in preparation for a real world test. The next iteration, maybe a more realistic test, is only a fork away.
Pedantic, but "pedantic" actually idiomatically means "not important to note".
In my dictionary it's "of or like a pedant"; a pedant being someone who is overly concerned with minor details.
Still... noted!
That is EXACTLY THE POINT of this. We use Ruby for a ton of real work, and we are sick of seeing it get pummeled on unfairly for the misconception that it is a slow language and you are guaranteed to have scaling problems if you use it. Productivity and performance are not necessarily tradeoffs, with Ruby and some good planning, you can get both.
Ruby is a slow language, certainly considerably slower than c/c++/c#/scala/java/etc. The question is whether the performance gap is big enough in your particular use case for it to matter.
What the fuck is the point of opening a million connections if it takes you 1.55 hours to process one request from each connection?
179 means that while app holding and communicating to 1 million persistent connections it is still able to process 100+ standard requests per second. pretty enough to accept new clients to your online game, audio/video chat, podcast etc.
this graph shows how many requests per second app may process depending on amount of established persistent connections:
https://raw.github.com/slivu/1mc2/master/results/requests-pe...
Are there other messages being processed too? Which direction are the messages going? How big are they, what do they contain, how is the data processed. Also, if you have 1 million clients how do you handle all of them arriving at around the same time? How long did it take your system to get up to 1 million connections? Also how sure are you that each micro instance is sending and receiving the messages at exactly the expected rate and is not getting overloaded. What server type are you using for your central server?
If 'pretty enough' means only allowing 179 new connections to your server per second, that's great. When you have a large event, everyone gets on your site at the same time, not to mention times like getting to work, lunch break, evening rush, etc. You ever hear of the slashdot effect? That's more than 200 requests per second.
Ignoring the poor connection time, you could only have 179 users actively using your app every second. Out of a million. %0.000179 of your user base. Talk about really shitty user engagement.
I'm not even going to talk about the incredibly bad idea it is to host a million connections using one server. The idea that "most websites" only do "about 100 requests per second" is laughable. Sure, the average may be 100, over a month, but that's nothing compared to peak times. Try tens of thousands per second. A high-traffic site might do something on the order of thousands of database writes per second. Which, when all those connections come in, will kill your database servers, which backs up your frontends, which is why you have to have fast forward-facing pre-loaded cache. But I digress.
Focus on scaling your application to actually handle traffic before you obsess over concurrent connections.
http://urbanairship.com/blog/2010/08/24/c500k-in-action-at-u...
> Starting with Ruby 1.9.3, the GC in mainstream ruby can also be tuned
http://www.web-l.nl/posts/15-tuning-ruby-s-garbage-collector...