The C10K problem
kegel.com
kegel.com
In modern times, RAM and CPUs are more than 10 times bigger and faster, but I am seeing people get around 25 times LESS out of them, because they choose terrible tools, don't benchmark, and generally don't care. Now (different company) our app server pool has 150 nodes and each serves 4 users. The application complexity is significantly smaller than what I did 13 years ago.
I sincerely doubt any of my coworkers have read this document. It shows.
It's often a lot easier to rely on horizontal scaling, especially on a strategic level when you're deciding whether to trust your organization's technical ability (which really means management's hiring ability) or your ability to call AWS and spin up another hundred servers.
"Easier" doesn't mean "cheaper" or "faster" in the long run, but it is lower-variance for sure.
In other words, horizontal scalability means being _able_ to throw hardware at the problem :-)
Anyway, I agree with fiatmoney's sentiment that if you're in a corporate environment where your goal is growth, it's often best to focus on ensuring that your systems can grow along with demand, in a confident, easy, operations-free way -- as opposed to ensuring you've eked every bit of performance out of the hardware you have. In many business projects, the cost of human engineering time exceeds the cost of hardware.
Even when considering large-scale systems, it makes sense to scrutinize optimizations. If you had time to implement only one of these two equal-effort projects, which would you choose? (a) An optimization project expected to reduce a recurring yearly cost by $1 million (b) A new feature development project expected to earn an additional $2 million yearly recurring revenue?
I've found it effective to look at all engineering decisions as business decisions. From a business perspective, the mistake at the center of fallacies like premature optimization is the failure to consider opportunity cost.
(On the other hand sometimes pushing a technical project as far as you can is its own reward. Not everything in life has to be about business. Whether or not the results are widely relevant to typical businesses, I look forward to seeing the research into the C10K and C10M problems!)
Impossible to answer without knowing the gross margin of your business. For a big-box retailer (low margins), take the first. For a high-(gross)-margin business, take the second.
I would agree though that for startups, the problem is finding the place to $$$, it is important to know what the cost per user/connection/instance is and being able to control or monitor it.
With a bit of tweaking you can get the plain http case a bit higher, but the TLS route will not get much better, because TLS isn't written with speed in mind (the handshake is slow and expensive) and the prevalent implementation (OpenSSL) is not written for high performance servers (it basically dictates that you run blocking sockets in a thread per connection).
Unfortunately spdy and various http2 proposals rely on TLS (in order to punch trough proxies), which means going back in server performance about 10 years.
So it is of little surprise that companies have started offering "cloud" solutions, because the typical VPS can't handle todays high traffic internet over TLS (worse than plain http by a factor of 10) and the typical cloud server is worse than a VPS by a factor 5, creating a 50x performance degradation, artificially. Obviously when faced with the question of running 10 servers, or 100, most small companies turn to the even worse "cloud" solution, requiring even more servers (500).
The whole affair is a sodding mess, and we're wasting massive amounts of energy and capital on insisting on doing things inefficiently. This is because by rights our VPS servers should easily be able to break trough the 10k limit in every way, but it can't because the OS wastes a lot of time running an inefficient network stack as well as that TLS and OpenSSL can't be bothered to get their act together.
And that is how, in the year of our lord 2014, more than 15 years after somebody writing the 10k problem article, and after webservers becoming at least 128x faster then back in the 90ties, most websites out there can still not stand up to serious traffic, take ages to load, and are hosted by an infrastructure (routers and whatnot) that easily buckle under even light DDOSing.
Our servers should be able to surpass C1k easily, even C10k shouldn't tax them. They should, by rights, only be taxed by the C10m problem.
The latency of main memory read is still around 100ns. It has been around that for over 10 years now. It means your CPU will have to wait for hundreds of clock cycles to get a read from RAM if it's not in the cache, and in huge datasets it probably is not in cache.
Another issue is the processor clock speed. Yes it is true that modern i7 can reach 124850 MIPS. However that number comes from having 4 cores with each of them being able to reach up to 8 instructions per clock. You are still limited in executing dependent instructions.
That sounds a lot. But one must remember that it reaches 8 instructions per clock only when the instructions are a good mix of float/int instructions, no branches and the instructions are not dependent on eachother. In practice you reach maybe 1-2 instructions per clock. In some code it can go even to 0.5 IPC (bunch of unpredictable branches and whatnot).
Writing a code that takes advantage of large memory bandwidth and poor latency combined with massive CPU performance if the instructions are not too dependent on eachother is almost like writing modern GPU programs.
It would be interesting to see what kind of an web server perf one could get by carefully writing it in OpenCL (using CPU target, not GPU).
But, blaming bad I/O performance on large datasets misses the point slightly. You can perfectly well write a program that doesn't use heap, and has less stack use than those processors have L2 cache... (of course that's a test program).
But the network performance is probably not bound by system latency as much as by an abysmally bad software stack, started with the kernel to the networking stack to the sheer idea of TLS and to the implementation of TLS (OpenSSL).
Indeed, there's been calls to get rid of it all, all the layers and whatnot, and bann the OS from all but one or two cores and get rid of the whole network stack and layers and implement the networking directly in the application that needs to do it.
You certainly know building and supporting that would cost more or less the same as building and operating a sizeable datacenter. If it succeeds.
Using all processing power a modern CPU offers on real code with real data is almost impossible. And it's not only memory latency and instruction interdependence - there are latencies all over a PC even before you leave the rackmount chassis. The supporting network is another source of uncontrollable latencies. Most apps I manage spend 99.99% of their time waiting for something to happen, be it the next packet, be the results from another server, which is actually a cluster behind one or more load balancers.
You may get some better cache hit ratios by tweaking thread/core affinity, but it won't take you to where you want to be.
If you really need that much performance, I'd suggest building your own VLIW architecture and generate the instruction mix on auxiliary CPUs as a single continuous thread on the fly based on all incoming requests for the VLIW core to devour. That would be a huge undertaking, but it would also be pretty cool CompSci.
There are plenty of solutions for that already. For the simplest case of the user-space networking, you can pick a number of "off the shelf" solutions for it:
http://lukego.github.io/blog/2013/01/04/kernel-bypass-networ... http://www.openonload.org/
I run a few thousand simultaneous connections on AWS small instances without any issues. (HTTP + HTTPS).
The key is writing good software. Which a lot of programmers are still really bad at.
From a quick glance on your post, I too suspect there is something wrong - your performance shouldn't be that bad, but there is not enough information to point where.
But, fortunately it's relatively easy to test. You get whatever server you prefer, and install nginx and install a bunch of performance testing tools (like siege, ab etc.) and then you test different concurrency load scenarios.
No custom software, and widely regarded the fastest webserver out there.
Yeah there's your problem...
Tux may be a bit faster, but it can't serve dynamic content nearly as fast as it serves files, so that's pretty much useless.
And you've written a webserver and you can't.
Interesting...
Lets stick to the facts...
You stated that Amazon instances couldn't cope. They can. They can cope just fine with thousands of simultaneous HTTP and HTTPS connections. If you can't get your amazon instance to do this, you're using crappy software.
> "The 10k problem is still not solved. A run of the mill VPS typically manages around 1k simultaneous connections and 10k requests/s. That is, over plain http. Over TLS this goes down to about 100 connections and 600 requests/s. A cheap amazon instance will usually be around 1/5th of all that."
Where the hell are you getting those poor figures from? Are you using Lisp or something?
> "and the prevalent implementation (OpenSSL) is not written for high performance servers (it basically dictates that you run blocking sockets in a thread per connection)."
Again - you're using crappy software. Don't use it.
The 10k problem is historically interesting, but for anyone who knows how to program, it's not really an issue.
Seriously. Writing a webserver isn't a big job. It's not very hard to program something that runs better than the ones you list, for specific use cases.
http://en.wikipedia.org/wiki/Comparison_of_TLS_implementatio...
?
The C10M problem
It's time for servers to handle 10 million simultaneous connections, don't you think? After all, computers now have 1000 times the memory as 15 years ago when the first started handling 10 thousand connections.
Today (2013), $1200 will buy you a computer with 8 cores, 64 gigabytes of RAM, 10-gbps Ethernet, and a solid state drive. Such systems should be able to handle: - 10 million concurrent connections - 10 gigabits/second - 10 million packets/second - 10 microsecond latency - 10 microsecond jitter - 1 million connections/second
This article[2] breaks down the issues fairly well. In order to handle this kind of traffic, you have to basically redesign huge swaths of technology that exist because we don't want to have to implement these things more than once. I don't see how anyone would invest in this without a specific itch to be scratched (like deep packet inspection).
[1] http://www.cisco.com/web/about/security/intelligence/network...
[2] http://highscalability.com/blog/2013/5/13/the-secret-to-10-m...
Also I think for files/directories, you can listen for any changes that occur.
Wish this was officially available in Linux.
"See Nick Black's execellent Fast UNIX Servers page for a circa-2009 look at the situation."
http://web.archive.org/web/19990508164301/http://www.kegel.c...
I think Dan wrote this page originally when he was at Disappearing, Inc. Later he was at Google (2004-10).
https://www.readability.com/articles/zrgovuxr
TW;DR: Too Wide, Didn't Read.
It's obviously possible to make sites nicer using CSS, but if you're just trying to let people read your words rather than show off your design skills, leaving it to the browser's defaults is absolutely fine.