Scaling node.js to 100k concurrent connections
blog.caustik.com
blog.caustik.com
Once you're rendering views, etc it's hard to maintain 50,000 req/sec. (250K connections @ 5 req/sec)
Well, it means my product is going to be superior since I went the better, if less well known, architecture. It's not like erlang is a little unsupported side-project of a language, it's actually older then javascript if you count the time period before it was open-sourced, and only a few years younger if you don't, and is used extensively by many industries.
Also, just because javascript the language is more well-known doesn't mean javascript the server architecture is more well-known. I would argue it isn't; when people want a highly concurrent, solid server, erlang is always mentioned.
Lastly, erlang is a pretty easy language to learn. I had the basics down in a day, I had a prototype pubsub server that could handle 50k connections in two. The syntax is a bit strange, and honestly it does get in the way sometimes, but it's not hard.
Probably Erlang is "better". BTW, Java & C are quite fast also. Java applications, well written, do scale. If they don't, they are folks out there specialized to make them scale.
Also, Erlang is probably easy to learn as language. But when you develop a web app, you have enough other skills to keep up with. Let's name CSS for one ;) The human brain is limited in its capacity to remember API and language specifics.
Also, more popular means more libraries, which makes the product in turn better. This is why so many folks turn to PHP. It's not elegant, but everything you need is already here.
Now I won't argue that you may have good reasons to use Erlang for yourself, may it be because you like the language structure, like to write libraries by yourself, or so on. But it doesn't make it "superior", foremost not as a platform.
Perhaps, but for many people, node.js will be "good enough".
http://journal.dedasys.com/2006/02/18/maximizers-satisficers...
And because of the big community, there probably already are, or will be, more libraries available for it, meaning you have more ready-made blocks to build with.
I don't write this as a "supporter" of node.js, either - I've actually known and used Erlang for the past 8 years on and (mostly) off, and would highly encourage any hacker to have a look at it, because its way of doing things is quite enlightening, and, IMO, is superior to node.js.
They probably shouldn't be using node.js at work either.
Putting a client-side javascript engineer (even a decent one) on a node.js project can be really dangerous.
Evidently this wasn't your actual reason for thinking that, which just leaves me even more bewildered by your position.
Node.js is so easy to screw up, so difficult to debug, and little things can take down your entire application.
The combination of the language itself in addition to the type of people who typically would choose Node.js over other more proven options would make me worried that a blind choice is being made based on language alone and not proper evaluation or understanding of the other options available.
Personally I believe it is far better to be a language-agnostic company that thinks of different server side components as services, which might be in different languages instead of trying to use a tool just because they know the language already.
Of course there are many companies large and small that do it differently. But having someone whose responsibility includes client-side javascript but not server-side code is not by any means unusual outside the valley, at least IME.
We also tended to have companies with small teams.
http://www.metabrew.com/article/a-million-user-comet-applica...
I'd also add that this shouldn't be taken as something to say that Erlang is totally superior to node or that Erlang makes scaling to 1M concurrent connections a piece of cake. If you're working at that level, there's no magical out of the box solution.
And Erlang is not the only sane option for this kinds of problems, there is also Go. And to a lesser degree you can do the same in many other languages given the right libraries and careful thinking, it takes more effort than with Erlang or Go, but almost anything beats JavaScript in both performance and code clarity (both at the 'low' code-readability level, and at the high 'project organization and design' level).
http://shootout.alioth.debian.org/u32/which-programming-lang...
Javascript with V8 stacks up pretty well.
Let's say they are a full stack programmer who knows some html, some css, some javascript, some java, and some sql.
I have a pretty good idea how fast I can bring someone up to speed on node.js -- I have to teach them some advanced JS concepts, some node.js conventions, and the APIs of my library. Async takes a little bit to wrap your head around, but it's not terrible.
Node.js seems like it is on the way to "worse is better."
Probably another week to get up to speed with OTP for all the promises of resilient Erlang applications.
What is often missed about Erlang is that it's not really about highly concurrent applications. That property is actually a means to accomplish its primary goal: fault-tolerant applications.
See: http://www.erlang.org/download/armstrong_thesis_2003.pdf to understand the motives of the Erlang designers.
A week to learn Erlang. Where do I start?
You can pick up the basic syntax in one full day easily (if you are an experienced developer and already understand functional programming) ... 3 or 4 days if you are new to functional coding or just very inexperienced.
It is a very brief / minimalist language from a syntax point of view. Then, it will take a week or two to get your head around OTP, which is the primary framework and has years (decades?) of mission critical work under its belt.
Then, at the end of your journey will be the really hard problems... dealing with massive netsplits at a cluster level, elections for new masters, and all the other hard problems that happen at the upper-tier of massive clusters.
If you are building an HTTP(S) app -- you can blessedly avoid a lot of these by avoiding a true massive cluster all together and using lots of individual "micro clusters"(note) balanced / routed by HTTP middle-ware.
Also, check out Cowboy [webserver] (https://github.com/extend/cowboy/), Agner [package manager] (http://erlagner.org/) and of course, the always awesome rebar [build tool] (https://github.com/basho/rebar/)
.. and I love lager [log tool, make those erlang logs less alien looking] (http://basho.com/blog/technical/2011/07/20/Introducing-Lager...)
(note) This is basically a strategy of using small clusters based on locations -- so if you are across lets say 3 locations, you would build 3 node clusters, 1 node per location and have them work as a unit localizing workloads and responding to requests, and then you allow your higher level middle-ware to deal with your many groups of "micro clusters". High reliability rather cheaply, but means you need your own system for pushing out updates.
I'd really like to see a story of someone really having 100k connected browsers. My online game currently peaks at about 1000 concurrent connections, and node process rarely lasts longer than 2 hours before it crashes. Of course, using a db like Redis to keep users sessions makes the problem almost invisible to users, as restart is instantaneous. I'm using socket.io, express, crypto module, etc.
I'd really like to see real figures for node process uptime from someone having 5000+ concurrent connections.
However, I do use Redis for one thing: user sessions. I turned persistence off as Redis seems to be rock-stable, and I really don't need sessions to persist. I was using a modified version of Node's MemoryStore, to which I added clean garbage collection, but with often restarts I mentioned earlier it has become pain for users to have to login again when in the middle of the game. Having a separate, dedicated Redis instance to handle the sessions made restarts completely seamless, as the cookie sent by user's browser remains valid between node restarts.
I was not willing to learn new db technology, but there wasn't really much to learn with Redis. You can set it up in minutes and it just works(tm). I highly recommend you try it.
It took a while to iron out most cases that can crash, right now I have:
web.1: up for 12h
web.2: up for 12h
web.3: up for 12h
web.4: up for 12h
web.5: up for 12h
web.6: up for 4h
web.7: up for 1h
web.8: up for 12h
web.9: up for 12h
web.10: up for 12h
web.11: up for 32m
web.12: up for 12h
web.13: up for 7h
web.14: up for 7h
Regarding crashes, do you know of any special things to look out for? I do crash dumps and log uncaught exceptions, but sometimes node simply dies without any trace in the log files.
Most of the crashes come down to stupid things, it's so easy to make a mistake when you don't have a compiler watching your back. External dependencies can hurt if they're laggy or unavailable. Unterminated requests are a really easy accident as well.
At this point I just use exception catching and dump the results into Redis unless I'm specifically hunting down a bug and want the crash to occur:
Thanks for sharing.
Is it common practise to have node face the web without nginx?
We mostly use Erlang on server-side and node.js + CoffeScript on client-side (where they rightfully belong ;)
In my case I have entire db tables and collections replicated in memory and kept in sync via redis pubsub, and the 100,000s of concurrent users I have are all sharing just a few dozen persistant redis and mongodb connections between them.
I don't mean to minimize this accomplishment. If you're assuming you need 100k database connections in order to scale, you might be solving the wrong problem. Scaling is a matter of moving data as close to the CPU as possible. This means in-memory caching is where real performance comes in. I don't care how good your language/framework is, you can't defeat the physics of slow I/O over a network.