An Easy Way to Build Scalable Network Programs
blog.nodejs.org
blog.nodejs.org
Is it single threading? Is it the weird typeof crap you have to do to check if a variable is defined? Is it the lack of integers? Is it the prototyping system?
Node.js to me looks like a slightly better syntax than the horribly ugly C# async calls. (Not the new async/wait system). Javascript completely pales in comparison to F# or Haskell in terms of readability of async code.
If you prefer non-functional languages it would seem that Go would be a much better place to start for performance than Javascript. Or clojure or scala.
Sure node.js outperforms rails, but rails isn't designed around being the fastest webserver ever.
Many of the people currently "bashing" node are more or less aware of its capabilities. They aren't mad at the technology itself, just that it's being sold as much more than it is.
You don't see Apple plastering "ONLY AVAILABLE IN THE USA" on their pages about iCloud music storage for exactly this reason - you want to get people to say "Hmm, that sounds cool" before introducing them to the caveats.
The captive portal, DHCP, DNS, firewall/access management, VLAN tag/trunk management, route management, VPN tunnel management, SNMP Poller, SNMP server, network device controller services running on the internet access gateway for sports stadiums on commodity hardware running Linux. Or those same services running on a resource constrained system with significantly lower traffic.
This is obviously a scenario where Node.js is infeasible (web based portions had to have C bindings for anything outside of presentation layer and all parts of the systems required significant profiling and optimizations to meet speed and resource usage requirements). However, I've had several applications recently where Node.js could have been a possibility, but the documentation and overviews for it are either so beginner that you can't determine what it's capable of or is exploring areas where there isn't precedent for what it can do. I couldn't find a straightforward list of supported features and the ideal use case for it.
After expecting a more detailed article covering ideal usage of Node.js, and that not being it, I figure I would finally break down and ask "What is Node.js' ideal usage? Under what scenario is it an ideal tool?" It's not a knock against Node.js at all, maybe a slight one against the documentation, but not the technology or its use.
I use it to serve up several hosts on the same ip, or split sections of my url namespace into several node.js servers.
Maybe look at NodeJitsu blog as the articles seem readable, and those guys seem to use node heavily for running their hosting platform.
If it was because you don't know (and didn't want to learn) Javascript I could understand, but in terms of your comment above it seems like you just didn't look hard enough.
A fairly popular small DNS system was written in node for the purposes of development. Anyone who uses "pow" knows what I'm talking about: http://pow.cx
Transparent enough?
The suggested approach is to separate the I/O bound task of receiving uploads and serving downloads from the compute bound task of video encoding.
I'm assuming by using something like child_process.fork to create a video encode queue separate from the main event loop.
http://nodejs.org/docs/v0.5.4/api/child_processes.html#child...
That's kind of the point of the Node.JS bashing, once you work around all it's pitfalls you're right back where you started except your now writing your app in a language unsuited for the purpose.
Node solves the problem of needing to write evented servers in javascript. Beyond that I can't see much advantage in it vs. existing languages. If I wrote something called "Node.NET" which was a JScript wrapper around completion ports and went around telling everyone that this was the future of webdev... what do you think the reaction would be?
thank god node.js was there to save me all that work.
2. You don't want to fork on every request, this is very vulnerable to fork bombs.
3. In worker threads model, you already have the work threads spawned and ready for crunching, which will reduce the system load because you aren't forking on every request.
(I'm not saying they aren't or that I know the answer.)
There are certainly advantages and disadvantages to both threads and processes, but it's not really a fair comparison to claim that processes are as fast or faster than threads because you can spawn them at a certain rate. The performance cost of separate processes is something you pay gradually, every time you have to take a page fault and copy 4KB.
Then again, as someone else said, the amount of time it takes to spawn a process to encode a video relative to the amount of time it takes to do the encode is probably trivial, and you would benefit from having the process isolation in case something goes wonky.
You would probably write some sort of "process gate". Though in a distributed architecture I'd do this with some sort of distributed work queue.... I did this in .NET many years ago for a similar service: http://blog.jdconley.com/2007/09/asyncify-your-code.html
When Python forks a process using the multiprocess module, that process can execute concurrently with the parent process. On a mutlicore machine, it can be simultaneous.
When Python spawns a thread, the thread and the parent process cannot execute concurrently. They need to grab the Global Interpreter Lock (GIL). Whoever has it can execute. Whoever does not must wait.
So, I suspect that what we are seeing is that even though the new processes/threads have very little work, the processes can exit faster because they don't have to wait for the parent process to give up the GIL. This is a misguided experiment.
I reran this test with Python 2.7 and it no longer appears to be true:
Spawning 100 children with Thread took 0.03s
Spawning 100 children with Process took 0.28s
I'm not sure to what extent the GIL was improved in 2.7, but it's possible that it was never the cause to begin with.Regardless, I don't think it's a misguided experiment – it was an objective observation. It shows that things aren't so black and white depending on your toolchain.
Two, threads don't buy you parallelism in Python, unless the majority of the work is being done in C modules.
Finally, this test is really just testing the multiprocess and thread packages provided by Python. I say this is misguided because the way the author talks about it, I don't think he understands that the difference between those abstractions and OS threads and processes. (Which, of course, are an abstraction as well.) I suspect the Python overhead will be more than the difference in cost between forking OS-level threads and processes.
Internet trolls do improve software!
To work properly, it seems like an event-driven program really needs to use mlockall() and hopefully get memory pressure feedback from the kernel.