Goliath: Non-blocking, Ruby 1.9 Web Server
igvita.com
igvita.com
How do I deploy Goliath in production?
* We recommend deploying Goliath behind a reverse proxy such as HAProxy, Nginx or equivalent. Using one of the above, you can easily run multiple instances of the same application and load balance between then within the reverse proxy.I am still wondering about how Goliath fits into both deployment architecture and application development. Traditionally these 2 have always been separated out.
* Thread safety. It is explicitly mentioned that middleware used must be thread safe. Doesn't this hold for all code?
* Can Goliath use multiple cores, or does one instance need to be spun up for each core?
* does it make sense to say, server a Sinatra app from Goliath?
Probably in a manner not entirely dissimilar to Unicorn: https://github.com/blog/517-unicorn
I generally do 2x the number of cores (e.g., eight Unicorn workers on a quadcore box), but it depends on your application. A heavyweight application with 150 models and lots of cpu-intensive tasks will suffer more from context switches than a lightweight one which spends most of its time idle.
As far as middleware goes, because you are reusing the same "app chain" between multiple requests, you just have to make sure that your middleware does not rely on any instance variables since those will get clobbered by other requests. Check the wiki page on middleware, we have a few examples around this.
Last but not least.. the EM reactor runs on a single core, so you're basically in the same deployment scenario as Thin or node.js. Having said that (always a caveat! ;)), Goliath can run on JRuby and Rubinius.. so in theory we have non-GILed environments there, which means we can start multiple reactors and run across multiple cores from within the same proc. This is not something I've experimented with in practice yet, but in theory its possible.
Having to choose between multiple threads and a better language is really not a decision you should have to make.
Rubinius is a little bit further behind, but I've been able to run our Goliath stack on it, and that exercises quiet a few syntactic changes. So, I think both are close.
For instance, I have a 20-line node.js app that does nothing but serve websockets for my rails app. I may replace that with Goliath in the interest of consistency.
Thin would be an alternative to Goliath as well, although we chose to switch to a different parser and also to add the Fiber logic/wrappers right into the framework.
The way I think about it is: if Thin is an app server, then Goliath is more of a minimal framework which you can use to go from start to finish in if you need to bring up an API endpoint. That includes, configuration, routing if you need it, validation, etc.
Now.. How you actually achieve that is a whole different story. You could, in theory, throw a job into some external work queue and poll that, or if your runtime permits, spawn some threadpool and periodically check that, or.. spawn a process and wait on that. In other words, it all depends on the actual operation.
In the case of PDF generation, if you rely on some external tool, you could use a mechanism like EM.system('shell cmd') to spawn a process and wait for it to return you the data.
I only make this pedantic comment because EventMachine makes it really easy to start down the path of "just event the process management and I/O", and it seems like you're almost always better off not doing that.
client request1 --> web server --> rpc server --> work queue
<process other requests>
client request1 <-- web server <-- rpc server <-- done queue
Of course you need to create consumers to operate on the work queue, process the jobs and put them in a different queue (eg. 'done queue') signaling the rpc server that you finished the job so that the rpc server in turn will reply back to the web server.
What I asked was whether these types of jobs can be done without the server blocking. igrigorik's answered that. The queue might still be needed but the polling can happen from inside the server code rather than from browser (the server will hold a connection from browser, keep polling queue and on success return the response)
Connections that you hold open while forking off and waiting for a process that will take many seconds to complete is just bad UX. There are better ways to do it.
You'll excuse me if after arguing with my response to your one-sentence "noob question: how do I make PDF generation asynchronous from an evented Ruby webserver" I am not chomping at the bit to get into a long architecture debate with you. I meant it: good luck with your design. Sorry I couldn't be more helpful.
Can someone enlighten me?
In contrast, a node user must specify a callback every time they make a query. I don't know if the library writer has to do anything special.
also noteworthy: postrank has been running ruby19 fiber'd webserver in production "for well over a year"?
cool!