Speeding up a Rails request by 150ms by changing 1 line
kevin.scaldeferri.com
kevin.scaldeferri.com
The regular occurrence of such issues in a project implies an immature, poorly profiled code base, and are the kinds of things that other projects fix without fanfare.
Our most recent database-backed dynamic web application can serve >10000 requests/sec on our mid-range hardware at <2 ms/req, which we achieved by running the most basic of profiling early on and throughout the development process, and using software that does the same. I don't expect a gold star for this.
Bragging about how fast you can serve a request is really meaningless when there's no way to know what's involved in that request. On my personal side projects (also using Rails), I serve requests in a couple milliseconds too. But, there's a vast, vast difference in what those sites do, vs what's involved in running a large-scale e-commerce operation.
There are quite a few stories of this nature, which could simply be summarized as individuals finding incredibly simple inefficiencies and then fixing them.
But it sounds like you were doing profiling and optimization early and often, which might have been premature, and thus might have been a waste of time.
Who's to say? Ruby isn't for everyone. Perhaps your ideal web stack is what you're using. (What're you using, btw?) I've been itching to use Haskell for a webapp.
He has stated in previous comments that he's using scala.
Interestingly, merb on jruby has reportedly reached 14000 rps on similar hardware to what he's refering to (8 cores, since he previously claimed to get 5000 rps on a 4 core machine).
Whether Rails is immature doesn't make it less interesting to see how people track down its problems and fix them.
I think a similar thing happened about a year ago when people were profiling rails but not profiling memory usage. After they realised GC was impacting their performance they shifted their focus and improved things in that regard. The most recent "X line Y% speedup" article I read here took this even further and examined the impact of the size of stack vs. heap memory on GC times.
My particular favourite example of this is Steve Souders work showing how much your web app performance is affected by the end users browser and its wacky network and cache behaviour. He has claimed these factors can account for 80-90% of the end-users' waiting time. I know rails core have been working to fix these issues as they did a presentation about.
That sounds to me like a bunch of useful lessons to be learned from. Certainly more info than just "profiling good", more like non-holistic profiling can be misleading. Your response certainly doesn't cover this as you talk about "basic profiling", so what is outside the profiling you do? How would you know if it is affecting performance and to what degree?
Properly profiling the full execution path should be an entirely obvious step to take.
Seriously, is it that hard to pull up dtrace and do a little analysis? They even have dtrace-ruby bindings on MacOS! http://www.infoq.com/news/2007/10/ruby-leopard
I know next to nothing about Rails, so I'm sorry if these are questions with obvious answers. I just think the decision to process body responses on a per-line basis seems odd.
What's frustrating is that apparently there wasn't any performance testing of the 2.3 release that ought to have found this problem. It's been in the code base for about a year at this point.
My guess as to the reasoning here is so that large files can be served without blocking while the whole thing is written to the socket. Granted, that's just my gut instinct looking at that code, maybe it is just careless...
In the article the author mentions it has to do with serving static files, which is something that nobody should be using Rails to do anyway. I'm not surprised nobody found it until now.
It's true, that the problem only becomes really glaring when the pages being served are quite large (this example was 1300 lines of HTML), but I still assert that good performance regression testing ought to have popped this up. With 100 lines of output, which is pretty reasonable, you'd still be looking at an extra 10ms on each request.
http://is.gd/xP8Y (github link)
http://github.com/josh/rack/tree/master :
* NOTE: String bodies break in 1.9, use an Array consisting of a single String instead.
p.s. this is definitely worth a 2.3.3 tag and release
Not wishing to troll, but I'm not sure that Ruby/Rails is you're looking for if you're obsessed with performance...
Profilers have been around for what? 30 years? I think the first time I hooked Ants Profiler up to an ASP.NET site was 2002. Are these other web technologies really that far behind the state of the art?
As that quote states, there doesn't appear to be a profiler as such, but rather a set of "timing and logging components" built into the framework. If it's true, then ruby today is essentially handling profiling the same way PHP did it in 1998. I find that surprising and saddening.
Also, what is the way PHP handles profiling in 2009 ? I haven't developed in php in quite a while, and I admit I never used anything more complicated then the profiling tools built in Zend Studio.
You're talking about a separate application that watches executing code and keeps statistics, right? As in, not simply a set of timing functions in a wrapper somewhere in the framework? If so, how does it come off the trolley after Rails finishes its job?