Web Framework Benchmarks Round 9
techempower.com
techempower.com
The Nimrod programming language is finally featured for the db tests (at least on i7. Not sure what happened with ec2 and peak, as neither jester nor nawak seem to appears in the results).
It fares pretty well when the database is involved.
Look for the nawak micro-framework, in the top 10 both for fortunes:
http://www.techempower.com/benchmarks/#section=data-r9&hw=i7...
and updates:
http://www.techempower.com/benchmarks/#section=data-r9&hw=i7... )
And there is room to grow. That micro-framework will not be the best for the json or plaintext tests, but once the database needs to be involved, it is trivial to add more concurrency: firing up more workers (1-3meg each in RAM) acts as an effective database connection pool (1 database connection per worker).
edit: Why should you care? Nawak (https://github.com/idlewan/nawak) is a clean micro-framework:
import nawak_mongrel, strutils
get "/":
return response("Hello World!")
get "/user/@username/?":
return response("Hello $1!" % url_params.username)
run()
Benchmark implementation in 100 lines here: https://github.com/TechEmpower/FrameworkBenchmarks/blob/mast...It's nice to see Nimrod making it high up on the list of the fortunes benchmark. I still have not had the time to properly optimise my framework, or to even write code for the other tests, but I am currently working on a new async module which will be a lot more competitive (an experimental version of it already made it into the latest Nimrod release, see http://nimrod-lang.org/news.html). In any case it's nice to see that others are creating their own Nimrod frameworks to compete with mine (thanks idlewan :) ).
If anyone has any questions about this round or thoughts for future rounds, please let me know!
Clearly, our project is about testing web application frameworks and platforms (we just say "frameworks" for conservation of typing). Testing NIO vs. InputStreams, for example, in isolation is a different project.
But as part of a web framework—let's say you had a framework that provided support for both NIO and InputStreams—we'd be happy to include the multiple permutations.
Obviously, this is within reason. If you submitted 512 permutations, I might have a small heart attack. Also, a general principal that we ask all participants to bear in mind is that each test should represent a viable production-grade configuration. This guideline might rule out a great number of wilder permutations.
But a question that keeps coming up in my mind is that there are metrics that would be much harder to compare, but might be more useful in my book.
For example, I'd love to see a "framework olympics" where different developers build an agreed upon application/website on an agreed upon server using their favorite framework. The application has to be of some decent complexity, and using tools that an average developer using the framework might use.
In the end, you could compare the complexity of the code, the average page response time, maintainability/flexibility, and the time it took to actually develop the app and the results could let developers know what they sacrifice or what they gain by using one framework over the other. I know a lot of these metrics could reflect the developer themselves vs the actual framework, but it might also be a tool to let you know what an average developer, given a weekend, might be able to produce. It would also help me to see an application written a ton of different ways -- so I can make good decisions about what framework to choose based on my needs.
In the end, speed only tells us so much -- and speed is not the only metric that we consider when we write applications -- otherwise it looks like most developers would be coding their web apps in Gemini.
Make it happen. :)
Getting an agreed upon application might be tough, but I'll try setting some stuff up to make it happen, as long as you agree you'll be a part of it :)
For the very first round, we included the relevant code snippets in the blog entry, and I think that added a lot to the context. With nearly 100 frameworks, the volume of code has become too large to simply embed all of that code directly into the blog entry, but we've put too much burden on the reader to sift through the GitHub repo to compare code complexity.
Things we aim to do:
* Pop-up iframe with relevant code from GitHub.
* Use a source lines of code (sloc) counter to render sloc alongside each result in the charts.
* Render the number of GitHub commits each test implementation has seen at our repository. At the very least, this would show whether a test has seen a lot of review.
* Introduce more complex test types [1].
And as the other reply has mentioned, we have also discussed the possibility of a larger test type that might include a multi-step process. I'd love to eventually get to that point.
[1] https://github.com/TechEmpower/FrameworkBenchmarks/issues/13...
But dont worry, it's still way faster than most PHP frameworks, that have ridiculous performances,yet their core developpers dont seem to care.
I think it really depends on the kind of app one is building. Video Site like Youtube can use caching to the max,most of the hits wont touch any server-side code,only cached pages. on the other hand a webapp that actually does something and need realtime capabilities might not be the right use case for Rails(like Twitter,though it helped them build their MVP quite fast,same for Iron.io).
It's a tradeoff, do you want to develop fast,at the cost of raw perfs,or get good perfs from the beginning without scalability issues at first place?
Most of the pool management code uses mutual exclusion which can get really expensive when 40 threads compete for it.
anybody know what's going on .. is a full PHP framework performing better than raw c++ under certain conditions?
[1] http://www.techempower.com/benchmarks/#section=data-r9&hw=pe...
Anyway, the tests show how hard it is to scale to many processors with a single process. It shows that the Java VM is really good at it (to be expected, Sun always had systems with insane numbers of cores).
Not really, it's more a benchmark of the runtime and its networking code probably. It would also be interesting to see a trace / summary of the syscalls for benchmarking runs.
They have a blog post about the results here: http://www.techempower.com/blog/2014/05/01/framework-benchma...
If you're running on EC2 and not dedicated hardware (probably most people reading this), be sure to toggle to the EC2 results at the top right of the benchmark.
These are the most realistic scenarios for the web app in my opinion.
It is quite disheartening to see it being 20 - 50x slower on the best. Or 2 - 5x Slower to other similar framework.
That said, if you need serving power, the fact that some some solutions can literally serve 100x more requests than others, and for likely less than 100x slowdown in development time and effort, may matter.
If on the other hand, you're building an application where performance is key, it's useful to know how much overhead your framework is adding. We have a real time bidding application which falls into this category, so these benchmarks quite interesting to me.
I am wondering if I should spend more time on Clojure instead, as it seems to be significantly higher in the ratings.
I was very surprised that Compojure did not score higher in the rankings. I have several deployed apps with Compojure and it seems very performant.
http://www.techempower.com/benchmarks/#section=data-r9&hw=pe...
https://github.com/TechEmpower/TFB-Round-9/tree/master/peak/...
https://github.com/TechEmpower/TFB-Round-9/tree/master/peak/...
Also can anyone share their experience using cpoll_cppsp?
For a typically webpage with multiple queries there appears to be around 5-10x performance disadvantage between slow and fast languages.Things like serving a plaintext or json response, where the slow languages are much much slower, Varnish is a good match for.
That said, we have considered including a measure of, say, nginx delivering the same content statically as a high-water mark. That could be interesting to see how well a framework compares to what might be an ideal case.
One thing is strange: The HHVM result on the "plaintext" test. How can HHVM only do 938 requests/sec if it can do 70,471 in the much more complicated "singly query" test?
However, http.sys is kernel mode in windows, and should have advantages that other programs don't have.
We're looking forward to having it back in Round 10, though!