Web Framework Benchmarks Round 7
techempower.com
techempower.com
Round 6: https://news.ycombinator.com/item?id=5979766
Round 5: https://news.ycombinator.com/item?id=5727012
Round 4: https://news.ycombinator.com/item?id=5644880
Round 3: https://news.ycombinator.com/item?id=5573532
Round 2: https://news.ycombinator.com/item?id=5498869
Round 1: https://news.ycombinator.com/item?id=5454775
Lots of solid discussion.
Honestly, now with all the great Scala frameworks, Clojure, and the ability to run Rails, plus Cassandra, Storm, etc, I'm a little creeped out that I'm actually strongly considering building my current new project completely on the JVM.
It is perfectly fine to not squeeze every drop from your hardware and pick Go or Lua. In the same vain, Flask may be x20 slower at a meaningless benchmark but you'd note that the ones that involve actual work, you know, the kind your complex app will actually perform, the more involved it is the more the difference shrinks. All the way down to x4, in other words not an order of magnitutde.
The real takeaway summary should probably be: a well made framework like flask doesn't have so much overhead that the productivity gains it offers are not worth trading a bit of performance for.
A terribly made framework (I won't name names) will bite you in the ass. Choose carefully.
I'll take that any day, over guessing which framework is fastest.
Rails, for example, despite always having been "slow," performs far more than adequately for the vast majority of web applications, scaling patterns are well established, and it's obviously hugely productive for many, many developers and companies.
Facebook and Twitter seem to be doing alright. Care to elaborate?
We have recently released an online market for physical gold trading capable of handling 10k+ concurrent users with horizontally scalable architecture and complex trading engine/accounting logic, and it's all Java. It has to be fast and there is not enough static or almost-static content to make caching effective.
As a related aside, we have an intention to eventually capture some additional statistics about the implementations such as source lines of code, total number of commits, and possibly lines of code of libraries (where available).
These are just additional data points to use as the reader sees fit, but they may provide some insights. For example: (a) developer efficiency, assuming you are comfortable using sloc as a proxy for developer efficiency, which is admittedly a hotly debated matter; (b) commits are a proxy for the level of attention the particular test implementation has received in our project, since a test with only 1 commit may be unrefined and in need of tuning and a test with dozens of commits may be considered highly volatile; (c) where applicable, the performance impact of additional code, perhaps even as granular as a calculated average per-line cost in rps. The last point might be particularly illuminating for languages such as PHP.
Doesn't look to be "an order of magnitude faster" - 200% faster in some cases - certainly nice numbers, but not a massive game changer for many (yet?).
And Java/Scala are disqualified pretty quickly.
I'm currently working on a new project (still research phase) where I was pondering going with Spray and a lightweight Postgresql wrapper because I primarily need to read data from a database, do some transformations on it, and write it out as JSON as fast as possible. I had it working in Spray, but I had issues with server crashes and the speed wasn't as I'd have expected. I fiddled with it for one day, and then, out of frustration, decided to give openresty a try. I've never written much Lua in my life, but after only a couple of hours, I had it working, and it was far, far faster than the Spray implementation. I did some research there, and it seems that the database stuff took a whole lot longer in Scala/Spray than in Lua. Now, of course, I loose type safety, so there may be hidden issues in there, but since I'm really just doing simple data transformations, I think I'm fine with lua / openresty.
As a first verdict, I really, really like what I've seen of Openresty so far.
Sorry, round 7 isn't helping much. Wait for round 8.
This is significant because, on round 6, Revel was performing better than raw Go in at least some of the test scenarios, which is not something you can say about many frameworks.
Was there ever an explanation why Go without any frameworks was performing worse than Revel? Simply because not much effort had been put into optimizing the Go code?
I've been too busy with work to look at the tests but my goal is to make a proper showing in round #8!
Yes, if you have time, please do help us resolve the issue (and I am guessing it is some configuration issue) for Round 8.
Slightly sad to see the Mono performance is still abysmal.
Edit: link to the blog. It was submitted earlier, but apparently got killed:
http://www.techempower.com/blog/2013/10/31/framework-benchma...
I'm impressed by the Java options. Definitely something to consider if absolute performance is the requirement.
[0] https://github.com/TechEmpower/FrameworkBenchmarks/pull/272#... [1] https://github.com/TechEmpower/FrameworkBenchmarks/issues/31... [2] https://github.com/TechEmpower/FrameworkBenchmarks/pull/339#...
In my own testing there was another ~60% performance to gain with better tweaks.
[1] https://github.com/damianh/FrameworkBenchmarks/blob/nowin/no...
If you notice the specific filters I've set, PHP actually outperforms Go for many cases. That's actually nice to see (PHP gets better with each release version). Well, if there was something like Coffeescript, but for PHP, then there's no better time than now to make it.
I really love the simplicity of Ruby, the performance of Golang/JVM and the massive popularity of PHP and it's compatibility with budget hardware (Shared hosting, etc). If there was a "Write your code in Ruby and we'll compile it to PHP" kind of generator/converter/service, that would be so awesome :D
(If anyone knows something like that that already exists, then I would be extremely thankful).
The one PHP framework that seems to do well is yaf. However, Yaf is actually written in C and is a php extension. I would expect it to do better than all the php frameworks, but that doesn't account for the huge gap between raw php and the other frameworks.
> For each request, an object mapping the key message to Hello, World! must be instantiated.
Why mandate how the benchmark is implemented? What if some frameworks could be faster by not instantiating an object? For maximum speed, you'd want to stream-process the input into output while materializing as few intermediate objects as possible.
Benchmarks should be defined in terms of their observable inputs/outputs. Otherwise it's like defining a car's 0-60 measurement with a rule that says "the car must then take gas from its tank and inject it into the engine." And then a Tesla could only compete by adding a gas engine that it doesn't need or want.
I'm not saying it's an especially useful benchmark, but it's intended to test something besides raw response speed.
> I'm not saying it's an especially useful benchmark, but it's intended to test something besides raw response speed.
It's such a small, silly optimization. It doesn't bother me because it's not going to affect the overall results much, but I do feel it's against the spirit of the test. That kind of thing would be prevented if we made the test echo a request parameter, as was suggested elsewhere in the comments.
These tests are intended to be stand-ins for realistic application behaviors. They are not realistic applications, but they should make use of the same functionality that a real application would. I want the tests to establish realistic high-water marks for the types of operations they exercise.
Running a single query per request? The high-water mark you might expect is what you see as our single-query test. It would be very difficult to make a single-query operation in a real application more trivial than what we have.
If I treated all tests as a black box, none of the database tests would actually query the database; the fortunes test would not add an ephemeral object and re-sort on every request; and so on. Everything would just send static responses back.
Bottom line: nothing beats benchmarking your own application. But the tests in this project are intended to give a first-pass filter of sorts, assuming performance is on your list of requirements.
A better way to avoid this would be to make the output depend on the input. Require that the reply include some JSON that contains portions of the request.
Likewise for databases, require that the reply contain contents of a previous request. If you want to, make the request contain an ID and make the reply return the data keyed under that ID.
If you want to require that the data is persisted durably, require that the benchmark continue to function even if all processes are killed and restarted in the middle of a run.
What really matters about a system are its inputs and outputs. If an innovative system comes along that works differently inside, it shouldn't be disqualified from a benchmark just for being innovative.
In any case, the way I read it, the requirement doesn't mean you have to make anything particularly heavy, just that you need to use a representation that's general enough to be usable in other scenarios without rewriting the entire app.
Of course, a little bit of extra encouragement in the form of a benchmark that's slightly harder to game wouldn't be bad.
For example, I wrote a little in-memory DB for use in a .NET framework to avoid having to deserialize an object graph per request. Instead, it navigated a byte array and plucked out just the data needed.
Another time, I wrote a compact in-memory representation of a trie, using bit-packing and a few other tricks, to get an order of magnitude reduction in memory usage, making it possible to cache a lot more of a data set.
Why go to the database and instantiate a classic object if you can avoid it?
Indeed. But if you want to actually force the server to encode some JSON, then make the reply include portions of the request. Then you don't have to have artificial rules about how it is implemented to get interesting/meaningful results: https://news.ycombinator.com/item?id=6650499
If you wanted you could also make the request specify the key at which a value must be written.
Be creative. :) But don't think that apps always have to be written the way they are now to be "legitimate."
Otherwise you won't be able to compare the results.
For example, the Benchmarks Game [1] specifies how the algorithms in those benchmarks should work. For some of the benchmarks, there are faster algorithms. However, it's not about comparing algorithms, it's about comparing language implementations.
[1] https://github.com/TechEmpower/FrameworkBenchmarks/issues/10...
We use uWSGI + bottle + gevent for REST APIs and a relatively humble box can reach flood a gig Ethernet link with _useful_ replies. And, of course, any decent implementation will cache computed results, etc. (but toss Varnish in - which we haven't, yet, since there's no need - and you'll outperform anything else for cacheable replies).
I fully realise the language itself is annoying and clunky and that Scala is not necessarily the answer. I'm just saying that there are some serious upsides to consider beyond language aesthetics.
In a lot of cases, performance of the server isn't the determining variable in the overall cost calculation.
In addition, the relationships between those variables likely changes over time.
I'm guessing that for most startups, HR expenses and opportunity costs are pretty important relative to actual server performance.
Obviously, if the point of your startup is some kind of high-performance computing thing, the equation is tipped more towards server performance. But, you probably aren't using off-the-shelf frameworks in that case, either.
Also, my point was that the newer jvm frameworks aren't much more complex or time-consuming than Rails so developer cost isn't such an issue. That said, not much compares to Rails for getting stuff up and running quickly.
Maybe these options will become really compelling once Kotlin is released. It compiles to both the jvm and JS so it will hit the performance, convenience and "fun to write" sweet spots all at once...or so I hope.
# WebSocket echo service
websocket '/echo' => sub {
my $self = shift;
$self->on(message => sub {
my ($self, $msg) = @_;
$self->send("echo: $msg");
});
};My only complaint is that I started with Mojolicious::Lite way back in the beginning of the project, and I really needed to port to Mojolicious proper with a better source file layout quite a while back, but Mojolicious makes it too easy to throw in another lite-style route. At least the templates were easy enough to separate that I did it long ago.
See https://github.com/TechEmpower/FrameworkBenchmarks/issues/49...
This less forgiving stance with respect to glitches comes from our intent to publish a round of tests every month from this point forward (assuming we have the manpower to do so). In order to pull that off, we need to be less forgiving with problems that we don't have the know-how to fix immediately. We investigated Rebar for a bit but ultimately conceded defeat, skipped the tests, and posted the issue linked above. We'll revisit it again for Round 8.
wsgi:
import ujson
...
response = {"message": "Hello, World!"}
...
tornado: obj = dict(message="Hello, World!")
self.write(obj)
What you have is the slower way to contruct a dictionary and then passing to the Python native JSON vs. a C optimized JSON for output. %timeit d={"message": "Hello, world"}
10000000 loops, best of 3: 103 ns per loop %timeit d=dict(message="Hello, world")
1000000 loops, best of 3: 298 ns per loop sys.version
Out[24]: '2.7.3 (default, Jan 2 2013, 13:56:14) \n[GCC 4.7.2]'
So, pretty big difference.At 103ns per iteration, you can do 9.7 million iterations per second. At 298ns per iteration, you can do 3.4 million iterations per second.
Let's look at the wsgi numbers for json serialization:
wsgi-nginx-uWSGI: 109,882 rps
bottle-nginx-uWSGI: 65,793 rps
wsgi: 65,755
..
Now, let's look at the best case comparison: 109,882 rps vs 3.4 million iterations per second. Is cutting iteration time down to 1/3 significant (298ns vs 103ns)? Yes. Is it significant in the overall context? No.With a 31x difference between 109,882 and 3.4 million* (and, since this is comparing to an i7, that 31x is probably closer to the ballpark of 100x), this simply isn't likely to be the place where optimization will help much. Put simply, the {"message": "Hello, world"} vs dict(message="Hello, world") cost difference is likely insignificant when compared to the cost of the rest of the request.
With that said, this sounds like a change that should be committed as part of a pull request! Why? A couple of reasons:
- It may still help the performance in a minor way (at the very least, it shouldn't hurt it)
- We prefer idiomatic code and, assuming the literal {} notation - vs dict() - is idiomatic for such a case, this would be a good change.
Your point about this is also interesting:
What you have is the slower way to contruct a dictionary
and then passing to the Python native JSON vs. a C
optimized JSON for output.
Please do consider a pull request for both. We love community contributions.* Yes, I know these aren't directly comparable numbers, but they compare well enough in this case.
Edit: Fixed the literal {} vs dict() mixup that e12e pointed out in my comment.
Good points on the overall (likely) impact on the benchmark(s).
> We prefer idiomatic code and, assuming dict() is idiomatic for such a case, this would be a good change.
I think you mix up two things here: dict() is slower (presumably method look up and maybe class instantiation? Just guessing here) -- and I'd say using the literal notation is in general more idiomatic:
http://docs.python.org/2/tutorial/datastructures.html#dictio...
And yes, I noticed that you weren't the one that noticed the original possible performance issue. What I do appreciate is that you actually went and tested the performance difference.
Maybe suggest some updates to the people doing the benchmarking. Or, god forbid, do the testing yourself and post some results for people.
You wanted to know why they didn't test whatever you thought was the optimal setup for your use case, and the answer should've been obvious - there's dozens of frameworks across multiple languages and OS platforms. If you look around on this thread and those for previous iterations of the benchmarks, a consistent response from the benchmarkers is that they didn't test some specific scenario because they lack the expertise to do so.
The tests are on github. Go write one that suits you and send a pull request, or, god forbid, do it yourself and make a post about it on HN and get yourself some karma. But as it stands, your post was lazy and trivial. I've said similarly dumb things in the past, gotten downvoted or called out, and we all moved on.
I'm sorry your fee-fees got hurt because you said something dumb, and I pointed it out in terms that failed to indicate respect for you. I upvoted your response as a gesture of love and understanding.
Ideas I had were: - Generating a nested JSON structure weighing in at 30kb (from a database) - Other more-real-word scenarios which I haven't thought up yet ;)
[1] https://github.com/TechEmpower/FrameworkBenchmarks/issues/13...
Instead, Pedestal's new interceptors abstraction decouples http requests from threads. This provides better concurrency support because it enables processing a single request across multiple threads (http://pedestal.io/documentation/service-interceptors/).
Incidentally, the Pedestal team gave their fans a free idea that I really liked in August [2]. Do it for fun and glory.
[1] https://github.com/TechEmpower/FrameworkBenchmarks/issues/58...
[2] https://twitter.com/pedestal_team/status/364945233814884352
This is the first time I've seen this benchmark, and it was really really interesting to see how Djangos performance degraded as more queries were executed. They mention that the lack of connection pooling is likely a big factor. I never realised how much that could affect an application.
These results are really interesting!
JRebel's web framework comparison this summer had GWT as the front runner, pitted against Spring MVC, Play, Grails, etc.:
http://zeroturnaround.com/rebellabs/the-curious-coders-java-...
Glaring omission for a comprehensive comparison, IMHO.
If you or anyone is willing to give it a try, however, we would gladly accept a pull request.
Surprised it came in so high (6), no wonder Spray was acquired by Typesafe...and why Spray will take place of Netty in Play stack.
We'd be happy to receive a PR to update these versions if anyone is interested in crafting one.
The errors column is a sum of non-200 HTTP responses and socket connect, read, and write errors. In practice, it tends to be predominantly 500-series HTTP responses, meaning the web server acknowledged the request but was too busy to assign it to a worker process or thread.
This can cause the latency of some frameworks to appear artificially low because the 500 responses tend to come back very quickly once they start.
Incidentally, if you hover over the error values, you'll get a success rate percentage.
[1] https://github.com/TechEmpower/FrameworkBenchmarks/blob/mast...
I'm curious to see what happens to performance when v.12 is released.
Rails is disappointing. Slow to be expected for interpreted language framework but that is really sh slow.
Speed is not the only factor, and usually not even a factor that has to be accounted for, lets be reasonable, 99% of any web apps dont reach the state that they even need to scale to 10000s of user actions per second.
Id be interested to see how hhvm/php whould compare though.
To think that a benchmark like this tells the whole story of the "best" framework is wrong. Of course it is. It's also wrong to think that it tells the whole performance story. Of course it doesn't. But that does not make it useless.
Id be interested to see how hhvm/php whould compare though
Now you're just arguing yourself :)
To the rest who are building apps that need to scale, the benchmarks are definitely useful.
With Grails specifically, a very late PR arrived that unfortunately broke the test implementation for us, so we removed it from the Round 7 test. We aim to push new rounds on a monthly cycle, so I hope that anything that is missing in one round won't be missing very long.
Because people just aren't going to read the notes. They'll skip to the results.
Also, anyone out there have any good real world experience with other scala persistence frameworks? I was surprised to see play-anorm do so "relatively" poor compared to other scala frameworks. Though I'm quite, admittedly, naive in how "big" each of the other frameworks are (only used play and scalatra).
Please forgive any misspellings or grammar errors. This was typed on my iphone.
I'd really like to see the chart showing RAM usage for these tests.
Time and time again though, these benchmarks show PHP outperforming node.js. I'm sure node.js will have lower RAM usage than PHP but as I'm more interested in reqs/s, I'm going to reconsider using node.
In my previous project I've actually used straight servlets and found them quite RESTful and elegant - HTTP GET calls a function, you run a SQL query, render results using a template. What else does a framework really need? I much prefer that to the beast that is Spring. In fact, that's pretty much what people use tornado+express for these days - map a URL to a function.
Our test cases are explicitly concerned with exercising performance when reverse proxying is not suitable, for whatever reason. If reverse proxying works for your use-case, definitely consider using it.
We classify object-relational mappers as follows: Full, meaning an ORM that provides wide functionality, possibly including a query language. Micro, meaning a less comprehensive abstraction of the relational model. Raw, meaning no ORM is used at all; the platform's raw database connectivity is used.
With Rails I might do:
1. Fresh rails app
2. Add page caching / CDN
3. Add database/query caching
(I have no apps large enough to require anything past here)
4. Sharding the DB?
Rails is awesome because each step is concise (via the language, ruby) and many other people have demos / documentation on how to do them.
Sure, your framework of choice may be fast at stage 1, but how easy is it to go to 2, 3 and so on when I need to?
Yep, that's the problem of these benchmarks. Apples and oranges.
I have answered this sort of question many times. I eventually wrote a blog entry about it: http://tiamat.tsotech.com/unfair-comparisons
I wonder of they setup PHP with APC cache (opcode cache) which is basic and easy setup to speed up PHP.
"Have you enabled APC for the PHP tests?" Yes, the PHP tests run with APC and PHP-FPM on nginx.
But when we're busy with other projects, this becomes a real bottleneck. We wanted to get Round 7 out and be able to do future rounds on a monthly basis. In order to pull that off, we simply need to reduce our own bottleneck and ask the community to do more of that work. Round 7 was the first with a preview round and the community submitted a bunch of PRs prior to the final run.
To be as clear as possible, I absolutely do not blame the missing Go data on EC2 on the community. It's just the nature of attempting to herd all of these cats, er frameworks. :)
I'm fairly sure it's a configuration problem either in the Go test implementation -or- our test-suite. And this part may be obvious, but it bears repeating: except in very rare cases, test absence should not reflect poorly on the framework itself.
http://www.techempower.com/benchmarks/#section=data-r7&hw=i7...
- Laravel Version 3.2.14
- PHP Version 5.4.13 with FPM and APC
Don't waste your time with this crap.
Wouldn't it just be better to compare caches/caching mechanisms?
They released it as part of the language in 5.5 because alternative APC cache got just good enough and everyone was installing it by default.
So, to answer your question: yes, APC should be used on any PHP benchmark if performance is in question.
I think we used to have a brief description of each test in the results section as well, and they somehow got lost in this round. We should add them back; how is a new reader supposed to know what "Fortunes" means?