TechEmpower Framework Benchmarks: Round 20
techempower.com
techempower.com
These days when I'm evaluating a framework, I look for strong documentation, integration with tools for observability and deployment, and compatibility with a well understood runtime that makes operations easier.
These things are way more important for building software that works, that scales, and is operationally efficient for a team of engineers to work with for several years.
Everything else is just vanity metrics.
I suspect the examples would be few at the top of the list and many near the bottom.
I don't care much about performance either, especially compared to documentation, but that does not mean these benchmarks are useless. For instance, comparing similar frameworks can show big differences which illustrate different software architectures. Among the Nginx/phpfpm/Mysql fullstack-with-ORM frameworks, Laravel has roughly 8% of the raw stack capacity, while the other mature frameworks can do 2× or 3× more. I suspect this reflects Laravel's abuse of magic (calls to non-existent methods are redirected at runtime to other objects created on-the-fly, and so on). This may not be the real cause, but such a poor performance makes me suspicious about the complexity or the quality.
Computers are stupidly cheap slave workers and are easy to spin up compared to dev time. Does it actually matter in almost any profitable business?
The speed argument is a much better argument imo, but we have to think if 100-300ms is worth the tradeoff of a great framework with a strong community. (I've not seen anything come near the beauty of Rails in Java, SpringBoot is not it. Last I checked there was no good SideKiq alternatives which feels kinda vital for most apps)
If a Rails app can handle only 300rqs that's over 1 million requests served an hour. A SpringBoot app handling say 2000rqs? Thats 7.2million a hour definitely a shit ton more but you know what, I'll probably just set the EC2 autoscale to 7 more instances and pay the extra $1000 a month because its still a magnitude of 10x cheaper than developer time. I can see your point in keeping server count low for simplicity being an advantage but not enough to persuade.
I say all this but I do like the speed argument, I think its important that we keep web requests fast for a better user experience. I mean Stripe, Shopify, GitHub all get by on the slowness that is Ruby but I will say that lately GitHub has been super slow for each page request but GitLab has been very fast which is also Rails.
There is another view point if you are operating on a slim margin business that keeping infrastructure costs tiny is important and that's another case for a faster backend for sure.
Now the interpreted lang vs statically type lang argument I can agree with, I hope for a language that sits somewhere between Java and Ruby. It seems Crystal fits that bill at the moment but has no backer and will be years off a great ecosystem if it takes off.
The complexities mentioned generally fall on the developers and sysadmins who have to be aware of them and work around them, those people are definitely not cheap slave workers. Screw ups handling these complexities can bring the whole system down.
Which is why Rails is only suited for SaaS where you get paying user early in the development. If you require traffics for ads revenue it simply doesn't scale. And majority of the Web is still based on ad revenue model.
Also, for a SaaS, keeping the cost per served transaction low can mean the difference between making a profit or a loss in a market that has a converged on a price-point for a type of service.
Honestly? Sane people.
I worked at 80 person company which broke and closed because of this. We almost escaped this demise by going from PHP to HHVM which unfortunately was very alpha at the time, but we simply ran short of money and time. If it had worked, we would have scaled down 15x our AWS costs and would have saved at least 40 jobs.
I wouldn't use TechEmpower as a reliable source for checking framework performance.
I’d love to see stupid simple benchmarks against each framework’s build-a-blog tutorial.
Seeing Rust climb, PHP/Ruby/Python/etc fail. Also what Node.js is capable of. It's fun!
And it's also nice to point at TFB when arguing with your architects that 100 rps is not "a lot".
I thought they were just hand optimizing their code to the point they submissions were totally unidiomatic. But now I learn even dependencies get this kind of attention.
Years ago I submitted a ticket asking for inclusion of memory usage stats (interesting in small environments) and start-up time stats (interesting in serverless environments). But nothing ever happened on that front; I guess it's because TechEmpower is a Java shop and that are not typical metrics where the JVM shines.
There is some value but it’s not a lot. At the end of the day it all comes down to how much time and motivation the person that implements the benchmark has, you can get pretty far in just about any language, after all they can all call into a native implementation so even an interpreted language can be fast. Then you can spend lots of time micro-optimising for the specific hardware the benchmarks run on.
An example scenario could be if a company has a choice between Java (Spring) and .NET (ASP.NET Core). Both are quite popular, have documentation, a big ecosystem around them, etc as you mentioned. If you might need to do low level/advanced web code at any stage the plaintext benchmarks may sway them to use .NET over Java; all else being equal. In that benchmark ASP.NET Core trumps Spring by a very large factor (Spring's 2% vs AspCore's 100%); you then obviously look into the benchmark to compare and contrast as you shouldn't trust them at face value. Although having said that if the .NET one is better tailored than the Spring one it does send a signal that the .NET one has a more performance orientated community willing to maintain the benchmark which IMO is a positive still.
We upgraded an app from Spring 4 to Spring Boot a while back and it divided our hosting cost by 5.
I started caring about speed when I worked on an app that became much bigger than anticipated. If you start with a really slow framework your only option is a rewrite
I read once on HN that a lot of modern day deployment strategies and tooling seems like it was created because of very slow apps. When you're running a rails app, you need to scale horizontally, and so you need to orchestrate that, and so you have autoscaling instances, kubernetes, and so on.
But if a single, beefy instance can handle all the traffic your app can reasonably expect to see (think, StackOverflow on just a couple instances), then a lot of operational complexity just dissolves away.
Managing 10 servers is perfectly ok, having to handle 100 servers would take a lot more work. A 1000 servers would require organizations to hire a lot more people and work a lot harder on automating server management, deploys, etc.
I agree so far as that just using some standard language like C# or Java is probably good enough even though there might be another language and framework that is faster. I will personally stay away from php, python and ruby for larger systems.
Great, let's see those metrics tested and compared in an easy to read table, then we'll have a matrix of qualities on which to base a choice of stack.
I’m a .NET dev primarily. And .net has historically been a terrible performer on these benchmarks. With .net core the asp.net team and community has been working on performance.
While the raw plain text benchmark is not really realistic of a real world scenario that we would use on a day to day basis. It is an indicator of the performance baseline that .net core can achieve before you begin adding all the fluff on top.
This coupled with sites like stack overflow showing the gains they get using asp.net and moving from 1 version to the next gives me confidence that .net is not a bad choice these days.
Prior to .net core I was more or less ready to go switch full time to something else. But now I’m happy using .net.
/my 2c
I also got a bit suprised over this other benchmark showing how fast C# can be: https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
I think what really hits home are the stories about people upgrading elements at the foundation of their stack—their platform or framework—after those foundations have seen optimization efforts. It is rewarding to see application developers realize dramatic performance boosts within their applications with so little effort; we feel we have contributed to this delight in a small way. For example, check out this post by the Azure Active Directory team [1] where they saw massive performance improvements by upgrading from .NET Framework to .NET Core 3.1 (which isn't even the latest version).
Some people focus too much on specific rank ordering and the optimization efforts made by those who jockey for top spots. It's better to consume the data in rough tiers of performance, however you choose to define those tiers. We tend to encourage people to consider performance as one part of the puzzle when selecting infrastructure software. Several high performance platforms and frameworks (in Java, Rust, C#, Go, Python, and more) are also ergonomic and easy to work with. And using high performance infrastructure software means you can avoid premature optimization in your application code and avoid premature architectual complexity. You can enjoy the trifecta of architectural simplicity, low-latency, and good scale headroom.
Obviously we know this project will always receive diverse and critical opinions. That's fine; hackers are a very opinionated people. For those who value the data, ourselves included, we are happy to keep putting the effort in.
[1] https://devblogs.microsoft.com/dotnet/azure-active-directory...
- Even if it's JS, just-js make almost no use of dynamic memory allocations.
- The authors rewrote himself a postgresql driver that support batching request.
- it wraps high performance c++ libraries
Even if the techempower implementation is far more verbose and complex than mainstream JS framework, the amount of work behind just-js and its preformances are just impressive!example:
"-d:danger : Turns off all runtime checks and turns on the optimizer."
https://github.com/TechEmpower/FrameworkBenchmarks/blob/mast...
That accurately represents how I've been forced to run applications in production too, when under huge stress and we're trying to squeeze everything out of the existing hardware.
Welcome to _Hacker_ News ;)
Edit: And absolutely yes, check out the code behind it. Every benchmark has their flaws and biases so it's important to verify before making conclusions.
https://github.com/rryqszq4/ngx_php7
https://github.com/TechEmpower/FrameworkBenchmarks/blob/mast...
Very similar to OpenResty, but with PHP instead of Lua. Which is funny, because OpenResty ranks lower.
And then Click on Filter, Full Stack Framework, And Full ORM only. Again, most people will be using some sort of Web Framework with Full ORM.
I am surprised the top results is from PHP. For many years it has always been Java. And it is from a framework called ubiquity [1] which I have never heard of it before.
There is Crystal and Lucky [2]. Which is exciting because they are getting very close to 1.0
Ruby Rails... is still at the bottom of the pack.
PHP
Impressed by PHP's performance - it makes an entrance at position 20. If you exclude Javascript, the next interpreted language is Lua at position 81, then Python all the way down at 206. Ruby makes its first appearance at 255.
Julia and Nim
These are both new, modern languages that tout their performance as a benefit. It's a shame they did not take part in the benchmarks.
***
Regardless of what you think of the benchmarks, the rankings do affect people's perceptions of languages and frameworks (both positive and negative). For example, I don't use PHP, but these benchmarks tell me that PHP has leapfrogged over Python and Ruby in the performance stakes for web development. And not just by a small margin, but by a significant difference.
In many languages, a bunch of web server code might end up being implemented in C and not the language itself. When you're creating a minimal endpoint for a benchmark, you might be exercising that C code and not the language itself. For example, the PHP implementation uses `htmlspecialchars` to encode the information which is implemented in C. That doesn't tell you much about the performance of PHP, but just that it has an optimized HTML escaper. The `asort` function used to sort the results is implemented in C and has more limited functionality compared to more general sorting functions that might be able to take lambdas in other languages. The PHP implementation even takes advantage of `PDO::FETCH_KEY_PAIR` which will only work if there are only two columns.
Likewise, the PHP implementation doesn't use templates, but manually builds strings. The fastest Python implementation actually renders a Jinja2 template: https://github.com/TechEmpower/FrameworkBenchmarks/blob/mast.... That's much more realistic in terms of what you'd do in the real world, but you're going to be carrying the overhead of a real-world template system like Jinja2. Part of Python's failure here is that no one wanted to implement an optimized Python version that would just build a string instead of rendering a template.
Changing the test constraints a bit would ruin a lot of the advantages that PHP used there. Let's say that you had to retrieve three columns: `id, sort_key, fortune_text` and sort on the `sort_key`. Now you need to read more information back rather than just being able to make it an associative array (hash map). You need to be able to sort based on that sort_key which means probably giving a sort call a lambda.
This isn't limited to PHP. A bunch of Go implementations do things like allocating a pool of structs and then re-using the same structs to avoid the garbage collector. A lot of implementations create result arrays sized so that they won't need to be re-sized (creating additional allocations and additional GC work). The rules say this isn't allowed, but they do it anyway.
So, before comparing tests, I'd look at the implementations to make sure that they're comparable. A Django implementation that actually returns objects and is rendering templates and looks like a canonical Django implementation is very different from an optimized PHP version trying to avoid running any PHP as much as possible. When we start looking at the popular PHP frameworks which will be executing a bit of PHP like CodeIgniter or Laravel, we start seeing performance similar to Python frameworks as the PHP code is doing similar things like rendering templates. It just happens that no one implemented a Python version that didn't use a fully-fledged template renderer.
And this is the weird thing: the benchmarks changed your perception of PHP while comparing things that weren't similar. I think PHP is often faster than Python and Ruby, but probably not to the extent that your perception might be given these benchmarks.
I actually find it fascinating to look at the implementations and see which communities care about realistic implementations vs. leveraging all sorts of tricks to win the benchmark.
Of the bunch, by only looking at the code, which of them would you like to get dumped on you 2 weeks before a deadline? Which is the most ergonomic?
Most of them check the Nth letter of the HTTP method to do routing, hard code the content-length, etc. This benchmark shows which language is the fastest to throw the hardcoded data into the socket.
Meantime, Django[1] is the simplest and a real-world application.
1: https://github.com/TechEmpower/FrameworkBenchmarks/blob/mast...
It’s kind of frustrating to spend a few hours investigating preview results of the latest techempower benchmarks just to come to the conclusion that (a) truly high performance requires tuning your code to linux, to the network interface cards, and a deep understanding of what your code is doing and (b) right now at the top of the benchmarks, common high performance Rust code is mostly “unsafe” memory-wise, and the same is true of C++ for obvious reasons. After that:
* When looking at garbage collected languages with very minimal implementations (Jooby mostly), Java can come out on top but requires 4 Gigs of memory which is more than I would ever like to use.
* After Java comes C# but performance slows down the second you want to do anything less optimized, like use third-party libraries.
* Finally, we’ve Go, which in a surprising twist is basically neck and neck with some pure PHP server implementations, but once I include those, I might as well mention there’s a JS implementation that skips the Node.js async event loop and offers performance on par with C++ and Rust, but heavily uses C++ internally and is thus more “unsafe”.
What this tells me is kind of what I expected, implementation matters more than language… at this point, memory usage aside for Java, an efficient HTTP/1.1 server can be implemented in basically any language and when tuned or stripped down tends to run faster than the commonly used web servers, like Nginx, Tomcat, Express, etc. Often this means writing plain text RAW HTTP to a socket and managing sockets efficiently.
Which brings us back to Go, though. Of all the implementations I’ve looked at so far, even the PHP one, only the Go benchmark is written like “standard Go” such that most people writing a service in Go will write something high performance without trying to make it high performance. Effectively, you don’t need to ensure all your dependencies use special optimized code routines to get something relatively optimized working quickly in Go with low memory overhead unlike Java.
I’m a bit shocked by this conclusion as I was really hoping that Rust’s high performance use cases would win out, as it’s true that Rust can get 3x faster than Go, but on the Rust side, both actix and ntex are too immature as neither’s hit 1.0 yet, while tokio has hit 1.0 but its server, warp, is slower than Go.
Irritating is that each of the techempower benchmarks use different implementations. For example, there’s a benchmark that’s supposed to measure 20 calls to Postgres, but one benchmark gets to the top by making only one long SQL statement that changes 20 rows. Another implementation uses pgx’s Batch functionality to send multiple queries in batches (a big timesaver, but not technically standard libpq), but then the standard Go variant doesn’t use Batching even when it could (which means we can’t compare custom implementations to generic Go ones fairly): https://github.com/TechEmpower/FrameworkBenchmarks/search?l=...
I only see a single "unsafe" in the Rust implementation[0]? And that's only in the "raw" Actix instance, not the pg. Or are you talking about Actix framework being mostly unsafe?
> only the Go benchmark is written like “standard Go” such that most people writing a service in Go will write
Maybe it's just me but the Actix apis seem pretty ordinary.
[0] https://github.com/TechEmpower/FrameworkBenchmarks/blob/mast...
Yep, that's what I meant. To me, until Rust unsafe implementations hit 1.0 as web frameworks, not just Tokio, I can't consider them as stable or bug-free as more mature web servers. It's arbitrary, yes, but it's also an important distinction. 1.0 generally means "you can rely on this to not change, and to have a useful set of functionality" and bug fixes equivalent to an LTS release. This is especially true if functionality is common enough that it makes its way into a standard library or conventional best practices. Generally, it has to hit "1.0" before it does so, though, in any stable manner...
Far as 1.0 in Rust, the predominant culture is to treat "0.X.x" releases roughly equivalent to a "1.x" semantic. I have run into very few issues of API breaking changes in "0.X.x" libraries, and even those are basically handed to you thanks to Rust's robust compile-time checks.
Of course your mileage may vary, but I've had a way worse time with even well-established breaks in Ruby, Python, and Go libraries. I understand your risk aversion; I am similar. But I've found Rust to be an absolute delight for stability. The shaky ground of the first few years has largely given way to impressive robustness over the last few.
I strongly urge you to consider Rust if it matches your needs. You may be quite pleased!
To me, that Tokio hit 1.0 is a significant milestone, one that encourages its use for new projects. I personally look forward to the first web server widely adopted that hits 1.0 and relies on an async implementation that’s also 1.0…
I agree that Rust is very attractive, but it’s hard for me to argue for Rust against Go’s stability if you can live with Go’s more limited syntax. If Rust could promise the same stability as you’d have writing a JSON API or HTML-templates website in Go, I’d start picking Rust for new projects.
My perspective, to be clear, is that the heisenbugs likely still exist in Rust servers but have been mostly ironed out in Go ones. I’m reminded of Hyrum’s Law - https://www.hyrumslaw.com/ - under the assumption that as libraries hit 1.0, their behaviour can be more reasonably relied upon than pre-1.0…
That seems like a clear cheating case, over-optimised to the circumstances of the benchmark.
With absolutely no snark intended: what alternative do you see (or imagine) to this? Applying quite a bit of mechanical sympathy to tune a program to the system and visa versa seems essential and inescapable.
Maybe so, but we've been running https://pernos.co on Actix for a couple of years now with no issues. We didn't have to use "unsafe" either, and I understand that Actix's use of "unsafe" internally has also been significantly reduced.
Tokio's "warp" is much less mature than Actix AFAICT. Version numbers can be misleading.
Part of it is that Rust as a language moves quickly, I agree, but there's right now a lot of confusion between picking the pre-1.0 Actix or the post-1.0 Tokio, with AWS seemingly endorsing Tokio but with practical web development restricted to Actix Web for now. This confusion highlights the risky nature of pre-1.0 web servers -- it's not that you can't be productive and stable, but you might need to keep living on the edge for a long time. If you have a small service, that's fine. If you're looking for an API that won't change for 3 years or more, it's harder to suggest Rust at this time...
Do the ones at the top (like Actix Pg for example) provide everything you need to do real development, or are they stripped down? In other words, is this comparing track bikes to cross country bikes?
It’s worth looking at low level HTTP, socket and concurrency management if you want faster performance, but that’s not really a language-exclusive feature at that point. And the more realistic you make the benchmark — the more communication between microservices, for example - the more your application architecture, deployment hardware, kernel and network tuning, and so on can play a role.
I am reminded of http://rachelbythebay.com/w/2020/10/14/lag/ for example, as something really low level you probably don’t need to worry about… until you do. http://rachelbythebay.com/w/2020/05/07/serv/ also. The same is true of most of the performance optimization at the top. Really fast? Yes. Useful? Depends on the rest of your code. Are you writing haproxy? How much does your app really need to do? Read the benchmark source code and get inspired, maybe.
My conclusions for the moment: Switch to Go if low memory usage and high performance is as critical to you as post-1.0 stability. Go is generally stable these days ;-) Otherwise if instability can be tolerated but you want high performance, use one of these newer web server techniques and write a web server in unmanaged C++ that has very minimal functionality with language bindings. Just-js can serve as an example. Heck, the PHP benchmarks show that if you use PHP to write your own HTTP server you can still achieve high performance. That tells me the advantage here goes to web servers that literally “do less” rather than picking a language as faster over another language… especially when most (all?) languages can interface with C or C++ to do the HTTP layer at high performance while writing your code in whatever language you choose…
I wonder if the current actix-core is the rewritten/safer one once a bunch of Rusters wrestled/annoyed the original author into handing it over. If so, it's still near the top.
Techempower classifies each one. Actix is what they call a "platform"
>a platform may include a bare-bones HTTP server implementation with rudimentary request routing and virtually none of the higher-order functionality of frameworks such as form validation, input sanitization, templating, JSON serialization, and database connectivity
Even if you use the top ones in the list, you will have to build abstractions in code to handle software complexities. Those abstractions will reduce performance.
I do wonder, if we could use AI to dynamically remove those abstractions, to the extent possible and transform abstracted, multi-layer code to simple functions / compile targets, improving performance. Sort of like looking at the project, at this instant, and folding all database / processing code into higher performance code that reduces function calls or object instantiations.
Yes they offer everything we need for real development at enterprise scale.
Not to say they don’t have value — they do — but keep in mind your code may look very different from benchmark code. I do think it’s worth looking at the code to examine why some approaches scale better than others, though!
Rarely is it the case that more code = faster performance, but sometimes features you need are sacrificed for extra performance even when a benchmark is classified as “realistic”. Viewing source code is how you can determine which frameworks actually best fit your needs, rather than just which ones performed well in this test…
For example, it’s easy to think that raw Go performance is terrible compared to Rust or C++, but it’s worth pointing out that the raw Go example uses commonly available standard libraries baked in to the language to accomplish almost everything it does. Most other examples at some level or another have to deal with epoll and concurrency manually themselves or via the web server replacements they use. Based on this, the overhead or raw implementation of socket management is often what is being benchmarked. Implementations like just-js use C++ and epoll under the hood to implement their server, like most of the other high performance examples at the top of the benchmarks.
The question often then becomes, do you want to write your own HTTP server or take advantage of an API someone else has written that might have more features or middleware? And remember that configuration matters, I forget off the top of my head, but the Go examples with a new web server had more performance optimization than the raw-Go benchmark had. Obviously I should submit a PR to fix this, but my ultimate goal here is to say that reading the source code will both help you become a better developer but also catch the shortcuts and techniques used by high performance implementations…
The master branch of rocket does have async (and stable rust, not nightly!) support.
https://github.com/TechEmpower/FrameworkBenchmarks/blob/f340...