Silly benchmarks on completely untuned server code
rachelbythebay.com
rachelbythebay.com
0. Performance isn't everything, especially to a lot of companies where having something at all is more important than just being the fastest
1. Dynamic languages are generally require less LOC, which generally equates to faster implementations
2. Rails, Django, <your favorite dynamic language framework> sure are generally bloated, but being able to drop in a library for nearly anything you need cannot be overlooked. Especially for small-midsize companies, spinning up another app server is generally easier than writing a whole bunch of multi-threaded code.
3. The appserver is generally not the bottleneck, rather, of course, its the database.
That's not what I observed. My understanding is dynamic languages promote people to think more higher level and care less about performance and program structure. While dynamic constructs can be efficient, the freedom they give is often misused by developers to create baroque structures that hurt performance mainly because developers don't know how to use them efficiently.
apr_socket_recv: Operation timed out (60)
Total of 49189 requests completed
Same thing when omit `-c 100`. Am I doing something wrong?Anyone can cook up a plain http server, making a real application out of it is the other 98% of the work.
package main
import (
"fmt"
"net/http"
)
func main() {
http.HandleFunc("/", func (w http.ResponseWriter, r *http.Request) {
fmt.Fprintf(w, "Foo")
})
http.ListenAndServe(":8088", nil)
}
and when I run it with `ab -c 100 -n 100000`, it falls over before 10k requests: $ ab -c 100 -n 100000 http://127.0.0.1:8088/
This is ApacheBench, Version 2.3 <$Revision: 1843412 $>
Copyright 1996 Adam Twiss, Zeus Technology Ltd, http://www.zeustech.net/
Licensed to The Apache Software Foundation, http://www.apache.org/
Benchmarking 127.0.0.1 (be patient)
apr_socket_recv: Operation timed out (60)
Total of 6388 requests completed
I'm extremely new to go so maybe I'm doing something wrong. Could someone help me understand?https://www.techempower.com/benchmarks/#section=data-r18&hw=...
As you can see, the fastest Python implementation is just 11% of the top solution in Rust.
And your key/value store choice matters a lot. Redis is well known, but it's much slower than Tarantool or LMDB. Which is slower than RocksDB or Aerospike (though take this claim with a grain of salt, they can perform differently under different loads (writes, reads, updates, deletes) and different number of concurrent requests).
Keeping the use case in mind (and it's possible evolution) will help pick the best tool for the job.
I'm unfamiliar with Tarantool, but comparing a "server" database like redis with embedded ones like lmdb or rocksdb is so strange. It's comparing different engines to other whole cars.
NB: For anyone not in the know, "server" vs embedded here is about how your processes communicate with the DB, not about the type of hardware they're suitable to run on. An embedded database tends to be more of a library that attaches to a file or directory, often from a single process or occasionally from multiple processes on a single host.
"Server" databases tend to be connected to from processes on a number of servers over the network.
I keep using "server" in quotes here because I don't think I've seen that word used to make the distinction vs embedded databases reliable in literature.
black sheep (python): 101,508 req/sec
actix (rust): 886,499 req/sec
yes, 101508/886499 = 0.11
but your "x is % of y" doesn't seem to capture the relationship very well here IMO.
you could also say, the rust library processes 773% more requests per sec.
or, the python library processes 89% fewer requests per sec.
(using: percent change = d2/d1-1)
personally I prefer multiples in terms of the larger thing. actix does over 8x the requests per/sec as the python lib.
Can someone simplify the argument that she's making about them? Something to do with maintaining separate, blocking threads in anticipation of requests vs spinning things up as you need them? Or is it a criticism of dynamic languages as HTTP servers in general (it's not clear what language was used in this article)?
If you tune and use a performance oriented language (on a big multicore machine), you can probably get about 100x better than the untuned not-python throughputs. So, for http hello world, 1000 untuned python servers ~= one beefy server.
Is this some C code written by the author? Is it Apache? Something else? Is the point just how much faster native code is than Python?
Apparently someone else also built an untuned setup from scratch, and got better numbers.
I think the point of the article is that if the results are surprising to you, then you should try writing one too.
2. Python has several different ways to do concurrency (because they evolved over ~25 years) that shouldn't be mixed but it's very easy to accidentally mix them.
3. Binary RPC (e.g. gRPC or Thrift) is more efficient than REST.
- The Global Interpreter lock - you can't avoid forking processes in order to handle concurrent python code.
- Threads are cheaper than processes. A default java thread carries 2MB overhead, a python process for a typical app can easily be 2GB without very careful memory consideration.
- The pure single-threaded python runtime is 10-100x slower than your typical statically typed language, even when doing everything as carefully as possible you'll be 1-2 orders of magnitude off the best implementation in a statically typed language. Conversely a sloppy implementation in a statically typed language will probably work about as well as the best python implementation.
- Foreign Function Interfaces(FFIs) are slow, in Python an int32 consumes 24 bytes of memory whereas in C the int32 consumes 4 bytes. Stringing together 2 C calls that manipulate an int with python would require 4 allocations and 4 casts. Applications that avoid this overhead have to adopt symbolic APIs that inform an underlying C program how to connect multiple function calls and will grind to a halt if there is any python control flow e.g. TensorFlow.
Most of these limitations tend to be common to other dynamically typed languages such as Ruby, PHP, and Perl. While one can theoretically "drop down to C", it's often not that straightforward in common development scenarios. For businesses that require high(er) throughput, or low(er) latency the time spent fighting the interpreter or the associated AWS bills may not be worth the productivity gains from dynamic types in 2020 when compilers can perform an incremental build in milliseconds.
I think (or at least assumed) that the post's discussion is more interesting than that. A sibling comment seemed to be zeroing in on the crux of the issue. "Python is an order of magnitude slower but it's web server is two orders of magnitude slower" is a meaningful and fruitful thing to talk about.
It may be worth looking at how this performance gap has trended over time. Anecdotally I recall the performance gap being much narrower 10 years ago.
Sorry, this is not true. I have a Python app that uses 2 threads. It's memory doesnt even exceed 200 MB.
Yes, Python threads generally suck because of the GIL, but they dont cause memory bloat like what you're describing.
For context, I use Python extensively for extract/transform/load (ETL) work. I deal with quite large files. My loaders all run in mear linear time relative to the number of records in the file, and none use more than 300MB RAM. This is against Python 3.7.3.
If your multithreaded Python app is using 2 GB of RAM, it's not because of the threads. Best look elsewhere. Maybe you're caching something large in thread local storage?
> - The Global Interpreter lock - you can't avoid forking processes in order to handle concurrent python code.
> - Threads are cheaper than processes. A default java thread carries 2MB overhead, a python process for a typical app can easily be 2GB without very careful memory consideration.
And it is indeed true that forking can be much more memory intensive than threading.
Threads are not cheaper than processes in Linux in any significant way. See the stack overflow question on that: https://stackoverflow.com/questions/807506/threads-vs-proces...
A _langauges_ use of those may have poor performance.
class TestApp
def self.call(env)
[200, [], ['Blah']]
end
end
run TestApp
ab -n 10000 http://localhost:8080/ --> ~1900 RPS (~8400 RPS)ab -k -n 10000 http://localhost:8080/ --> ~1900 RPS (~26000 RPS)
ab -c 100 -n 10000 http://localhost:8080/ --> ~2500 RPS (~14500 RPS)
I have no idea why it performed better on the concurrent benchmark. This was on my 2.3GHz i5 mac book pro. I think about 1/5th the performance of something closer to the machine is quite decent.
More likely I'd fork a process in a safer language and direct it to make HTTP calls.
I've also seen interesting things from Abseil and Folly.
[0] https://isocpp.github.io/CppCoreGuidelines/CppCoreGuidelines [1]https://www.dre.vanderbilt.edu/~schmidt/ACE.html
I'm guessing you don't count cURL as half-decent. Because it's API is C++k let alone idiomatic, modern C++?
Basically anything with an event loop and good non blocking IO support for client requests and you will get loads of traffic through without too much hassle.
Of course a lot of this really depends on what your app needs to do over and above handle HTTP requests and what sort of library support you get from language / framework X.
Java, for example, has some great reactive libraries, but any JDBC driver blocks (as it has to, enforced by the JDBC spec). There's a couple of non JDBC drivers that support non blocking IO, but they are still evolving.
It's a similar story with other languages, there's always something necessary that isn't quite there yet.
That said you can still squeeze vast performance out of the blocking IO things with a solid engineering base (Tomcat, IIS if you can stomach windows) you just hit the thread limits earlier and harder than with the non blocking stuff.
Close rivals might be PHP and Node.JS, there's pros and cons to all of them.
If something needs to be very performant you pluck it out into a little golang microservice.
Node.js is much faster because V8 is much faster, but it's still basically single threaded, so you need to run process-per-CPU, which is what the original blog post (to which this post is a follow-up) was complaining about.
Node.JS is much faster both because it has a fast VM, and because it has concurrency much better dealt with in its standard library than Ruby does (i.e. a proper event loop is built in and its use is idiomatic).
That's not what the original blog post was about though, it was about how bad Gunicorn is. I've written Python, but never web apps (since Ruby's superiority there is obvious), but I don't think it's really fair to judge the whole language based on the use of some popular library that's shitty. I think it's symptomatic of Python that the library is shitty, but that's my dislike of Python shining through. The reality is that there are most likely great Python libraries for dealing with concurrent web requests, and she didn't bother to learn them.
On Ruby the community would without hesitation recommend you to run Puma or Passenger or even Iodine or Falcon or whatever fancy stuff you have nowadays. These webservers all deal with concurrency, optimizing memory usage, load balancing and managing queues correctly. They don't have silly things like fork'ing before importing libraries, because good software architecture is something that's highly valued in the Ruby community.
Yes, the post starts out describing an issue with how Gunicorn listens for connections. Like you said, there are better libraries than Gunicorn so that's not a reason to jump to Ruby.
However, the article goes on to talk about other things, like the lack of real multithreading, import-time code execution, and overall efficiency, all of which also apply to Ruby.
I'm not saying Python or Ruby are a bad choice. It's just that MapleWalnut was soliciting an alternative that doesn't have the problems described in the original post, so Ruby doesn't qualify.
Not sure if my machine is at all comparable (i5-8265U unplugged laptop CPU @ 1.60GHz) but also the phoenix hello world does a hell of a lot more, like setting session cookies and producing a full front page that renders through two templates. They also took a bit longer to get to in developer time. While you go get a website at :4000 with two commands -- five if you have to install elixir from scratch, it's not really representative of performance: I had to go to prod to disable the live code-reloader and debug/info level logs, disable SSL (releasing to prod defaults to SSL-enabled) and perform a release (elixir these days really wants you to have devops hygeine).
1) I was pretty pleased with the performance. (1900 rps in the base case, 2600 rps in the keepalive case)
2) Unlike Rachel's platform, Elixir is better with concurrency out of the gate (went up to 7800 rps with -c 100), which tells me, the erlang VM is really doing something right.
3) Paying for the cost of having correctly implemented, difficult things like sessions and no XSS would be worth a massive performance hit, IMO.
4) Doing the right thing with Elixir is crazy easy, there's way fewer footguns than python, and code is typically extremely well documented and tested, and inspiring enough to make you want to document and test (though there are fewer libraries).