Scaling to around 15K requests per second with Java
medium.com
medium.com
Our Tomcat servers don't see 15k requests/s, but our load testing has shown a pair of them is good up to 3k/s (each) on dual-core 2gb servers with 512mb heaps. We really didn't push it much farther because that exceeds our needed capacity.
This is pretty boring technology. I'm fairly certain we could hit 15k req/s per node with plain Tomcat if we actually tried. I do think if you wanted to stretch past 15k req/s you probably do need more exotic programming models like Vert.X, but that does come at complexity.
I think the ultimate in throughput would be to have Tomcat listen on a unix domain socket (This is actually supported in Tomcat(https://github.com/apache/tomcat/commit/a616bf385a350175a33a...) and have TCP terminated with HaProxy (https://bz.apache.org/bugzilla/show_bug.cgi?id=57830). We set this up and it is _really really really_ fast, but we'd be on our own for support since nobody else is really doing it. Tomcat also has some strange behavior with the sock file, you kinda have to manage it on your own which is strange, but it could be done within the startup scripts.
It plays to the strengths of both becuase the reverse proxy handles long-lived but slow connections very well, and the Tomcat server gets to have fewer, but hotter threads.
In our limited tests with HaProxy -unixDomainSocket-> Tomcat, we saw a significant increase in throughput and decrease in latency (when it worked). It got rid of our wiretap point though which is pretty handy, and we didn't _really_ need the performance increase. If I were to speculate, HaProxy is a rocket and beats even the mature codebase of Tomcat it seems when parsing http, and probably takes advantage of kernel features that haven't made its way into Tomcat yet. It wasn't prod-worthy stable, sometimes execution threads seemed to hang.
In a real application, with database connection pooling and auth sessions, it went down to 15k requests/s.
And that was PHP7. PHP8 introduced JIT so it's probably significantly faster these days and hopefully fully typed.
Not that they shouldn't have used Java, I understand their application was already written in Java and it wouldn't make sense to use an unfamiliar language, but I do wonder about Javas' efficiency here.
Just to pick a random stack overflow post, as a little anecdata to go with anecdote https://stackoverflow.com/questions/44708704/netty-as-high-p... has someone who's complaining they were only hitting 50k POST/sec with Netty.
> You can buy a 1000MHz machine with 2 gigabytes of RAM and an 1000Mbit/sec Ethernet card for $1200 or so.
I think it's less about the web server choice and more about the architecture around pre-computing and caching certain results in memory and then asynchronously writing via a queue.
In other words, having 32/64 concurrent threads you won't get the throughput you want so agree there.
However, I'll elaborate on my point and why it's not just "use event loop".
If you block the event loop you can get catastrophic behavior, like your entire server acts as if it's a single threaded synchronous runtime.
However, even with not blocking the event loop, you need to have optimal IO __given a coroutine__. If you are waiting a bunch of ms for multiple network calls and also a db write within a coroutine, your latency is going to go up and you might not hit your latency requirements.
I was specifying that the architecture was taking these concepts into account and getting low latency on each request in a high throughput way by not blocking on writes, precaching info in server memory, etc.
If they had to go to a db and/or do other calculations per query instead of pre-caching regardless of 32/64 or using an event loop, it likely would have either created another bottleneck or the web server would bottleneck and they would not hit their RPS
That said: it’s a good bit of writing & interesting
However in practice it’s more around 4-6k because it’s doing database operations. Some end points are only 1500 rps while some of my read ops are in the 10-25k.
See TechEmpower benchmarks, or gRPC benchmarks as a starter. JVM workloads on those are always around top 5, at most top 10.
I agree!
It's written in php, running on a $4 a month shared hosting from namecheap.
Everyone says I'll need to move it over somewhere better if it grows. How soon? Even people with experience seem to have trouble explaining when. It always fascinates me and makes me wonder how many small webapps are on AWS for no real reason.
Personally, I need cpanel and phpmyadmin because I don't know anything else really. I still ftp files up manually from my windows computer. Going to a real server seems like it would make things 10x more complicated, and no one seems to be able to predict at what point I'll need to.
One of the core offerings in AWS is EC2. Which are just shared (up to dedicated) virtual machines, and they start at a few dollars a month. I'm in no way suggesting a change to your setup, only pointing out that AWS isn't some mythical complex ecosystem.
Is namecheap running PHP8? That upgrade alone will net you a massive performance improvement over 5.x
No doubt AWS has options that could work for me at a similar price point. But I don't think it would be at all easy from what I've seen, and I'd have zero support. Compared to namecheap which basically helps me do anything I need via live help with actually useful people in my experience.
I once upgraded from a $2.50 a month server to the $4 server for a few other benefits I needed, and performance actually went DOWN at that server IP. I asked them to move me to a different one, just based on my "trust me bro, this is slower than before" and they did move me and performance returned.
Reactive programming introduces complexity, one which may not be needed in the future due to improved core tech.
Also, vert.x are already exploring how to adapt to the new threads model. (2)
1 - https://gist.github.com/vietj/fe9f886d489853ab07111ba5715b13...
Not trying to cast doubt... just noting that I have exactly the same feelings (hopes?) but I have yet to personally confirm it, in large part because it's not yet fully GA so we haven't jumped in with our production workloads.
https://github.com/netty/netty/issues/12816
It'll be interesting to see who (if anyone) picks up Netty's mantle in the Project Loom world.
Hot reloading (to the extend that change-save-refresh would feel like working with a python/ruby projects), shallow stack traces, less 'magic', , excellent documentation, plethora of modules, less memory footprint, milliseconds start/restart (owing to compile time wiring), first class container support, dev tools, always running tests are few that makes DX with Quarkus amazing.
Ironically this would likely be too expensive to deploy on AWS serverless infrastructure (AWS Lambda, API GateWay etc) even-though it touts endless scalability as it main feature.