The long life of Apache httpd 2.4
utcc.utoronto.ca
utcc.utoronto.ca
Another note on this 10 years anniversary is that, though Apache 2.4 was released in 2012-01, it was a rushed release and some features were still lacking. In my opinion, the most important was that mod_proxy could not talk to unix sockets. It made Apache 2.4 hard to use in many cases. This feature only appeared with 2.4.7 in 2013-09.
Nowadays I rarely see Apache installed on web servers that I work with. In one case, it was because the sysadmin was used to it (though 2.4 broke the compatibility with 2.2's config syntax), he didn't consider changing and even kept the awful Apache prefork engine. In other cases, it was because of some modules (e.g. shibboleth) that were packaged by the OS for Apache, but not for Nginx.
While setting up Apache 2.4, I came to like certain things about it. For example, https://httpd.apache.org/docs/2.4/mod/mod_macro.html. Configuration seems less footgun-y than that of nginx (see https://news.ycombinator.com/item?id=30435282). I have also adopted https://tcl.apache.org/rivet/ and use it on the LAN server where someone else might reach for PHP. It is admittedly quite niche.
I'd love to be shown a better way on that, but Apache did rack up a W for that one use case.
> Caddy 2, my go-to HTTP server, lacked a WebDAV plugin at the time. [...] Since then https://github.com/mholt/caddy-webdav has come out for Caddy.
1. I have discovered that Matt Holt published the WebDAV plugin shortly before the launch of Caddy 2. So there was never a time when Caddy 2 was officially out and without a WebDAV plugin. I found out about the plugin later.
2. Caddy 2 actually launched in 2020, and I set up my server in 2019, when I only used Caddy 1. Caddy 1 did have an unofficial WebDAV plugin, https://github.com/hacdias/caddy-v1-webdav. I remember finding some compatibility issue with it, too.
Eh, the prefork engine isn't awful. It's not without limitations, but it's an excellent fit for some loads and a bad/terrible fit for others.
The main limitations are:
The child ramping behavior: you should set MinChildren = MaxChildren = StartChildren. Ramping children on demand sounds nice, but it's easy to run out of memory when all the children start or spend too much time on childinit instead of serving requests.
Anytime the child spends connected to a client socket but not actively working. Some of this can be mitigated. Turning off http keepalive (or handling it elsewhere) and setting large socket buffers allows the server to generate a response, give it to the kernel, and move on to the next client. If you're often waiting significantly for disk i/o, Apache pre-fork isn't a good fit.
Http keep-alive is a big deal for some uses, Yahoo had a nice hack for that, which I don't think made it out: have sockets wait for complete requests in another process, and send those sockets to apache over a unix socket; after apache is done with it, send it back to the 'cheapalive' daemon. That adds a bit of complexity, and I don't know of a public implementation, so not super useful to people, but it's possible to build.
I personally think prefork apache + mod_php is simpler than some other daemon + fastcgi php. But that's debatable.
Edit: oh and one more... .htaccess is clearly a performance suck. Perhaps a useful one depending on your application, but there's good reasons nobody else does it.
nginx is at: 26% Apache is at 20%
It doesn't sound like nginx is even close to having this wrapped up; which is a really good thing because I've tried nginx and the config and design sucks. Apache does so many things well.
From the wiki.
Personally I use apache as I have very few users (internal facing stuff), it has a nice simple mod_auth_openidc config which remains the same from one year to the next
I do have to recompile ws_tunnel though as I proxy a broken product I proxy which doesn't accept the correct standard for camel-capped "WebSocket" -- my "fix" script runs
sed -i 's/WebSocket/websocket/g' ./modules/proxy/mod_proxy_wstunnel.c
and recompiles
I'm one of those who always just stuck with Apache because 1) I'm used to it, and 2) it's never been close to a bottleneck in our stack. Especially now that the vast majority of static content requests don't even hit the servers at all. Buy I am curious! I haven't looked at comparisons in years, but last I did, everyone liked to compare nginx to prefork Apache, which is pointless.
Its like having the same waiter vs a pool of waiters
Sources: https://serverguy.com/wp-content/uploads/2020/10/Apache-Vs-N...
https://www.google.com/search?rlz=1C1CHBF_enUS874US874&sxsrf...
People sometimes like to have opinions about stuff. But there's no reason to google for secondary sources when the documentation is good.
Read the manual instead. The documentation for both projects is correct, short and to the point. They are both multi threaded, multi process, with an event loop. The documentation describes how to set those parameters. Look at the thread_pool directive for nginx and AsyncRequestWorkerFactor for Apache for example.
You could of course configure one thread per connection, but that only makes sense under specific circumstances.
TLS kills this, alas, and, I think, with TLS OpenSSL will be bottleneck for static content, not server.
OK. But you might as well use a fast setup if there's no cost in complexity to do so.
Since then computers (and presumably also Apache) have become faster and I don't really know how it compares today, but at least in the past it really did make a significant enough performance difference to care about.
Which use case is that?
> like fresh snow under foehn
I am from none of those places, speak no relevant languages other than English, and I had no trouble with the phrase in question.
I appreciate your effort in trying to stick up for me, though (: Perhaps if you want to be more helpful for those people who do conform to your assumptions, you could share a definition, in addition to your criticism.
"Fohn, an Englosh derivate of 'foehn', means 'a warm, dry wind', and is known for being particularly effective at melting snow."
Which coincidentally happens to have a similar contention as standing on fresh snow compacts it, but only from the comments did I guess it was supposed to be a melting analogy.
* applications are built to change them dynamically, making them part of application code, instead of you know... just having proper request router in app
* It is not "enter URL in, get rewritten URL out". You can do conditionals based on existence of files and a bunch of other stuff which means it can't be isolated from the app overusing it easily
Rewriting on loadbalancer to mangle some paths is to be avoided but overall not too bad to manage but what can happen in mod_rewrite by badly designed apps make app basically glued in into apache.
nginx is very powerful and fast.
nginx could send static file with one syscall and zero-copy if OS supports it. It could exploit all system async mechanisms - POSIX AIO, kqueue, whatever you have. Linux was worst in this regards before io_uring, and FreeBSD or Solaris were much, much better platform for nginx.
For example, Igor Sysoev (mastermide behind nginx) pushed patches to FreeBSD kernel to shave off one syscall per file on this system.
Now, when everything is HTTPS, Linux have io_uring and Igor left nginx after F5 acquisition, it is not very useful - nginx is stagnating, support for KTLS is absent, support for io_uring on Linux is absent (though there are some forks and patches for this)...
But compare nginx and lighttpd is simply stupid, it is apples-to-oranges comparison.
Each one has its niche, of course.
Even Youtube used Lighttpd at some point.
Both are http servers, and lighttpd came out before Nginx, they both tried to solve the c10k problem.
if (st && S_ISREG(st->st_mode)) return HANDLER_GO_ON;
Change it to if (st) without a regular file check, and it works as expected (great) with frameworks like Joomla etc (/administrator URLs, for example)
It had not seen any bugfix releases for quite a long time at that point, and nginx was available by then, so I made the swap. I continued to serve that website off of a single server for the rest of its lifetime.
It could probably replicate Linux's solution to this and stop fixing "major" version numbers. Make the next release 3.0, then 3.1, and so forth. And every so often, bump that first part so it's not on "3." forever either.
On the other side, who remembers the version of Chrome, Firefox, Thunderbird or even the Linux kernel since they started moving fast? They basically become versionless, we just upgrade to the latest one prompted to us by their autoupdates or by the OS.
Incidentally, that means that Apache is not, in fact, using SemVer. They're using the old-school even-odd system.
The x number (in x.y.z) was still used to denote a major upgrade but jumps between y were usually considered breaking changes too.
It’s also worth noting that it used to be common to have 4 sets of numbers, x.y.z.n with n being automatic build numbers.
Lastly, there’s still another common trend to have version numbers based on release date. Ubuntu do this, for example. I personally quite like this approach as it communicates a little more than just a release version. For internal projects I often go even further and recommend a git hash included as part of the release version.
So yes, versioning schemes are older than the web. However agreed standards in how to use version numbers have always varied a significantly (and still do now). And the Apache httpd project does predate the rise in popularity of semantic versioning.
I do like when first number change signals need to read the changelog (as in "backward-incompatible changes"). Semver or not.
For something like Ubuntu, signalling the age of a release to end users is more valuable than following semver strictly (not to mention that the patch part of semver wouldn’t even apply).
Whereas if your software is being consumed by other software, such as an API, then you need to communicate where interfaces might change in ways that are verifiable via other software interfaces.
So it really depends on who is consuming your software interface. You might even want multiple different version schemes for different parts of your application. You’ll sometimes see applications have one versioning scheme for their end user facing interface while their developer APIs, which are still bundled as part of their application, might follow another.
Going back to the original point about dates as version numbers, the biggest appeal for me is it takes human error out of the equation. I’ve lost track of the number of companies I have joined which were very good at supporting versioning but terrible at remembering to increment that version number. So having a system which can auto-increment version numbers takes one problem away from developers. And if you’re auto-incrementing private packages then the question becomes: “how strictly are you now following semver?” At which point you might as well have a timestamp and then have the bonus feature of communicating to your team when a specific build was committed.
I don’t think there is a right or wrong way to do these things though. You just need to pick a standard, document it and be committed to following it.
I wrote "changing primary version number" instead of "using semver" on purpose. I agree semver can be too narrow for bigger apps and that is why I only said about signalling "potential sysadmin action required", not signalling new features or bugfixes in version number.
> standalone applications,
can certainly have "we changed how configuration options work" at the very least
> distributions of packages
Distros are so far managing that extremely well all things considered
> nor operating systems
Cue loud Linus yelling about breaking userspace.
The OS problem is problem of having different components with different lifecycle.... but my iptables rules are working despise the fact iptables in kernel was entirely replaced by nftables so it certainly can be done just fine.
I was making a general point. I wasn’t suggesting you were advocating semver specifically.
> can certainly have "we changed how configuration options work" at the very least
Any major change to your versioning scheme can signal that. Including having a year number as your major iteration.
But what’s more valuable to your end users in these types of scenarios is documentation about the change.
> Distros are so far managing that extremely well all things considered
You’ve completely lost me now. It sounds like you’re trying to defend Linux distributions but my comment was neither a criticism nor was it about Linux distributions per se. These are plenty of instances where several packages are distributed as a whole. Like with multimedia applications that might ship ffmpeg, a Python scripting environment and a multitude of other components.
Linux distros, of course, is another example (and one on a much bigger scale). It’s worth noting that no major distro uses semver either. Some don’t even have a minor release let alone a major version number. But their signalling around versions satisfies a different problem to what an API needs to signal. Hence my point that versioning is about sticking to a standard more than it is about trendy schemas.
> The OS problem is problem of having different components with different lifecycle.... but my iptables rules are working despise the fact iptables in kernel was entirely replaced by nftables so it certainly can be done just fine.
I don’t really understand your point here but I’m happy for you that your config is still working :)
They could adopt it, but that'll mean the next release is 2.5, and if a feature is added, 2.6 after that.
Meanwhile nginx was the authors own homepage, complete with hammers and sickles, with a personal anecdote about how he read the c10k paper (describing how you should have 10,000 connections on a Pentium III ) and decided to use nonblocking io make a fast Web server.
Nowadays nginx is the safe choice and Apache is seen as a little bit old hat.
Like how example in Apache ordering of vhosts doesn't matter, except for choosing the default vhosts
Non-conformant to what? Is it following a spec somewhere?
My brain's shot and I thought they were talking about some sort of server spec (not sure for what? conig? request handling?)
https://www-archive.mozilla.org/party/2002/flyer.html
I actually liked the commie vibes in the free software movement. In some ways the movement was attacking big-capitals, and was attacked by big-capital.
Too bad Red Flag Linux[1] (from China) and Red Star Linux[2] (from North Korea) are not actively maintained. Nova Linux[3] (from Cuba) seems to be the only distro still somewhat maintained.
1: https://en.wikipedia.org/wiki/Red_Flag_Linux
> That was complete bullshit, of course. Yes, I absolutely branded Mozilla.org that way for the subtext of "these free software people are all a bunch of commies." I was trolling.
> I trolled them so hard.
No direct link to source because jwz.
https://dereferer.link/?https://www.jwz.org/blog/2016/10/the...
[1] https://w3techs.com/technologies/history_overview/web_server...
I declared that it would be the end of innovation in httpd.
Microsoft, at the time, was making big moves in IIS in support of then-new ASP.NET web framework.
Although there was so much that could have been done to improve the DX/UX of Apache httpd, it stopped moving ahead, and the OSS community fragmented into nginx, lighttpd and others - often losing velocity because of the need to not only achieve feature parity with what httpd offered already, but also the slew of CVEs security issues that cropped up again and again as the Internet grew exponentially during that period as more people went from dialup to broadband, then into mobile.
yeah, there was a huge push for a breaking change from 2.2 to 2.4, but it was very well thought out. proof is that we had 10 years of security and features and nobody even noticed while updating.
that's actually great. and apache can keep with ngix without breaking a sweat if you know what you're doing.
the fact ngix took over apache metrics overnight is because few people know what they are doing, and while keeping apache, put ngix in front back in the days.
also nowadays every language packs their own server, which learn all the old lessons again, slowly...
Nginx is a Russian (in its inception) product. The reason one should not be worried about this (as one is worried about Kaspersky products) is that Nginx is open source.
Do I have this right?
As far as benchmarks go, nginx is much faster using less memory.
Apache2 could save energy on the scale it is still used.
I stopped using it for 7 years now