Nginx Is Taking Over the Internet
wired.com
wired.com
I should mention that Apache had been good to us in the past, despite being the 1.3 series with patches bundled with OpenBSD, but we were expecting far more from it than it was originally intended to do. In that regard, configuration and deployment became much simpler since we moved to Nginx.
You may have seen the same or related problem, since it takes a bit while for packages to come to OpenBSD. But either way, you ended up with something that worked well :)
And that's how a lot of countries will feel about using proprietary (or even open source) software from US in the future, after all the NSA revelations.
There's a lot of reasons to prefer free software over mere open source software, but privacy is none of them.
Given Apache's track record and massive deployment, how could this be possible? Isn't it more likely that the Wordpress people were doing something wrong? Not that Nginx isn't great, but I'm bemused by the occasional suggestions I see that it's saving us from the suddenly broken Apache.
http://www.iis-aid.com/articles/my_word/difference_between_p...
"Another option is to configure IIS to use PHP in FastCGI mode which allows PHP processes to be recycled rather than killed off after each PHP request and also allows you to run several PHP processes at once, making PHP much much faster with the added bonus that as it is using the CGI interface there is little or no incompatibility issues with PHP extensions. This is still the fastest way to serve PHP, and is the way the IIS Aid PHP Installer is configured to install PHP on your IIS Environment."
The problem with WP and PHP is not stability, but large, complex and suboptimal code. One of the reasons is because WP supports everything and a kitchen sink, and all the features were just patched over it as time went by.
I run custom PHP applications for thousands of users on Apache without problems. I am using nginx in front of it for KeepAlive and static files, which reduces RAM and CPU usage drastically. And it also makes everything faster for website visitors.
There's nothing wrong using both Apache and nginx, each has its purpose.
You misunderstand. You don't need to write threaded code to write code that can't be used in a threaded environment. Think about using and accessing global or static variables in a threaded host environment vs a prefork process based environment.
PHP devs have adopted the practice of using Apache pre-fork rather than simply fixing their bad code because it hides their mistakes. On top of that, Wordpress being a huge plugin based community, you get tons of poorly written plugins that assume they're running in a single process environment because the authors simply don't know any better.
As far as I can tell, most PHP devs seem to think Apache only has a pre-fork mode; they write lengthy articles about how bad Apache is on memory not realizing the only reason it's that bad is because they're using the oldest and most resource hungry worker mpm. They don't seem to be aware Apache has other worker modules like threaded and evented that don't suck down ram like the prefork model does.
From what I've seen essentially Wordpress is just super heavy and tie that in with a more memory hungry MPM+PHP it quickly leads to Apache lockups. I don't think this is a problem with Apache at all, nor even a PHP problem but simply needing to have lighter-weight PHP applications, otherwise you are going to have to start resorting to tricks to get some good RPS counts out of your configuration.
My end of story: Wordpress is awful to scale in Apache and it's not an Apache problem.
----
Hi,
Barry Abrahamson here. A bit of context to my comment about the stability of Apache. We started WordPress.com using the Litespeed web server, but decided that we wanted to maintain as much of an open source stack as possible. We looked at lighttpd but had lots of issues with memory usage (leaks?) so we naturally turned to Apache. It worked mostly as expected performance and functionality wise. Keep in mind this was back in the Apache 2.0 days. Our problem was that deploying configuration changes which required a graceful restart under production traffic levels didn't work very well. About 10% of the time it either failed to reload the configuration, Apache stopped responding completely, or connections were dropped. Keep in mind that we were reloading ~1000 machines at once, but 10% is a pretty high failure rate. Yes, I know we could have removed machines from production, reloaded the configuration, and then put them back, but it's slow and error-prone, even if automated.
This is when we started looking at nginx. The nginx configuration reload works as expected all of the time, so that was enough to make us invest more time in the software. The transition was non-trivial since we had thousands of lines of .htaccess files which had to be translated to nginx configs, but it was worth it. Today, nginx is the "Swiss Army Knife" of our software stack. We use it for everything we can. I joke that we would use it as a database server if we could :)
------
As a side note, this may be the first useful comment on a Wired article I've ever seen.
Coincidentally, you probably could use Nginx as a rudimentary database at least since it can be extended with Lua. You can probably make it function as a database proper by using OpenResty. In fact, there are quite a few modules just for that purpose http://openresty.org/#Components
tl;dr the claim that Apache was the problem is Marketing Bullshit.
No, the threaded worker is the default, not the prefork.
If you consider a slow client that consumes 1mb of html in, say, half a minute, an Apache thread(2.0)/process serving him would be effectively locked during this time. And thread/process has a relatively big overhead, both memory and cpu wise. Nginx, on the other hand, would just switch to serving the next client/clients, whos socket is ready (using epoll/select).
At first, I poured so much into learning Apache I didn't want to learn this weird obscure Russian alternative but I am sure glad I did.
Under load testing, I can have double the amount of concurrent users without even causing my server to hiccup.
When I was running wordpress and apache in AWS I had to use a costly m1-small instance to maintain the cpu and load needs of apache.
When I made the switch to nginx I doubled the performance and was able to move to a free micro instance in AWS.
In conclusion, moving to nginx saved me money and allowed my server to run more efficiently.
I am a believer in this software. Hopefully, going commercial doesn't ruin it like money usually does.
The legal profession has the bar. Doctors need accreditation. Builders need licenses for their jurisdiction.
But software? It doesn't matter who wrote it or where because bits are just bits, especially when you tread into opensource land.
We're a meritocracy because that's the nature of the industry, not because we're particularly benevolent.
For now nginx is first choice for Webperformance, but in some years the next thing will arrive.
See http://wiki.dreamhost.com/Web_Server_Performance_Comparison for instance.
Immediately followed by:
> Apache can work with PHP-FPM just like Nginx using the new evented worker mpm
seems contradictory. Can you clarify this bit?
Prefork - if you want to use mod_php, you're stuck with this. I think this is what you're referring to as "the process model"
Worker - Uses threads to handle connections. A big step up from prefork, but not as awesomesauce as nginx.
Evented - The new MPM which use epoll
It's been possible to run PHP(-fpm) under "worker" (via FastCGI) for years, and it performs adequately. I never benchmarked, but I was using this setup back when nginx was still too bleeding-edge for my tastes, and it certainly beat the pants of prefork/mod_php
[0] http://w3techs.com/technologies/history_overview/programming...
I use nginx too and see the difference. There is a place for apache, but not the leadership of webservers.
found here: https://en.wikipedia.org/wiki/Igor_Sysoev
"This is an interesting interview with a creator of one of the best web servers out there. The original interview is in Russian and the translation (by google) is fairly difficult to read. I took a crack at writing a better translation and I think this might be a bit better."
This was another one I just googled: http://www.freesoftwaremagazine.com/articles/interview_igor_...
An example. Assuming you have the domain “example.com” then you might have an app that handles user login and you spin that up on, say, port 30000 and you map it to port 80 such that is appears as:
You might also have an app that handles user signups which you spin up on port 30001 and you map to port 80 such that it appears as:
You might also have an admin app that you spin up on port 30002 and you map to
You might also have an app that allows users to update their profile information and you spin that up on port 30003 and map that to
http://www.example.com/profile
And you might have an app that publishes much of your content as static HTML files, which you spin up on no port, as it does not accept TCP/IP requests — instead it queries the database and then creates static html files, which you save to some directory such as /var/www/example.com/public_html/ and you map that to
You can see how this gives the sysadmins a lot of freedom to spread load across servers in creative ways — if 1 app becomes especially popular, or resource hungry, the sysadmins can rather easily move it to its own server (or set of servers). This is one of the main reasons most big sites move to an architecture like this — it facilitates fine grained control of what sort of requests go to a particular server.
This is a flexible style and it allows small, maintainable apps. By contrast, consider web development circa 2001, when many people felt it was enough to put a big blob of PHP in a directory and let Apache serve it as one big app.
Nginx's fast reverse proxy allows developers to focus on building their apps, without having to worry too much the server details. It also offers a cleaner separation between concerns that should worry programmers and concerns that should worry the sysadmins.
(A final point: in my own apps, for information that needs to be shared quickly across apps, like which users are logged in, I use ZeroMQ to knit the apps together.)
The problem people have with Apache usually isn't really apache, it's with PHP forcing them to use the process based mpm which makes it a pig. This is PHP's fault, not Apache's. use the threaded or evened worker mpm's and it's not a pig at all.
It is also remarkably easy to set up SSL (minus generating certificates) -- it's only 3 lines.
I really can't tell if you're being deliberately obtuse about this. Don't you see all the people going "Yeah, it was faster to do this with nginx than it was with Apache"?
Are you the sort of person that's surprised PHP got so popular?
I prefer nginx' configuration syntax over apache's though.
I didn't say it wasn't simple. But simple is just sufficiently managed complexity.
This page should be all the proof I need: http://www.apachetutor.org/admin/reverseproxies
where as reverse proxies in nginx: http://www.cyberciti.biz/tips/using-nginx-as-reverse-proxy.h...
So maybe you could make a case for nginx being less modular (or obviously so) -- but 1 config files with the right magic spell beats downloading x module and setting up x directory PLUS the magic file.
Also, I think that in general, nginx fits more into the many-apps-different-implementations-same-server vision of the modern day web app.
That's 26 lines of configuration, compared to the nginx example, which is 54 lines. Both are kind of arcane looking, if you haven't used it before.
True, they're not night and day, but it's still stuff to do. Also, the page I sent, the actual reverse proxy is only 14 lines, half of it is extra stuff... The section where it says "start with example.com".
And that's the whole server, whereas the
Configuring Apache is easy as there are tens of thousands of tutorials showing you exactly how to add those two little lines of code to enable proxy and/or a cert. Both Apache and Nginx are easy to configure, Nginx is easier, that doesn't make Apache hard.
> Are you the sort of person that's surprised PHP got so popular?
Nope, I've worked with tons of PHP developers and I know exactly why it's popular. Wide support, easy copy to server deployment, and in large part, dumber programmers who want you to give them the codes because it's a copy and paste community in large part. All languages have a wide variety of users, but I've found the worst of the worst go to PHP like a month to a flame exactly because it's so simple to get started and despite how terrible a language it is.
Maybe it's just the wrong tool for this time, and it used to be right tool in the past.
ProxyPass /foo http://foo.example.com/bar ProxyPassReverse /foo http://foo.example.com/bar
I've used it for years, it's not hard.
Edit: Not to mention sticky sessions per proxy which simplifies user management.
Didn't you just contradict yourself? How can it be simpler and all be simple? Simpler implies one is greater in simplicity so which is it?
FWIW, we were running Apache 1.3 with patches on OpenBSD. For us, using the 2.x branch wasn't possible without staying atop the modules for code issues/vulnerabilities. Considering other client expectations, we didn't have the resources to audit everything like that. The fact that OpenBSD 5.2 included Nginx by default was a nice plus for us too, but we started the move before then.
"which is it?"
He just said which one it was. I don't see how anyone could fail to understand this.
No, I did not.
> How can it be simpler and all be simple? Simpler implies one is greater in simplicity so which is it?
Is English not your first language? I can choose 5 simple things, and then compare them to each other and say one is simpler. They're all still simple.
> FWIW, we were running Apache 1.3 with patches on OpenBSD. For us, using the 2.x branch wasn't possible without staying atop the modules for code issues/vulnerabilities.
Sounds like your OS was the real culprit for not keeping up with newer releases. Again, if Nginx works for you, great, you can promote that without putting down Apache which is quite simple and rock solid stable as well.
As someone who's apparently more fluent than I am in English, you should have noted that I'm not "putting down" Apache as much as showing that, out of the box, Nginx still is a better option than Apache. This isn't my opinion as much as it is empirically observed fact both in our environment as well as our clients' environments.
I didn't insult you, it was an honest question as I can't assume everyone here is a native English speaker and your confusion could have easily been a language thing.
> And English is my first language, thank you.
OK, great, then I don't know what was at all confusing about what I said.
> Apache as much as showing that, out of the box, Nginx still is a better option than Apache.
Such blanket statements are simply unsupportable. It depends entirely on the workload, needs, language, and configuration. Out of the box, Apache isn't configured for the process model that everyone complains makes it a pig.
This is the stupid HN system at work: Someone who doesn't understand English being a native speaker (simple vs simpler) can downvote you because he has "earned" that right, by submitting a gazillion articles with 1 point each.
While you're here, you may want to take a browse on the guidelines http://ycombinator.com/newsguidelines.html
Newer releases are quite bad from both a security and a stability standpoint. They did not make configuration simpler, so the idea that using the last reasonable release of apache was the problem doesn't make any sense.
There was no need to be rude.
Two feet is not very far to walk. Three feet is farther, but it is still not a great distance to walk. Two things can both be short/simple even if one is a bit shorter/simpler.
I used to use Apache all the time way back when. The memories that stick with me are of a massive tangle of configuration rules, mod_*, and htaccess files. I've never looked back since going the way of Nginx. It's just so lightweight in comparison, both performance- and maintenance-wise.
Apache httpd is still bloody awesome.
It just struck me how starkly the difference in resource consumption between it and Nginx was illustrated in that case in particular, as the rPi handles the latter just fine.
All the discussion around Nginx recently, has me wondering if I should revisit my own proxy project: http://switchflow.org/
Nginx has taken over some of Apache's market because it's lighter on resources, i.e. with the same amount of resources it can process & deliver more stuff.
Regarding functionality and configuration there is not much difference, actually I think Apache still has the edge here.
That's precisely why we moved to Nginx. We do mostly Java so Apache was not the problem there, as it was simply proxying over to Jettys.
But the blog is wordpress that means PHP, so we had to use modphp modules with Apache. That was Okay initially, but as we scaled, we saw some 50 instances of Apache and each taking 25 Mb memory. And note, most of the bloat in memory was due to modphp being loaded in-memory in each instance.
We moved to Nginx with an external php-fpm for PHP, and just 4 Nginx instances take up the entire load and each uses some 2.5 Mb of memory.
And at the same time Google webmaster shows a significant drop in average latency.
So that's our experience of why we moved to Nginx.
Edit: typo
And I sort of felt sad doing this. Because I have a huge respect for Apache. I hope just like what Chrome did to Firefox, Nginx does to Apache, and its a win-win for all.
"More with less" here is a judgement call. You've increased the number of processes and subsystems that need to be managed (nginx + php_fpm), which is more. You've increased the context switching when dynamic requests need to be passed to another process and proxied through nginx. The communication between the frontend web server and php_fpm needs to be serialized in some fashion, so you're now incurring a greater than 1x cost in processing and parsing things like HTTP headers and the request (serializing and deserializing the data to pass that between the processes). It seems you're ignoring the actual cost of running dynamic PHP code because the cost has been moved out of the "web server", the thing that listens on port 80, and into something that isn't considered "a web server", but this is just a change in bookkeeping.
Do you know where that reduced average latency is coming from? Is it because nginx is serving static assets that would tie up an apache worker that could be running/serving PHP in apache? If so, you can do the same thing by offloading all your static asset serving to another server (or a completely external CDN, which has other extremely beneficial client-side advantages too) using apache too. That is, the choice of software here isn't as big a win as presented, it's the architecture of the ecosystem that is the big influence.
It's great that you found something that works better for your workload, and it's appropriate to use what works and not get bogged down in why all the time. But the way you've presented it doesn't indicate well that all the causes and effects are completely understood. And I think this ends up giving Apache a bad rap.
Yes, I understand that. I have studied about the C10K problem[1]. And have coded another service which just uses libevent[2] directly and some C++ code to serve a feature for our site very low latency. So I understand quite well why Nginx is offering low latency.
As you rightly observe later on, I don't have the luxury to get obsessed with all the Whys, so often shoot for the major architectural gain and take any side effects that come along in the stride.
>.. But the way you've presented it doesn't indicate well that all the causes and effects are completely understood. And I think this ends up giving Apache a bad rap.
I am surprised by your this observation. I have been 100% honest in what I wrote above, and I repeat have a huge respect and thankful for the Apache team and Apache web server software. Just that, by the page loading gains we got, were clearly very good. And also there were less errors in general reported in the web master. So we stuck with the change.
Also a point regarding the switching costs and extra processing between php-fpm and Nginx processes. As I said in my first comment, we do mainly Java and PHP is just for the blog. But I very clearly remember seeing all the Apache child procs bloating to 25 Mb after the first hit to the PHP code was done. So I am actually saving by externalizing on the resources. Trust me I know what I am doing.
Another tangential thing not related to Nginx, which I am doing for low resource consumption. Is moving some relevant code from Java to Go. And there also am seeing huge memory gains (i.e. savings). Perhaps will share more about it at an appropriate time.
[1] http://www.kegel.com/c10k.html
[2] libevent.org
Edit: grammar
Sorry, I didn't mean to suggest that you weren't being honest or that you didn't know what you were doing.
I'm sure we've all had to deal with "Well, this guy on HN converted to X from Y and saw <insert generic gains claims here>, so why are we on Y again?", which the somewhat hand-wavy details in your original comment can end up contributing to (obviously we're limited for space and attention here, so leaving out some details can be desirable). What I failed at was communicating that my comments were not intended as an attack on your specific methodologies or choices, but was more meant to make the reader of our comments consider the wider implications of blindly following the herd and not making informed decisions. I have had people forward me lists of links to isolated, comment-less HN comments as "support" for their position.
But I very clearly remember seeing all the Apache child procs bloating to 25 Mb after the first hit to the PHP code was done. So I am actually saving by externalizing on the resources.
Yes, sorry, I didn't consider the exact traffic ratios between the (proxied) Java requests and the (internally handled) PHP requests. If you've got a wide spread there favoring proxied requests, then, as you've experienced, it would be advantageous to get rid of the internal handling, proxy all requests, and outsource the bloating to other processes/systems where it is more isolated, and the upper bound on the bloat be more influenced by the lower total requests for the resources that consume more. And once that's done (the getting rid of the monolithic parts) then it's a lot easier to experiment with replacing the different parts to see if there are other gains. You may have seen similar resource consumption benefits by nginx proxying to Apache+mod_php or even to bare PHP cgi.
But one shouldn't even follow that method blindly either, because it's all based on workload, and everyone's workload is different.
There are so system management and monitoring benefits and drawbacks that go with these kinds of changes. We both know that there is no software selection silver bullet, but it's not necessarily the readers of these comments that know that.
A quick google search finds a nitty gritty page that might help - http://foertsch.name/ModPerl-Tricks/Measuring-memory-consump... there's probably better examples.
Yes nginx does have lower memory usage, but on most non vps/shared servers this is a lot less of a concern than most people think vs Apache.
php-fm works great with Apache as well and probably accounts for the huge resource gain over mod_php that Apache uses by default.
Apache, and specifically mod_php, doesn't have that level of control over memory usage (mod_perl (and, incidentally, mod_python) has deeper integration with Apache to encourage, with the right configuration, more of the memory to be shared between processes, but it's still not that great). The Apache parent process, primarily exists to do process, signal, and socket management, it's not really possible to do a lot of application level (pre)processing before subprocesses are forked off, which is what would be required to share a significant portion of the address space.
If you have a .php file run via mod_php that looks like this:
$x="";
for ($a=0;$a<(1024*1024*100);$a++) {
$x.="1";
}
This will produce 100MB of non-shared memory in a child process, when the request is made and serviced by that child. And, because of the way requests are dispatched to children, at a low request rate the same child could end up serving all (or a majority of) requests. This unbalances the memory usage between processes.A quick google search finds a nitty gritty page that might help - http://foertsch.name/ModPerl-Tricks/Measuring-memory-consump.... there's probably better examples.
That is a great link in general for how shared memory and COW works.
https://httpd.apache.org/docs/2.4/mpm.html#defaults
In httpd-2.2.x, however, the default MPM on Linux is prefork, i.e. the "bad" one:
https://httpd.apache.org/docs/2.2/mpm.html#defaults
And those would be the "factory" defaults. Distributions can still put in their own defaults, e.g. Ubuntu 12.04 LTS supplies httpd-2.2.x with the worker MPM.
Anyway, Apache 1.3.x (built-in with something similar to the prefork MPM) + mod_php was the de facto (or only?) way to deploy PHP scripts, as you can just throw the scripts into the htdocs directory and they will just work.
Edit: Here http://www.aosabook.org/en/nginx.html
Strong Russian accent : check
Over 40 years old : check
No revenue stream : check
http://en.wikipedia.org/wiki/Non-English-based_programming_l...
> In English, Algol68's reverent case statement reads case ~ in ~ out ~ esac. In Cyrillic, this reads выб ~ в ~ либо ~ быв.