The Architecture of Open Source Applications: Nginx
aosabook.org
aosabook.org
I would like to mention agentzh and his team that did an amazing job in releasing OpenResty[1] which makes it easy to extend nginx with custom Lua functionality, which also happens to be the backbone of CloudFlare architecture, and the core technology being used by projects like Kong[2] when it comes to microservices management.
I built a nginx+luajit RTB bidder that did 168k qps on 8 cores. It smashed the C eventlib version and of course it crushed the Java version. Golang came in a close second but at the time with 8 cores I had to do crazy things to get it to pin to CPU's since the golang thread scheduling didn't seem very scaleable beyond 4 cores.
This benchmark shows Ngix+Lua (openresty) is more than 5 times slower than the leading ones, which can do more than 6M requests/sec.
https://www.techempower.com/benchmarks/#section=data-r11&hw=...
FTA: In February 2012, the Apache 2.4.x branch was released to the public. Although this latest release of Apache has added new multi-processing core modules and new proxy modules aimed at enhancing scalability and performance, it's too soon to tell if its performance, concurrency and resource utilization are now on par with, or better than, pure event-driven web servers. It would be very nice to see Apache application servers scale better with the new version, though, as it could potentially alleviate bottlenecks on the backend side which still often remain unsolved in typical nginx-plus-Apache web configurations.
I'm using Apache 2.4 with mpm_event + mod_proxy_fcgid and it's doing fine - 99% of the work and time spend is done in the FastCGI application anyway and for static content mpm_event is good enough. I wouldn't run a dedicated static CDN box on Apache but for everything that can run on a single server Apache can also do the job... even HTTP/2 with mod_h2 works fine as of 2.4.17
A problem with nginx is to figure out what matches in a complex config... it's not straightforward. .htaccess is nice and simple for a shared server with lot's of users.
I really like nginx but I guess most people just don't really need it. Migrating to 2.4 and mpm_event should be good enough.
As an example, I would refer you to a discussion down this thread about Apache 1.x forking for concurrent connections.
It's nowhere near that simple. You can fork lightweight like threads or heavy. One of the terrible things about Apache performance was all the heavy forking on every single request for generated pages. That was back when it was really expensive.
Nginx requires an openssl version that supports sni -- that was added in 0.9.8f on Oct 2007 (if you compiled openssl manually to enable it), and was enabled by default in 0.9.8j on Jan 2009.
Apache supported SNI using mod_gnutls since 2005; and in 2.2.12 (July 2009) using openssl.
They should never be used unless you're in a shared hosting environment where you don't have a choice.
Debian e.g. only received 2.4 with Jessie. Meanwhile, a feature-comparable nginx was available at least in Wheezy, maybe earlier.
If I have to choose between "compile my own Apache" and "just install nginx", well… Actually, I had to choose, and I picked nginx.
The way nginx handles requests and responses in an implicit event loop reminds me of a recent talk by Brian Kernighan, in which he mentions the ubiquity of the “pattern–action” model in many domains. I think it’s a very useful architectural pattern to have in mind when you’re designing a configuration system or a DSL.
I also liked this quote:
> …it is worth avoiding the dilution of development efforts on something that is neither the developer’s core competence or the target application.
Sometimes there can be a long gap (even many years) between first reading about something and first using its name in conversation. A long time for mispronunciations to stew away in my brain (assuming the person I'm talking to even knows how to pronounce it themselves.)
Just as I used to think misled (miss-led) and misled (my-zled) were 2 different words - the latter implying an element of malice
Rather like Larry Wall:
> I started trying to teach myself Japanese about 10 years ago, and I could speak it quite well, because of my phonology and phonetics training–but it’s very hard for me to understand what anybody says. So I can go to Japan and ask for directions, but I can’t really understand the answers!
Just for anyone interested: https://www.hiawatha-webserver.org/
The dev doesn't do much advertising, so word of mouth on a place like HN really helps.
Either way, I think Hiawatha is a great webserver that should get more attention. Especially since I am GPL proponent and Hiawatha is one of the only currently maintained GPL webservers.
As a bonus here is a benchmark of nginx and hiawatha under attack. (notice the service drop on nginx)
Source: I have to support that SW stack.
Fail at what? What were they were trying to "fix"?
> Yes, setting up a reverse proxy looks simpler with nginx or haproxy for that matter.
People don't choose nginx and haproxy because of the configuration syntax, they choose it for performance and features.
On Benchmarks and timelines: I agree that 2.2+ was not widely available for many years due to the update cycle of linux distributions, but at the time nginx wasn't in the distros... so it become a thing where people would yum install apache2, get a 2-5 year old version, and they would benchmark that against an ngnix from their latest dev download.
The original work for the Event MPM started around 2004:
http://mail-archives.apache.org/mod_mbox/httpd-dev/200411.mb...
The version in 2.2 was mostly focused on Keep-Alive requests. Apache 2.2.0 was first kicked out on December 1, 2005.
To go beyond Keep-Alive requests, is a set of features/patches called "Async Write Completion". Much of this work was done in 2006-2007 by Graham Leggett:
https://mail-archives.apache.org/mod_mbox/httpd-dev/201510.m...
Timing wise, most of that work did not find its way into a stable release until 2.4, which came out February 17, 2012. This is the date the article references.
That's not really correct:
https://httpd.apache.org/docs/2.2/mod/prefork.html
Nginx makes it easier to handle a bunch of concurrent connections, but it's not as if Apache simply forks for each new connection.
Back in the day, Apache 2.0 took a long time to gain substantial market share, for various reasons. The situation was not too dissimilar from the current Python 2/3 split.
Edit: But I guess Apache never forked for every connection, even in 1.3. It only forked if it needed a new child and the existing ones were busy.
Apache 1.3 did not fork a process per connection. It forked to handle additional concurrent connections, but that's very different than forking for each and every connection.
When was this written? Is it still too soon to tell? 2.5 years seems like enough time to tell?
[0] http://grisha.org/blog/2013/11/07/mod-python-performance-rev... [1] https://bitbucket.org/yarosla/nxweb/overview
location ~* \.(jpe?g|png|gif|ico)$ {
...
}
That would be pretty messy inside of JSON or YAML. You couldn't have the location line be a key for a map/hash, because you can have multiple blocks with the same "key".Besides, as bad as nginx configs are, just trying to understand which block "traps" a request is where you can spend most of your time; see ifIsEvil [1].
[1] https://www.nginx.com/resources/wiki/start/topics/depth/ifis...
However Apache just can't do some things that help you shooting in your foot. Maybe the C++ vs. Java comparison is not too far fetched.
I'm in the middle of a project built entirely in nginx and it's astoundingly performant. The restrictions on what I can do in (mainline) nginx force me to think through how I structure the blocks and directives with better logic representative of a web server and not an application server, which is what I'm used to.
DigitalOcean has a decent tutorial on now nginx decides on server and location block [1]. And, related to ifIsEvil, this blog post [2] goes a little into explaining how nginx "traps" a request. If someone has better resources, I would appreciate them.
[1] https://www.digitalocean.com/community/tutorials/understandi...
[2] http://agentzh.blogspot.com/2011/03/how-nginx-location-if-wo...
The tool interprets the nginx conf, rather than compiling any of it (as nginx does with the rewrite rules), makes it easy to log which lines are involved in the processing as it hits them.
YAML's initial release was 2001[0], likely not yet popular or stable enough for consideration.
JSON was likewise "released" in 2002[1] with a similar story.
In all likelihood, we're probably just lucky Igor didn't go with XML.
Okay I'll bite, what's wrong with writing nginx to go with XML?
XML is easily translatable to JSON, so... there's that.
Not really. XML is a document markup language, where order is important. JSON is a serialization format, and is a bit more bare-boned. For example, how would you transfer this XML to json, and back to XML?
<foo bar="baz"><bork>foo1</bork><bork>foo2</bork></foo> {
"#" = "foo",
"bar" = "baz",
"." = [
{ "#" = "bork", "." = "foo1" },
{ "#" = "bork", "." = "foo2" }
]
}
Or, for something more verbose (but maybe more intelligible), you could use "$element" as key for element names, and "$children" as key for child elements / text. (The point of choosing $ as a prefix, is it is not a valid character in attribute names, so cannot conflict with them.)I think it could be reasonably easy to come up with a config file standard in XML or JSON, but that the format will have to rely on the strengths of each. Translating between the two just becomes an unreadable mess. If anything, if I were to write an application that allowed for either format, I would come up with a separate standard for each. More code/upkeep, but when the config files are intended for humans and to be hand-written, the focus should be on the user.
First, a quick disclaimer: Apache conf format is not really XML. It leverages XML-like syntax but it's mostly not XML and avoids most of the serious problems that XML tends to bring, which I'll explain in more detail below.
XML was designed to provide structure to documents, it was not designed as a configuration syntax or a data serialization format. XML is meant for a document that already exists in its own right as a document, where the XML is added on as a layer to aid automated semantic understanding of that document's structure. It is not meant to directly represent programming data structures. As such, XML tags are designed to pop out and be visible from significant amounts of text that is not metadata. When there's more tags than text, as is usually the case when you try to use XML to write programming data structures, XML winds up being hopelessly verbose, and it's hard to avoid errors writing it (like misspelling end tags, forgetting a slash, putting end tags in the wrong order, etc.)
When not used for its intended purpose, XML winds up being hard for humans to write directly and hard(er than yaml and json) to write programs to parse it. In yaml and json, there's a standard, mostly direct mapping to common data structures in most high-level languages. With XML you have to make a lot of trivial decisions to make use of features that weren't designed for what you're trying to do. The most obvious examples are the distinction between attributes and tags: what does each one mean? What do tagnames represent? What do attribute values represent? What do attribute names represent? How do you handle CDATA that has more XML in it? XML is designed to elegantly handle something like this:
<A>first section <B>marked up section</B> second section</A>
But this kind of structure is horrible for a configuration file, unless the CDATA sections are a parsed language of their own and parsed externally, which is essentially what Apache does. If you're trying to use XML to specify data structures like lists and trees, it's messy. Consider this example: <VirtualHost>
<ServerName>my.server.domain.com</ServerName>
<DocumentRoot>/var/lib/www/my.server.domain.com</DocumentRoot>
</VirtualHost>
You might envision "VirtualHost" to be an item in a list, where the value of that item is a dictionary with subkeys specifying "ServerName" and "DocumentRoot." But in fact, there's more to it than that. An XML parser also gives you all the whitespace in between those two tags. You can discard it, you can write checks to ensure that nothing ever ends up in that unused CDATA area by mistake, you can write tools to generate the XML-- but no matter what method you choose it's something you have to think about that just doesn't come up if you are using a language designed for writing programming data structures instead of abusing one designed for marking up text documents.And that example highlights another problem with XML which is a flat out lack of support for common programming data types such as lists and integers. In the example above, how would you know that "<VirtualHost>" represents an item in a list, but "<ServerName>" should be a key in a dictionary? XML doesn't help you there, every parser decides for itself.
The fact that you have to pick a convention is the crux of the issue. "Easy" is a subjective term, but the fact is there isn't a direct mapping between XML And JSON.
It means that if you're using XML and converting it into a programming data structure, you have to make a bunch of decisions about how to handle the XML. It means that if you're converting a programming data structure to XML, you have to have a bunch of specific rules for how to generate that XML.
With JSON, you only have to make those decisions if you need to use data structures that aren't supported by JSON.
1. They created a paid version that has some additional basic features such as cache purging, dynamic upstream name resolution, and a few others. Charge for support, charge for some fancy management interface, monitoring, but for basic features (most of them available in Tengine [0]) - thanks, but no thanks. You lost me as an evangelist! In fact, in may aspects, they now are catching up with Tengine!
2. Instead of making LuaJIT integration standard and avoid the need to escape Lua in the configuration files, they invented some subpar JavaScript. People already use Lua widely, it's fast, it's great - don't you have anything better to do than invent yet another language!? I really can't believe pragmatic people would have done this, honestly! Speaks so badly about their thought process! I know can expect anything stupid from them!
3. The configuration language is not very intuitive. If they embedded Lua, the whole configuration could be a Lua script that initializes some internal state. This would have been a dream come true!
Web servers can have notoriously complex configs, up to the point where designing a mini-language might a worse idea than stripping down an embeddable language, such as Lua.
Lua, especially when sandboxed, seems like a fitting configuration language: http://stackoverflow.com/questions/1224708/how-can-i-create-...
I actually have some special corner case API endpoints that nginx just simply can not handle in the manner I would prefer. Further, the regex based location syntax, and even the prefix based ones, are not really what you want; you want a "Path" object that's aware of what / in a URL means, and does the right thing if you do/don't add it to the URL. (And it's not as easy as "/foo/bar/baz/?")
[1]. Conditionally streaming an upload to a backend (i.e., if auth fails, don't stream) is impossible; nginx will buffer the entire request body, either in memory or on disk, and there is no way to change this behavior.
(You're definitely not the only one to ask that question.)
I've been toying with getting a couple home servers going (replaced home service w/ business service, just installed two router based DMZ, looking at lightweight hardware -- probably will be Fit-PC products). I was going to run a separate reverse proxy and Lighttpd or similar but a quick glance makes it seem like Caddy could be used for both and more easily. Thanks for the link.
(Edit. BTW, you're not here: https://en.wikipedia.org/wiki/Comparison_of_web_server_softw...
You can do things like
{
"comment": ["blah blah comment message"],
"k1" : "v1", "k2" : "v2", ...
}
Or some silliness like that.https://groups.yahoo.com/neo/groups/sml-dev/conversations/to...
Well ok. But I'd still argue that Nginx was invented before the popularization of YAML. I think I didn't see YAML until I got in touch with Rails in 2007, which used the format extensively. And even then, almost everyone I met didn't take it seriously, saying that it is a joke until Microsoft, IBM, etc support it.
Note that it's not really fair to call Apache's configuration language "XML". Apache relies on XML for some structured data in its configuration file, but all the individual directives are parsed separately from the XML.
import json
import sys
sys.stdout.write("%s\n" % json.dumps(json.load(sys.stdin).get(sys.argv[1], None)))
This will accept JSON on standard input and will return the value of an object with the key name specified by the first command-line argument. $ echo '{ "value1" : { "sub-value": 5 }, "value2": 99 }' | python json-test.py value1
{"sub-value": 5}
$ echo '{ "value1" : { "sub-value": 5 }, "value2": 99 }' | python json-test.py value2
99
Such a program will look similar in any language with a json library that maps objects and arrays to native data structures. Granted, my simple tool will fail if the JSON isn't an object, but it's a very simple matter to extend it to handle lists and literals. For many applications this is a huge advantage over XML, especially if the point of the JSON isn't configuration but rather inter-application communication (aka data serialization).(1) https://www.nginx.com/blog/launching-nginscript-and-looking-...
(disclaimer: I work at NGINX)
content_by_lua_block {
ngx.say("hello, world")
}1. Nginx
2. IIS 6 (strange metabase thing)
3. Apache (XML, mostly)
4. IIS 7, 7.5 (XML, but with some of the files scattered through your Windows directory, and also some values aren't valid in some of the files).
5. Tomcat (XML plus madness).
I'd say the thing that makes all web-servers a pain is debugging which rules are passing/failing, and where are they sending their results to.
I think the thing which makes Nginx easier than the others is probably that it doesn't try to support the 'shared hosting' scenario, which adds a lot of mess.
I would prefer it if it were more like openssh or supervisor. Though I suspect those styles of configs are would make some of the more advanced configurations a pain.
I could see for example nginx in front of cowboy. It would strip away ssl, serve static pages, maybe authorization/authentication and then proxy connections to cowboy servers in the backend for application logic.
I find it sad that people expect good services built on top of it to be free as well. Without an enterprise/paid offering how else do you suppose people fund nginx? Right now the state of open source funding is abysmal.
(disclaimer: I work at NGINX.)
The Internet[0] has a much richer history and larger ecosystem than just the World Wide Web. The Internet started nearly six decades ago, the web has only been around for a bit more than two.
Architecturally and usability-wise, we're more or less at the same place as a decade ago:
In 2004, both nginx and Gmail were released, and people were going ape about exciting "Web 2.0" technologies like DHTML and AJAX, whose paradigms more or less still underpin all modern development. There have been a lot of additions and streamlinings, but "dynamic pages/apps in the browser without a pageload" were the modus operandi then, and are the MO now.