This classic 2007 tutorial starts by pointing out that NGINX parses the HTTP verb by looking at the second letter first, so that if it's O it knows to check for POST or COPY!
https://web.archive.org/web/20070505051653/http://www.riceon...
This classic 2007 tutorial starts by pointing out that NGINX parses the HTTP verb by looking at the second letter first, so that if it's O it knows to check for POST or COPY!
https://web.archive.org/web/20070505051653/http://www.riceon...
I think that if you want to support all verbs, you face at least 3 'ambiguities' whether you first check the first, second, or third character of the string. (It must be at most the third, as the shortest verbs are 3 characters long.)
First checking the first character is ambiguous between POST, PUT and PATCH. First checking the second character is ambiguous between HEAD, DELETE, and GET. First checking the third character is ambiguous between GET, PUT, OPTIONS, and PATCH. [0]
edit As danachow points out, the verbs are not all used with the same frequency. If real-world performance is the goal we'd presumably want to optimise for the GET case, which presumably means first checking the first character, as the 'G' is unique to GET.
[0] https://developer.mozilla.org/en-US/docs/Web/HTTP/Methods
https://github.com/nginx/nginx/blob/7587778a33bea0ce6f203a8c...
There are several points to note: This trick only works if the architecture supports unaligned memory access. Macros are used instead of functions, so you don't have to rely on the compiler to do inlining. The shifts and ors should get optimized out by the compiler. I once tried something similar in C#, but sadly the JIT doesn't do that optimization.
That method does look better. My idea assumes that the method has already been checked to see if it is a valid method, which has its own cost. Otherwise, POSABC would be parsed as POST. Their method does that more or less on the fly.
Presumably then it's relying on undefined behaviour.
https://stackoverflow.com/a/47622343/
https://stackoverflow.com/a/28895321/
https://blog.quarkslab.com/unaligned-accesses-in-cc-what-why...
if (m[0] == 'G') {
r->method = NGX_HTTP_GET;
break;
} else if (m[0] == 'P') {
switch(m[1]) { ... }
} else if ...
this should be fast enough.Something I don't see many people talk about, but I've setup web instances with both read and write clusters similar to db's, and have the nicety to getting more umph from the read cluster by making sure it only hits read db nodes and disabling any framework stuff for tracking changes on a model.
I expect anything that talks HTTP to reject requests with invalid HTTP methods. HTTP is a well-known standard and servers do not need to accomodate for sloppy, wrong, or malicious implementations.
Hopefully the actual code does a full parse after guessing what to try by the character test heuristic.
'Should' implies uncertainty; 'fast enough' implies there's no speed constraints / requirements. This particular bit of code - checking what request type is coming in - will be executed trillions of times across millions of servers across decades; it's the kind of thing you want to be as fast and secure as possible, so it's worth taking away these insecurities and guesses about performance.
First, the async model was literally years ahead of Apache. It leaned heavily on interfaces like epoll to manage large connection pools with a small number of processes, while Apache still used a thread or process per connection.
Second, it removed exactly the right features - those with minimal benefit and high performance impact. The classic example is .htaccess, which adds (at least) one stat to every single request, but in practice was only needed for the horrible multi-tenant LAMP reseller setups of the day - everyone else was fine with static centralized configuration.
The Architecture of Open Source Applications - volume II - nginx:
Here's there I ended up: https://github.com/nginx/nginx/blob/363505e806feebb7ceb1f9ed...
PS, I'm not totally sure, but they definitely use the count of letters as an optimization, and it seems they increment the bits associated with each type, so the order of the bits behind each NGX_HTTP_GET etc seems to matter...!
https://github.com/nginx/nginx/blob/67d160bf25e02ba6679bb6c3...
Someday I will understand :)
ngx_http_v2_parse_method iterates through all the tests, starting with the first test (GET).
They compare request method string length to test string length then on matching lengths compare each character in the request method string to the test string.
Fully matching strings set the request method numerical value from the test value and returns OK.
Any non-matching characters GOTO the next test.
After that it does a sanity check on the request method string characters (A to Z or _ or -) and returns OK or DECLINED as appropriate.
For the macros defining the HTTP method numerical values I think they're set up that way as bitmasks for bitwise operations.
For example, they do things like[1]
if (!(r->method & (NGX_HTTP_GET|NGX_HTTP_HEAD)))
Here HTTP request method value is bitwise AND against (0x00000002 OR 0x00000004)Any non-zero value here would be true and any zero value is false. So if the request method bit value AND matches either GET or HEAD bits then this conditional is false.
[1] https://github.com/nginx/nginx/blob/a64190933e06758d50eea926...