Nginx: a caching, thumbnailing, reverse proxying image server
charlesleifer.com
charlesleifer.com
secure_link_md5 "my awesome secret $uri";
This means anyone can extend the URI (eg. adding "/../../../some/other/file" at the end of the filename) and compute a valid key, and force the server to access an arbitrary image file on the backend. If the backend is a private server not publicly accessible, then this may be a security issue.This is why in the ngx_http_secure_link_module's documentation the secret is appended (instead of prepended), see http://nginx.org/en/docs/http/ngx_http_secure_link_module.ht...:
secure_link_md5 "$secure_link_expires$uri$remote_addr secret";
The Nginx doc is also at fault because it fails to explain why it is important to append the secret and not prepend it. I just emailed security-alert@nginx.org to let them know.[0] https://en.wikipedia.org/wiki/Hash-based_message_authenticat...
But yes, when in doubt, and to future-proof your code, it is good practice to use an HMAC anyway. Or use SHA-3 which is not vulnerable to any attack when prepending/appending a secret.
However, I have one nitpick.
The caching server's main location is a bit of a cargo-cult unoptimization. They have specified:
location ^/(.+)$ {
When the equivalent: location / {
Is shorter, clearer, and does not incur a regular expression capture on every request that is never used.That allows "/" to be used for other purposes, e.g. redirecting to the main page of the site, or displaying a nice error page. (However, in this case it simply shows the default Debian Nginx installation page: http://m.charlesleifer.com/ )
location = / {
# Just whatever happens for "/" requests goes here
}
location / {
# Every other requests ends up here
} location = / { [ configuration A ] }
location / { [ configuration B ] }
To learn more about this, I recommend http://nginx.org/en/docs/http/ngx_http_core_module.htmlThe Nginx docs are not always great but they are far more correct than most blog posts about Nginx so that is why I have turned to use only the official docs when I have a question about how to do something with Nginx.
My ideal web server would have a configuration that resembles a decision procedure in a programming language, so I can have a mental model for what makes it choose one route over the others, and how settings are combined, etc.
Somehow Nginx repeatedly makes me and my coworkers very confused—are we crazy or just mildly incompetent? Should I read an Nginx book?
(One time I wanted to do something fairly straightforward, I think using a conditional inside some block, and when I looked in the manual it warned very strongly that this might cause segmentation faults. I was like WTF?)
I highly recommend this document to start, which describes how NginX processes a request: http://nginx.org/en/docs/http/request_processing.html
For getting the mental model right, you have to understand that NginX configuration is declarative, like SQL or PROLOG. That means as opposed to an imperative or procedural language you might be more familiar with, in NginX you describe "what" you want to have handled, not "how" and NginX worries about the "how" under the hood, for the most part.
Of course like any abstraction it sometimes breaks and leaks, but this is practical software we're talking about. One specific big headache for people is when using "if" in NginX, which is a decidedly imperative construct! As such you do have to tread lightly with it; observe If Is Evil: https://www.nginx.com/resources/wiki/start/topics/depth/ifis...
http://lh4.ggpht.com/sTP-3BqnirkHm40qfb496w85A1bf7BpeXthFJ92...
http://lh4.ggpht.com/sTP-3BqnirkHm40qfb496w85A1bf7BpeXthFJ92...
Not sure if they have improved the native library but I doubt it has the flexibility.
The important part is caching the results because of the expensive cpu time.
ps. googling about the current state of this came up with an interesting nginx module which talks to imagemagick (or gd) directly https://github.com/cubicdaiya/ngx_small_light
We ended up using Thumbor, which was able to give us much better looking images (but at the cost of being extremely difficult to deploy on the CentOS 6.x servers we had at that time).
Python is definitely an OK solution for this as it turns out. We resize a few million of images a day on-demand on two n1-highcpu-8 GCE instances. Although, they could easily be a fourth the size, as CPU generally peaks out at 20-25% during peak hours.
We've actually tried more specialized services for this, like sharp + a threadpool using node.js and it actually performed terribly. As it turns out, blocking the gunicorn event loop by processing image resizes created a good amount of back-pressure to load balance incoming requests to each worker process.
How many of those same people are now going on about how nginx has finally got some support for loadable modules and built in processing like this?
I would doubt many of them are because early adopters tend to also be the kind of people who don't mind compiling their own webserver for whatever reason: time, patience, availability of dedicated hardware to do it on, etc.
These are often people who would compare Apache + mod_php to nginx + PHP-fpm, and not understand why it's not a realistic comparison of the web servers.
We use varnish in front of nginx
In terms of this blog post, you'll only need the "resizing server". This is a very simple and stateless way to get image thumbnailing to work. I'd also personally not bother with all the API key stuff - how 'malicious' can it be to generate a thumbnail of a public image?
This will even work with 3rd party images (the case of the app I used this for, it was product icons from Apple's App Store). If you somehow want to use images hosted elsewhere inside your app, but worry about speed or about hitting their servers too much, proxy them with Nginx and let Cloudflare handle the rest.
In Cloudflare's case, they operate 76 data centers. So every "short period of time" x up to 76, your server will need to regenerate each image.
So while a CDN is useful, a proxy cache can still be helpful.
(None of this applies to NuevoCloud though.. NuevoCloud has a global cache, with dedicated caches for each customer. So it's entirely possible to keep an image at the edge indefinitely)
Maybe this is typical with 'website accelerators' as opposed to CDNs that focus on caching truly static content, but even so features like dynamic acceleration and global caching are par for the course with any traditional CDN.
Cloudflare, which I use mind you, is non-traditional and their business model seems more suited to DDoS prevention and edge SSL than anything else.
So what sets NeuvoCloud apart?
Global Cache does not exist at any CDN that I'm aware of. We spent a lot of time developing this. It's true other CDNs will cache files at each POP, when that POP sees a request for it... but that isn't what we mean here. (I understand the confusion.. I get a lot of questions about this, and we should probably rename it.)
Global Cache means we have a single cache that is used/managed globally. Let's say a visitor in France gets a cache miss for a new file on your website. Every PoP at NuevoCloud now knows of that file and has it cached. So if the next visitor to your website is in Tokyo, they'll get a cache hit and the file will be served from the Tokyo PoP, even though it's the first time the Tokyo PoP has seen a request for that file. A cache hit at a traditional CDN means that PoP has the file; a cache hit at NuevoCloud means every PoP in the world has that file.
The cache is managed for each customer individually with guaranteed space at each PoP. Which means, even if you have a small website that receives no traffic.. you can still keep your files at the edge. The only other way to get this is to pay for a dedicated CDN.
Finally traditional CDNs send requests from the edge directly to your server (client > edge node > your server) which is a long distance connection. Connections through NuevoCloud are routed through our network: client > edge node > edge node > server. So both the client and server are talking to an edge node near them. This speeds up SSL and connection negotiation between the edge and your server. We also use this setup to dynamically route requests around high latency, network partitions, etc.
I put our 10 points against other CDNs daily... and it's faster. You would be surprised at how dumb the average CDN is when handling requests. This is the reason they scatter edge nodes all over the place.
Apparently, it is possible, great! One of the google results: https://gist.github.com/hilbix/5921589
Image filtering is an expensive operation for the CPU, with large latencies as well.
You're better off performing it on a CPU, of course SIMD optimized.
Besides, web servers probably won't have GPUs anyways.
And these Intel Xeon's with embedded GPUs don't need to transfer data to the coprocessor, since they operate off of system memory, so transfer latencies are a non-issue.
I'm genuinely curious because I've never stood up Varnish before so maybe there's an anti-pattern / gotcha to learn from.
I feel like I am missing something.
He's basically just moving functionality from Varnish to nginx and thus simplifying his stack a bit.