Abusing Amazon Images
aaugh.com
aaugh.com
By encoding all the relevant information in the URL, like Amazon does, we offload such work asynchronously to different web requests. If there's info in this URL we want hidden from users (spam-protected email addresses) we encrypt it using Base64-encoded AES. If the file store were to go down--which hasn't happened since we switched to using S3--users would only see broken images instead of the entire site crashing. Further, we store the images on S3 using the exact same URL as we fetch it from, and have our static web server check there first. That means if the image is cached it never hits our application stack.
Now, when we generate a page with one of these dynamic images, our app stack only has to generate a URL instead of relying on the vagaries of our file store and image processing libraries. It's made our site more reliable, easier to manage, and faster. If we suddenly start serving up many more images, we could easily replace our image generation service with something higher-performance or more reliable, since it's just a web service. Right now, it works great.
Clearly Amazon has seen the benefits of the same approach.
If you go this route, be very careful making the URLs too human readable. When it's obvious that you can hammer on URLs to use server resources, it becomes a natural vector for a denial of service. If we went this route again, I'd go one step further than encrypting sensitive data, and encrypt anything that can somehow result in a new image.
Unfortunately, doing this safely sorta kills the user friendliness of it. It was fantastic to be able to tell authors "if you want a 100x100 image, add /100w/100h/ before the file name". I think they'd be less inclined to respond to "if you want a 100x100 image, generate a hash from the width x height x secret key and add /100w/100h/the-hash/ before the image".
It's the laziest possible approach and it has one huge hidden benefit, which is that if a bot hits your pages you're not going to end up pre-generating a bunch of scaled images that nobody is ever going to use.
A simple list of allowed sizes in the handler makes sure that someone can not use URL injection to generate a large series of nonsense images for sizes that we'd normally not use.
Typically a url looks like http://host/resourceid_widthxheight.jpg , where width and height are replaced by the desired size.
I'll be the first to admit this is a nasty bit of code but I haven't found anything else that comes close in terms of efficiency and flexibility.
This is even documented in the apache docs, albeit with rewrite rules, which is more heavyweight and less efficient. http://httpd.apache.org/docs/2.2/rewrite/rewrite_guide_advan... I believe that using ErrorDocument for this purpose is explicitly mentioned somewhere in the apache docs, but I can't find it right now. I'm pretty sure that' where I first learned of it (maybe back in the 1.3 days).
I have a feeling that if ErrorDocument had been named (or aliased to) GenerateMissingContent, then this would be a widely accepted method of dynamic on the fly generation with filesystem caching. I've even heard a rationale that having a script generate the content on error and write it into the filesystem where apache can then read it directly on later requests is reimplementing a reverse caching proxy, and you should just stick squid in front of your web servers (with the, ahem, additional overhead of having to run and mange two services). This seems like significant effort of trying to avoid something that has "error" in the name for non-error cases.
The resulting image is indeed written to the filesystem, right next to the non-scaled one (so when we remove we can remove all of them in one go without having to visit another directory).
The really nasty trick is that when all is done I redirect the browser to the same url and this time I'm sure it will find it.
And if the source file is missing it really does 404, of course.
The only tricky situation is when you have a whole pile of people requesting the same image at the same time and it hasn't been saved yet, the first party to do the 'fopen' gets dibs on doing the transform, everybody else simply redirects if the file exists, by the time the browsers have processed the 3xx the file is ready for consumption.
On very rare occasions the race is so close that the image gets rescaled and saved twice, I've logged the occurence of this over many millions of rescales and it only happened a handful of times, not worth improving on.
This is a pretty nasty trick (seems like a risk of a redirect loop if the file can't be written), but unnecessary. You can emit content out of the script used for ErrorDocument the same as you can with a CGI. So you generate the content, write it to the filesystem, emit a "Status: 200" header, which replaces the 404 status apache was going to generate, then send a content-type and the generated content in the response body.
Your way may be easier if you have complex cache control or expires headers that apache is configured to ad to static files but not dynamically generated content, since the client is only ever given the file by apache directly from the filesystem. I've never used that though.
the twist we use to handle race conditions is to create a lock file. if the 404 handler is called and the lock file exists, someone else is generating that page as we speak -- it goes into a timed loop to check for the file being available. when found, it serves it directly with the 200 header, like you mentioned.
we have a maximum loop value (a second, i think) -- if you hit that, it means the other process died for some reason, and never deleted the lock or the page. in that case, generate and write the page yourself, and remove the lock. all future requests are served off of the file system.
one nice part of this caching mechanism is that to remove the cache, you just rm -rf the blog directory. tada!
That way the browser will see the exact same headers it would on a request where the file did exist.
1) The web server checks to see if the image exists in a file store and, if it does, serves it directly.
2) If not, the request goes to your web application stack (Rails, etc) which fires off an ImageMagick process to do whatever magick the image requests, operating on either a base image or something which you magick up at request time.
3) Your application twiddles its thumbs for a second or two. (This has potentially negative consequences for other requests coming into the same mongrel, at least in Rails. Consider doing load balancing which is aware of the mongrel being tied up -- I think Pound does this fairly easily.)
4) After you've got the image, save it to the file store and tell the web server "Send the user the file at this path".
5) The next request for the same URL will hit the webserver, find the file on disk, and stream it automatically.
In my case I use a cron script every day to blow away all the images that haven't been accessed in the last X days, but depending on your needs it might be easier to just persist them indefinitely. (I use GIFs to give users a live preview of the PDF they're essentially editing, and generate one every second, so if I didn't do this I would drown in half-written bingo cards. My users go through a couple of gigs a week 20kb at a time.)
(Answer: it means today will be a long, long day.)
IMDB's format for resizing images: [server hostname]/images/M/[image identifier]._V1._SX[width]_SY[HEIGHT]_.[format]
eg: http://ia.media-imdb.com/images/M/MV5BNzEwNzQ1NjczM15BMl5Ban...
Unfortunately I didn't have any luck wringing meaning out of their image identification strings...
I also found out rather late in the game that they use an interesting cookie based scheme for defeating people trying to link to their images from external sites...
I'm going to have to read this article in more detail and see if it can clean that up for me and move image serving back to Amazon.
If you are going to use Amazon's images, why not sign up for an associate account and provide links back to the product pages so you can generate some income from it.
Just because the URL describes the operations performed on the image doesn't necessarily mean they always cause those operations to be performed each request.
In addition, I think that any given day a small percentage of Amazon's products are viewed so it may not make sense to generate an image for every possible combination.
They may also have a longer term view and plan on using this graphic manipulation in another product.
It seems pretty smart to me. Rather than having some external process that tries to figure out every permutation of image that might be necessary they just have a simple HTTP API that always gives you exactly the image you need.
As long as they're caching the image after the first request I don't see any problem.