How to use Amazon’s S3 web service for Scaling Image Hosting
blog.teachstreet.com
blog.teachstreet.com
Perhaps a clarification, but Amazon does not guarantee your data as part of S3. They will do their best job (and have a great track record), but just like any other service -- failures will occur. Ultimately their SLA may provide for reimbursement of downtime, but that won't get your data back.
That being said, S3 is still (probably) magnitudes better than what you can homegrow.
(From http://www.allthingsdistributed.com/2010/05/amazon_s3_reduce... )
EBS volumes have a theoretical "annual failure rate (AFR) of between 0.1% – 0.5%, where failure refers to a complete loss of the volume."
(From http://aws.amazon.com/ebs/ )
Considering how many EBS volumes there are, that is not that low of a percentage for us to see many cases. It happens and the case at that link is not the only one.
But that person didn't take the advice, he did not snapshot his EBS volumes to S3. It says it right there in the description (as well as in the user's guide): "The durability of your volume depends both on the size of your volume and the percentage of the data that has changed since your last snapshot."
Further, there is not protection against your account becoming errant in Amazon's eyes resulting in you getting locked out. And d2viant's point remains, Amazon could mess up or be subject to something out of their control.
I'm not trying to diminish the excellent service and good track record they have, I use it myself as a piece of my total backup solution. But just a piece.
It's the equivalent of not getting insurance. Yes, it saves you money, but on the other hand if you're really unlucky the consequences are truly dire.
In case of an epic S3 disaster, we could backup our S3 buckets to another offsite storage, but we're not going to improve much further than what their giving us.
That way deploying more photo servers for speed just requires nginx and a few lines of config.
A couple notes:
1. Make sure you partition your data on your local disk properly so that there are never more than ~5000 files per directory, you can use nginx rewrite rules to do this for you (and make nginx save the actual url to nested folders).
2. If possible, make thumbnails smaller than 4kb.
2. Good tip - we try to use our judgement here balancing good design/good performance. We tend to opt for a pretty website first, then go back and make it faster with optimization.
location ~ "/(.*)" {
root /whatever/$1;
You could even do: set $subdir $1;
so that it can be used in another location handler.You've done 99% of the work to get your images serving quickly to your users, but then you stopped. Why??? Cloudfront takes exactly four minutes to set up, and it's a full-blown CDN.
Nobody, not even Amazon, recommends serving content directly from S3 anymore. Either this article is 2 years out of date or the author doesn't understand his subject as well as he thinks.
And speaking of cheap, Cloudfront is as close to free as S3 itself. It will roughly double your S3 bill, leaving you paying a grand total of $8.04 per month for your image hosting on a ~1,000,000 unique/month site. Remember the 90's when that would run you $400/month? It's so cheap that it's not worth calculating. Just flip on Cloudfront, configure your CNAME and get on with your life.
The primary advantage here is that you don't have to coordinate generation of new images just because someone, somewhere requires a specific size in some client of your image system.
Of course, just stick a CDN in front of your on-demand scaling implementation, potentially backed by S3, and you're done.
Cloudfront is really enticing, though. For a cheap CDN, if we really had the need, we'd probably migrate to cloudfront, then come up with a task to batch resize our S3 images, then stick them back into S3. This could probably be done on a one-off task on ec2 by spinning up a few instances, or even with Hadoop.
If anyone's taken this approach, I'd love to see it!
Right now I'm working on doing streaming intermediate resizes that never load the full decompressed image into memory, because I need to be able to take uploads of 50 megapixel images and be able to respond with new versions based on user input in a timely manner. RMagick may be an extremely shitty implementation, but even a more solid one is stymied by the algorithm.
The solution is to just do something else -- for JPEGs, libjpeg has the capability to do a cheap streaming resize to 1/2, 1/4, or 1/8 by partially sampling the 8x8 DCT blocks and for PNGs MediaWiki's image server includes a utility for doing streaming resizes, though it's finicky about the exact scale factors it'll allow: http://svn.wikimedia.org/viewvc/mediawiki/trunk/pngds/
We use a plugin I wrote (http://github.com/elektronaut/dynamic_image/), which seems to be based on similar ideas on syntax. For caching we just do regular page caching, with no expiration. Since the request doesn't have to hit the backend at all (except for the first view), this is pretty damn fast.
Have people written different datastores (like S3?). I love the idea of a modular, pluggable image server that can choose different backends based on business need/scaling.
Luckily, we deal with pretty small pictures since they are just pictures of our users & their classes. We don't really have any need for handling larger sizes, and our biggest use cases surround generating thumbnails (similar to many social networking web apps).
Based on your steps listed here 1.Handle request 2.Fetch original source image from S3 3.Resize/apply effects 4. Return result back to user. Why not just resize/apply at time of upload and store in S3... are you resizing and applying effects on every request for an image ?
"For performance, each of our image servers cache the source, and any resize, locally to disk. Since images are never updated (only created), and get a unique ID for each one, we don’t have to worry about cache invalidation, only expiration. We can then write a simple script to remove images from this disk cache with files of an access time greater than a certain threshold (say 30 days). That way, if we change from one size thumbnail to another, eventually the old thumbnail sizes will get purged."
The reason we resize and apply effects at request time instead of at upload time is that our design needs may change over time, and this way we don't need to go back and batch reprocess things.
A nice pain free, embed the uploader and set some settings and then the 3rd party processes the images and then uploads them to your S3 and pings you. With the ability to reprocess old images to a new size and integrate with attachment_fu or paperclip (or has it's own similar plugin)
For use cases where you expect only to need a fraction of the possible image "shapes" for any given asset but can't easily predict it ahead of time, on-the-fly generation is very attractive.
I've been back-of-the-mind sketching my ideal service along these lines for a year or more -- if my imaginary ideal existed, it would save me a ridiculous amount of grief.
edit: Sorry just found out they only do resizing/rescaling, you still have to take care of the initial upload yourself.
It would be really cool to have an EC2 instance handle these resizes and put the images in new buckets for you - Might be a really neat implementation.