An analysis of Facebook photo caching
code.facebook.com
code.facebook.com
In other words, I wonder how much of this efficiency boost is due to FB's abilities (both in people and technology) to scale. The paper seems to imply that it's relatively simple, now that the data has been gathered, but for a one-person team like mine, I wonder what benefits I can take away from this.
There is certainly a lot of infrastructure that this is built on that many small teams don't immediately have access to in other companies. Whether it's self-service hardware provisioning, Scribe logging infrastructure, tailer frameworks with checkpointing and retries (and job systems to schedule them), and large amounts of available space on Hive for experimentation. But most of the software parts are available as open source, so it doesn't need remain unavailable.
This is my first team at Facebook that I've been heavily involved in this scale of data capture and analysis, but it only took a few days to get up to speed through a combination of great tools and good documentation. Being able to drop a Python file in a code repo to ensure that some complex data warehousing task takes place every day after that is pretty powerful.
I'm not sure whether a "refresh" (ie, 304 not modified due to expiration) is counted as a miss in this data.
And then you get fired for looking at private user photos.
https://www.youtube.com/watch?v=ENaQScyvOzY&list=PLn0nrSd4xj...