Facebook's software architecture
muratbuffalo.blogspot.com
muratbuffalo.blogspot.com
It looks like most of the savings here is due to optimization of the replication level of storage, according to the usage patterns of the data.
Also found it interesting that it doesn't talk about CDN usage.
It can be more/less than this depending on your rack configuration and how big drives you use, and how dense you pack them (e.g. "standard" off the shelf servers typically peak at 24 3.5" drives per 4U, while top-loaded specialised storage enclosures reach about 50 x 3.5" drives per 4U; full racks vary between 42U and around 47U etc.)
But you could if you wanted to get this much storage into about a 5m x 5m server room....
Of course you wouldn't, since you'd want multi-datacentre redundancy etc. Often there's also no point using the biggest drives available because you need higher throughput. Or no point packing them as tight as possible because you need more processing power per drive.
(I work on the team that builds and runs the Facebook CDN infrastructure.)
Increasing the replication factor would allow for faster reads, but there usually wouldn't be a need for that if you offload reads of hot data to a CDN service (could be an internal CDN...).
Hope that makes sense - let me know if not.
[1] http://www.cs.cornell.edu/~qhuang/papers/sosp_fbanalysis.pdf
[2] https://www.usenix.org/legacy/event/osdi10/tech/full_papers/...
[3] http://muratbuffalo.blogspot.com/2010/12/finding-needle-in-h...
I think it is a very logical solution. Naturally when I think of facebook, its the freshness of the data that is important for me, as a user I like to know whats going on now, as opposed to older timeframes. I think it seems quite logical that it is much more optimal to manage their data this way.
I read High Scalability and the occasional company blog. Is there anything else out there that might be even better? Blogs, books, forums, doesn't matter.
You're also putting the burden on the folks in charge to know what's out there and pass it down to you, instead of actually doing your own research.
Also, sometimes those people don't know what they're doing and they're struggling just as much as you would be in their position, so now you're learning from people who are making it up as they go, which you could have done equally as well by yourself.
Wow TIL what Tao is and what it can do.
This part is very sexy to me: Facebook's new architecture splits the media into two categories: 1) hot/recently-added media, which is still stored in Haystack, and 2) warm media (still not cold), which is now stored in F4 storage and not in Haystack.
http://muratbuffalo.blogspot.com/2012/10/building-fault-tole...
http://muratbuffalo.blogspot.com/2013/04/aws-summit-nyc-day-...
http://muratbuffalo.blogspot.com/2013/04/aws-summit-nyc-rest...
http://muratbuffalo.blogspot.com/2013/04/aws-summit-nyc-day-...
System architecture on the other hand describes the different applications/services that compromise a system, including the ones that you write yourself, plus the ones that you just "use", e.g. database, storage, load balancer, etc.
I would argue that the article mostly talks about system architecture with focus on storage and I guess the parent you replied to would too.