Tumblr Architecture - 15B Page Views a Month and Harder to Scale than Twitter
highscalability.com
highscalability.com
Holy schnikeys! The graph changes are two orders of magnitude larger than the content additions.
Somebody from Tumblr (though on a different team) was nice enough to respond and guessed that it was because all media on S3 was not counted.
https://twitter.com/#!/adamlaiacano/status/16914079173274009...
http://amule.org (or http://mldonkey.sourceforge.net/Main_Page if you're into functional programming)
http://www.torproject.org (especially the hidden services stuff)
And you probably want to read this: http://pmg.csail.mit.edu/papers/osdi99.pdf
In addition to solving the problem space using simple, focused applications (in the classical Unix style), it solves other practical issues. Like PHP having shitty clients for most modern data stores (Cassandra, Redis, hell, even Memcached) compared to what you get in, say, Python, or a JVM language.
[1]: http://mongrel2.org/
http://gearman.org/ http://gearman.org/index.php#how_does_gearman_work
"Don’t hire people based on their survival through a useless technological gauntlet. Hire them because they fit your team and can do the job."
I think so!
Perhaps I need to use hair coloring.
I don't have C# on my resume. I wouldn't shrink from working on a C# codebase though. Would I get an interview though? Unlikely.
Apache, PHP, Scala, Ruby, Redis, HBase, MySQL, Varnish, HA-Proxy, nginx, Memcache, Gearman, Kafka, Kestrel, Finagle Thrift, HTTP Func, Git, Capistrano, Puppet, Jenkins
I dunno, it seems they might do well to hire based on someone's ability to survive through a technological gauntlet. But the word useless makes the statement kind of....useless.
The Akka folks do a good job of keep the actor model on your mind when building Scala systems, but it's obviously not always the right approach. I'd love to hear more about how they ended up abandoning the Actor model.
$ curl -i 'http://assets.tumblr.com/images/favicon.gif?2 HTTP/1.1 200 OK Server: nginx/0.8.53 Content-Type: image/gif Last-Modified: Fri, 15 Apr 2011 22:13:30 GMT Accept-Ranges: bytes Content-Length: 635 X-Varnish: 1795992965 Cache-Control: max-age=2196224 Expires: Sat, 10 Mar 2012 05:35:30 GMT Date: Mon, 13 Feb 2012 19:31:46 GMT Connection: keep-alive
So it looks like their application is being served by Apache (and PHP), while their static assets are served by nginx behind Varnish.
$ host assets.tumblr.com
assets.tumblr.com is an alias for
assets.tumblr.com.edgesuite.net.
assets.tumblr.com.edgesuite.net is an alias for
a1092.g.akamai.net.
a1092.g.akamai.net has address 69.31.106.32
a1092.g.akamai.net has address 69.31.106.50Get in touch and I'll add you to the list.
From the article: ... • MySQL (plus sharding) scales, apps don't. • Redis is amazing. • Scala apps perform fantastically. ...
Some of the organizations (Including Tumblr) using Scala were discussed as part of this talk: http://www.youtube.com/watch?v=qqQNqIy5LdM and slides here: http://mrkn.co/s/video_martin_odersky_what_s_next_for_scala,...
Is that anything to do with my demographic (41, M, UK based)?
"dashboard" is their admin area?
So in theory 700 of their 1000 servers are for people to make a new post?
If I remember correctly, wp.com has a similar problem.
Tumblr is not a blogging platform like Wordpress, it's much more a community like Reddit. Yes, you can use Tumblr to host your blog, but the majority of people use the dashboard to interact with others and use their blog to share their content. Very few people will ever see username.tumblr.com, they'll see the posts via tumblr.com/dashboard. Like Twitter, very few people visit twitter.com/citricsquid, they follow me and see my tweets in their stream.
For example, my dashboard right now: http://i.imgur.com/0YAYv.jpg (potentially nsfw material, didn't check)
Lots of read-only content, and then a decent amount of interactions and events (messages/new follows/like notifications).
It's going to be a little fragmented, but nonetheless, I found this a hugely interesting article.
I use Google Reader for two main reasons. One is that once I subscribe to a feed, I get access to all messages, not just new ones. The second is that I can search old messages as well.
So if I understand Tumblrs inbox model correctly, that's exactly the kind of usage pattern that isn't supported.
Most blogs have disqus comments on each post; every page-view on tumblr is at least one fetch on disqus; many more for index pages.
Very clever strategy for tumblr to make comments someone else's problem ;)
Interesting.
Anyone notice this ??? logical?
Well, luckily for them, that's not actually the case.
Yes, I was being extremely sarcastic.
1) there are far fewwer than 175m active users of twitter 2) a tweet compresses to far less than 140 bytes. Text normally compresses to (easily) a third its size 3) Wikipedia says there are some 300 million tweets per day...or 13 gigabytes per day (compressed).
A single 1 TB hard-drive is almost certainly good enough to store a couple of months of tweets. A year of Tweets should fit on a couple of thousand dollasr of hardware.
and 4) I added the Amazon line so that you guys would know I was being absolutely, 100% sarcastic in every way possible. I guess the clue still wasn't enough.
My point is that there are very few things that are as easy to scale as a platform built on sharing 0.13 kilobyte tweets (0.04 KB compressed), in an age where rendering many common front pages requires a browser to download 600-800KB in static and dynamic content. Come on.
600-800KB +++ I was chatting with my brother in law recently about his website (fashion photographer). His homepage was about 5.6MB. When I asked him about it -- he figured it was normal. Sure enough -- he rattled off a few names for me to check. All of them had front-pages >6MB.
They're still phenomenal numbers, but IMHO should be much closer to a 100-server environment than a 1000 server one.
At the startup I work, we've got 25-30 Million users, many stats similar to Tumblr, and we're running it on about 250 Ec2 instances of varying size. I think if Tumblr's numbers are high at all -- due to rapid iterations and no time to focus on deep optimizations -- it's maybe 10-15% high, not 90%.
I'm saying this because I've seen periods where our usage numbers are somewhat flat, even falling, but our hardware demands rise as we provide more features. When there's just one or two primary ways to use a service (eg "I post status updates and comment on my friends' status updates") it can be quite easy to optimize. But add features. Photos. Chat. An in-house ad serving platform. I18N. Etc. You have different types of interactions with different acceptable service levels and varying storage requirements.
A request for static content != a request for dynamic content != a request for inter-user messaging != a request for a recommendation engine.
Its high time everypost about insane pageviews and servers be coupled with solid revenue numbers.