Why Can't Twitter Scale? Blaine Cook Tries To Explain
alleyinsider.com
alleyinsider.com
http://romeda.org/blog/2008/05/scalability.html#140141155247...
"Scaling Twitter as a messaging platform is pretty easy. See Mickaël Rémond's post on the subject. Scaling the archival, and massive infrastructure concerns (think billions of authenticated polling requests per month) are not, no matter what platform you're on. Particularly when you need to take complex privacy concerns into account."
Sounds like they are having the kinds of problems that Friendster had years back.
How come sites like MySpace don't have these issues? They also seem to have a pretty complex social graph.
http://www.baselinemag.com/c/a/Projects-Networks-and-Storage...
Are they handling more than 64GB of user-generated data per hour? If not, why not just store everything into RAM on a big 128GB RAM server and query that?
The fact that they chose to use RoR anyway hints that they may not be one hundred percent technically competent.
Scalability is a second-order issue: if you go from 1 server to 2, but they need to go from 100 servers to 400 servers to handle the same load, then not only are they slower than you are, but they can't scale as well as you can.
I get why using a framework for the UI and business logic could make them 100x slower than something custom-coded. And from the original post, I get why a conventional database+read cache may not be appropriate for a messaging application.
But what I don't get is the connection between RoR and scalability. Unless you are speaking of its default configuration, namely RoR+ActiveRecord+MySQL. Which speaks more to the architecture choice (tables, rows) than to the framework choice (views, models, controllers).
or am I missing something?????
Thanks for explaining your reasoning.
I believe someone has made the defensible claim that most of their activity occurs via clients (twhirl, alert thingy) and SMS over the API rather than through the website. Unless they are using ActionController and ActionView to do the rendering and output there, it seems they would be using "traditional RoR" in a very limited capacity.
I'm not being an apologist, but I think if people are discussing Twitter and scaling for the purposes of learning something useful and not just framework bashing, language bashing or avoiding doing homework, we should probably avoiding talking as if they use the full stack in a significant way as it's a red herring.
Someone else in the comments mentioned 400K messages a day, plus personal messages, say thats even another 400K.
I have zero experience working on a high traffic web application, but I work on a serious database application were we can push somewhere in the region of 1M transactions in a 6-8 hour window. Each of these transactions comes in the form of an XML string and results in as many as 20 plus selects and maybe 10 - 20 inserts.
Our application is build on top of Oracle and implemented in PLSQL - not too sexy, but it seems to get the job done.
The point I am trying to make is that with a bit of a caching layer it takes some serious throughput to reach the limits of database scalability (certainly with Oracle, where is where my experience lies).
It seems like a lot of these start-ups are the founders' first tries at something big, so they make a lot of mistakes; but since it's new and exciting to them, motivation wins out and they actually accomplish something.
I really admire their enthusiasm and stick-to-it-iveness.
That'll be the Second-system effect ( http://en.wikipedia.org/wiki/Second-system_effect ).
If you've previously had a manager who vetoed your doohickies then there's a risk that some get implemented by your startup. Any of them could be a scalability nightmare.
I don't know how much the private messages add to this, but it can't be an order of magnitude higher.
Perhaps they are continuously doing DB writes instead of keeping a write cache and storing content in batches, who knows.
If they are using rails (and certainly they started with it) that is almost certainly what they are/were doing. Thats the Rails Way. Store everything, load it up, remember nothing at all.
So now that metadata caching is out of scope here, we turn to frameworks, architectures, and application design. The issue Twitter has is in aggregation of their meta data. People who propose a distributed solution for Twitter obviously miss the inherent nature of Twitter; it's a centralized system. Twitter should remain as it is, but ditch RoR and become more modularized. This is where they start developing real systems; the type of systems they talk about in those mundane CS classes, like programming in C -- that stuff. Modularize the applicaion, develop systems that scream for aggregation, cache the hell out of everything, and start applying some computer science.
Twitter was born as one of those next-gen Web 2.0 "keep track of your friends" hip Ruby on Rails insertbuzzwordhere application. Now it poses actual architectural challenges. It's so similar to the evolution of Facebook. Just take a step back and look at it.
I really hope my initial post wasn't interpreted the way it was to validate that response. I don't take back anything I said--and I'm sure it's not what I want to hear. Architectures scale, not languages... Twitter simply cannot be distributed... it's not news, I've written about it earlier.
I think caching can work, but at a different level of granularity. Rather than cache a person's full timeline, which is composed of multiple sub-feeds (each of which requires a database query), cache the data from the sub-feeds themselves, then recombine them on every page load. This would significantly lower the number of database queries, as each cache element would be invalidated only when its "owner" sends a tweet. This solution would be much more CPU intensive on the application servers, though, and Ruby may not be the best tool for the job if that were the case.
On the other hand, there's gotta be something fishy going on in Twitter-ville. Although I agree "the idea that building a large scale web application is trivial or a solved problem is simply ridiculous," by now Twitter should have enough performance data to know exactly which part of the process is causing the high-load issues they're having. If we are to assume that after each outage they at least "throw more hardware at it," then, theoretically, it's a problem that horizontal scaling cannot solve and the issue is deep-rooted in the system -- somewhere in the basic architecture of it.
Is it Ruby? RoR? Poorly optimized queries? Improper caching? Lack of domain knowledge? Leprechauns? I don't know, neither does anyone else outside of Twitter... but I guess speculation can be entertaining.
That's it. Now, sending a message takes O(n) (n=followers) time, which is really cheap. On my machine, it takes about a second to create and sync 40,000 files (there's not much data, so replicating this via NFS wouldn't be that expensive either). With that out of the way, all you have to do is ls your "twitter directory" to see all of your friend's messages. This is another incredibly cheap operation. It's easy to distribute, and there's no locking.
Anyway, just look at the mail handling systems at huge universities and corporations. They scale fine, and they're much more complicated than twitter. Twitter is just a subset of e-mail, so it should be implement that way, not as a "SELECT * FROM tweets WHERE user IN (list, of, followers) ORDER BY date". That is the wrong approach because it makes reads (very common) expensive and writes (very uncommon) cheap. That's why twitter doesn't scale.
API conections, given their nature of having to do a security check each time should be based on a slave copy read-only db which is near-realtime, or potentially dirty read but who cares, its for their API. The security lookup should be cached since you rarely change API security accounts more than once..if you do, then you update the caches.. etc.
Granted I'm no expert, but it just sounds like they're overloading their DB with polling.
Their web pages for each user should be cached since they're not really hit as much as the api or rss
I dont' care what anyone says, having THAT many API connections constantly polling their DB along with what blane said about having to do authentication requests on every singel hit is taxing on any setup you can put together.
Oh and if they have a single table holding all their tweets...big issues there. I'd have a 26 "tweet" tables, one for each letter of the alphabet and my switch in the business layer. Then simply ship tweets older than a month, or patition based on that.
Responding to reasonable complaints by using a different definition of word "scale," makes for a weak argument.
It's all a question of what you believe to be the scarce resource. H=e obviously believes that for a well-founded company, server CPU cycles are not a scarce resource. This implies that he believes that RoR addresses some issue raised by something else being the scarce resource.
You can get 10 servers bought, delivered, installed and configured in one week. These servers can be deployed to do anything from load balancing, web caching, database caching, being database read slaves or application servers. Regardless of the quality of your architecture, 10 servers will probably add some scalability to your system. Furthermore, they can be re-purposed as the situation demands.
Hiring a better programmer (or just another warm body) takes longer. Furthermore, the improved code they produce isn't as flexible as surplus hardware.
Of course, the situation changes when you've got thousands of servers.
That's just one idea. I'm not really an expert.
is there any publicly available documentation for twitter's architecture?
did they use consultant help, did they contact SUN, IBM, Oracle or any other respected consultant when they started facing those problems?
i recommend you watch this video: http://www.infoq.com/presentations/qcon-voca-architecture-sp...
really ... we need to have a look at twitter architecture before we discuss this further
Are you implying that Sun, IBM, and Oracle are respected for the prowess of their consulting organizations? I'm sure they have some Fine People working there, however the view that "writing a seven-figure cheque to these companies guarantees a positive outcome" is not universally held.
http://teddziuba.com/2008/04/im-going-to-scale-my-foot-up-y....
In the good old days, every message over ICQ was sent over their servers, and they probably had more messages to handle than twitter does nowadays. And there wasn't a single problem then.
Not thinking long about it (5 minutes), but this is how I would do it: Not going muhc into detail (I don't use twitter, just had a quick look at it).
1) 1 replicated system where you can fetch messages by id (complete body with who sent it, to whom, time, etc...) 2) User page: list of ids.... 3) Private messages: list of ids... 4) When a new message comes in, write it to a queue. Process those ids and append the ids to the different pages. Multiple processes doing this (each process has a subset of users with the complete list of people following those users). One could add another layer to do bulk inserts.
Could be easily done in Memcachedb. One page view takes x + 1 memcachedb requests (x number of items on page). One can still optimze this by caching (static html pages which are deleted when a page is updated for a user). When inserting, replace existing data by adding the ids.
Everything is nicely seperated. (Eg pages for user 1-10000 are on server 1, etc... Messages can be nicely sperated as well).
Any thoughts on this? To twitter: hire me not him ;)