India's SMS GupShup Has 3x The Usage Of Twitter And No Downtime
anand.typepad.com
anand.typepad.com
Downvote me all you like, but I'm embarrassed this story made 21 points.
Moreover, Twitter's load is a joke compared to what leading telecoms have been dealing with for many years. If there were serious about recruiting they could have hired someone with experience to do just that. Without this typical "distributed everything" mantra. Twitter isn't freaking Google or Amazon, they have ZERO computational complexity, they simply move (and store) up to 400 bytes of data for _not even a million_ of active daily users, how hard could it be?
I suspect the reason we see such embarassing technical failures is because software engineering these days is less about engineering and more about playing Lego with pre-existing open source parts, even if they don't fit together or even don't fit at all.
SMS by itself is slow, unreliable, expensive, and only goes point-to-point.
And don't exaggerate for the sake of exaggeration: how many Twitter users have "thousands" of followers?
I used to run a ringtone website and dealt with these sorts of things.
Perhaps SMS is slow and unreliable in the US, but not in europe. Each message has with it a message longetivity which says how long it should stick around in the system trying to deliver itself. Also you have successful send notifications for that rare instance a message could not be delivered.
I don't feel that it's completely different from twitter, then again I could have missed something
UPDATE: I'm wrong - people with only cell phones aren't going to browse SMS update histories often if at all. You're right; though their solution isn't that far off given the existence of Jruby + Rails
I know that the "culture" of Silicon Valley is not to criticize, so let me put it this way: HIRE CAREFULLY for your startup, especially when it comes to engineering. Even if you're solving a trivial issue.
Now, misunderstand me correctly; I don't think this is anything new. My former boss built tools (over 20 years ago) that ran such systems as the NYSE, Swiss Air-traffic control, etc that did fault-tolerant multicast. And the uptime there was somewhat better than Twitter.
What led me to post the original comment was this: They are not solving the same problem, but if they were their architecture would be just as flawed as Twitter's.
Fwiw, here in India, in the initial days of cell phone adoption, some phone cell service providers tried providing plans with "receiving charges", but no one used those plans so they (mostly) faded away.
I assumed something like that happened in the United States as well.
Another interesting phenomenon here is that (most) cell phones are not locked to a particular vendor. The cell phone manufacturers (Nokia , Samsung et al) compete (fiercely) on phones (almost every week a new phone launches) and the service providers (Airtel, Vodaphone et al) compete on service (bazillion plans with different mixes of features. You can shift a plan to another one i a few hours by sending a free sms to the service provider)
You get a sim card from the carrier you subscribe to and buy a phone from wherever you want, insert the card and you are good to go. You can change or upgrade the phone and/or service provider independently of each other.
"Twitter is, fundamentally, a messaging system. Twitter was not architected as a messaging system, however. For expediency's sake, Twitter was built with technologies and practices that are more appropriate to a content management system."
I remember having a discussion with DHH about RoR on IRC a long time ago. Probably somwhere in 2004 when RoR was still very young and not so much known. Coming from a Java/Spring/J2EE background I asked him about abstracting database access in a third tier. He said he had never heard of that and did not know and understand why people were doing that.
Could you explain the advantage of having a middle man between the server and db? Does the middle man caches the requests? From what I can read up it seems too be just a logical layer that executes the request depending on some rules...
So you put something in the middle, that multiplexes those sessions down into say 100 session on the database, checking connections in and out of a pool as necessary, queueing requests asynchronously if there are no free connections in the pool. You avoid the expensive creation/destruction of database sessions, as you start the sessions when the middle tier starts and keep them, and you keep the session state the database has to maintain at an optimal level. Cleverer architectures add effectively another layer between the middle tier and the database to cache query results (because you the developer can know what data you can cache like that, but the database can only make a best guess).
In "sharding" I suppose the middle tier also has to do some logic to figure out which "shard" to direct the query to. Note that this logic must be done, whether you do it yourself in your code, or you let Oracle do it for you in the query optimizer, picking the right partition(s) to actually execute the SQL on.
What I mean when I talk about a 'data-tier' is the code that deals with accessing data without knowing in whatever datastore it is contained.
If you properly hide for example Twitter#getRecentMessagesForUserById(userId) behind a (Java) interface then you can easily change from an implementation where you do direct database calls to a sharded solution or a cached solution.
Other tiers that use this API will simply work because the interface has not changed. One day your app could be talking directly to a single MySQL database and the other day your app could be talking to a 60 node PostgreSQL cluster. It would never know since the details are hidden.
This is totally against what you see in every Rails book or example app where the first thing that is done is direct ActiveRecord queries.
Yes it is more work, but it pays off in the long term. It also greatly reduces local hacks for caching. All that is done in the right place. And automatically for all users of that tier.
S.
Specifically, I would need to change my database.yml file with my new database info and run a migration.
Did I miss something?
Although it also depends how many destinations each of those messages has.
Twitter's real problem is that they've had a pittance of hardware until very recently. It's not my place to go into why this is (and I'm sure the version of the story I know is a bit biased by the teller), but suffice it to say that a lot of Twitter's problems were, until recently, more business-oriented that technology oriented.
One one hand, they know their architecture isn't perfect and will be rebuilding it to cope with more users.
But on the other hand - they are a small team which is trying to keep it up in its current state, and cant just decide to close Twitter for a few weeks/months while they work on the new architecture, people will defiantly move to another service, and Twitter is dead.
Probably the fact that those drives are local while sharded databases are physically seperate (clusters) of independent machines?
Some apps totally suck at this model. Others are perfectly suited. I doubt Twitter needs complex joins, so it should be easy to partition their data.
S.
The same model as the far east with cheap manufacturing of electronics goods in the previous decades. After a while of outsourced cheap production/manufacturing, these players learnt about the markets/products and how to actually design and build themselves and now many of them are major manufacturers and producers in thier own right.
Hence instead of working for $10k~20k (USD) on average in a multinational/call-centre, they will form tech/consulting startups. This is already been demonstrated in the macro-scale with names like InfoSys/TCS/MindTree/HCL and now will happen in the micro-scale too, especially after media coverage of people like the Scrabulous founders etc (who many media sources claim earn about $20k~$30k a month in just facebook advertising).
Only a matter of time till we see a web superhit from India, imho.