Absolutely! You could probably get a single rack to handle storage and broadcast of millions of tweets per minute, perhaps hook it up to a Teradata appliance to handle search and storage at that throughput. Throw it in a nice distributed database indexed by user. Then just another couple of indices for the hashtags and mentions, and the follows/follower graph. Throw in another few tables for handling API permissions, managing schema migrations while ensuring that data from day 1 is still as transparent as yesterday's logs, a highly replicated ad server hooked up to a realtime recommendation engine that analyzes an active user's current stream and recommends the most relevant ad (to maximize revenue of course), set up a few staffed policy teams to comply with the differences in global regulations, handle lawsuits and legal requests, and of course keep competing for more revenue with the likes of Facebook and Google.
It's pretty trivial.
When you have several hundred million users, that adds up. An effective monetization model, if unusual in SV.
Are you suggesting that WhatsApp's model of organization is the norm for a sustainable organization? I think it's a clear outlier in many respects.
Are you suggesting that WhatsApps business model was based on annual revenue?
Silicon Valley mentality at its finest.
The company existed to make a return for the owners. The owners exited with $16 billion. That was the motivation of the owners/employees, not $14 million in annual revenue.
My point is that the transaction must be analyzed from the viewpoint of the owners (as that is what the poster I was responding to cared about, viz a viz total profit/revenue to Twitter). When you do that, WhatsApp cashed (well, assuming FB stock can be treated as cash) out more per employee than Twitter.
Since WhatsApp was a revenue-generation corporation with a board of directors, like any other corporation, before their acquisition, and since I don't like to indulge in conspiracy theories, I'm going to assume that yes -- their business model was based on annual revenue.
Then you'd assume wrong. The goal of the board is to maximize return for investors. In WhatsApp case, it's not hard to see that their goal was to grow as fast as possible, revenue be damned, so they could be a very attractive takeover target for a company like Facebook who could eventually hope to monetize that user base.
They're nowhere near relevant. I keep getting local-to-US ads, and I'm in UK (and geolocate as in UK too). Not as bad as showing alcohol ads to 5 year old children, but still pretty bad.
This doesn't mean that there aren't teams of people trying to constantly make those ads more relevant. And since not everyone at a Twitter can statistically be incompetent, I'm going to assume that the problem is actually quite difficult, and perhaps requires a large team.
WhatsApp had ~30 when they were acquired, and were also huge. ("Huuuuuge!", says The Trump)
Those teams are gods of scalability. Twitter is.... kinda the opposite.
Edit: My apologies if I've offended the HN mystique surrounding the god-like status of WhatsApp. Ask Jan yourself, they weren't geniuses. I know Jan. He won't claim to be this exalted persona.
Edit: I know b/c I've done that before in $previousjob. Storage is hard. Storage is where the money goes. Durability is expensive. Storing ephemeral stuff in a non-durable KV store is easy, cheap, and trivial to scale. This is why WhatsApp didn't have so many people, not because they were geniuses.
Disks are slow. If you don't need to wait for disk, then you can handle a great deal more req/s per node, which means the point at which you need to become a distributed system is further away. Without the state synchronization problem, converting to a distributed system consists of little more than adding a load balancer.
Even in an environment that needs synchronized state, the services that can be stateless are much easier to deal with (and you try to make as much of your infrastructure stateless as possible).
Storage costs are not relevant compared to the costs of engineering distributed systems.
Deciding which ad is the most optimal to show benefits from every input (metadata?) you can make available and analyze.
That must require some multiple of effort vs. making sure 140 chars get shown where they're directed to go.
This isn't a knock on WhatsApp -- I use WhatsApp and don't use Twitter -- but Twitter's scale challenges are vastly more complex than WhatsApp's, which doesn't have many features that can't just very simply horizontally scale.
Back of the envelope calculation: If your messages are 256 char messages, on a 128GB ram server you can fit 500 billion of them in memory; incidentally, that's one day's worth of tweets these days. I recently got a quote for a beefy 256GB ram server from Dell for $6,000. To make things easier, let's assume $10,000 with redundant 1TB SSD and everything else you might need. The data is extremely easy to shard.
With proper caching, you can probably keep 99.9999% of all tweets shown in just one such servers, and serve them within a few ms. I don't have recent experience with web frontends, but a properly configured nginx could probably server on meager hardware can do a few tens-of-thousands requests per second. a Seastar server on beefy hardware can probably handle a significant percentage of Twitter's load.
Search is a little harder, but not by much.
Is this engineering trivial? Far from it. But I've worked with stunt programmers who can build this in a few months with a 3 person team. At least one of them commands a salary of about $500K last I heard.
Is that all Twitter is doing? No, I'm sure they have 10 backend "things" for every user visible "thing". They needed something like Bootstrap for internal use, after all. I don't really believe you can run Twitter with just $1M/year in capex, just $1M/year in opex and just $1M/year in engineering salary - you need a whole lot more than that; In fact, to pull in revenue of $2B/year (which Twitter seems to be on track for), you probably need 50 people just in accounting.
What I am saying, though, is that the technical challenges faced by Twitter (at least those that are visible to users) are highly overrated. If you know what you are doing, Twitter can be run on $50,000 worth of hardware; This has been true since 2007 - Computers were weaker and had less memory then, but Twitter had much less traffic too. I've looked at the numbers every single year since 2007, and at any point in time, the "bare bones essential hardware for interactive response" was in the $50,000 range. Which probably means if they spent more than $1M/year on hardware (or depreciation-equivalent-hosting) at any point in time, they should have hired better engineers.
You will also need space for fulltext indexes, indexes per hashtag (a set which grows constantly and never shrinks, but has a power distribution, meaning that some hot tags will overwhelm any single machine and it constantly changes), the ability to call up any tweet from any point in history (growing by one beefy Dell server per day), the ability to fan out to 1 or 10 or 10 million followers and so on. The fanout problem turns sharding into a lumpy problem because again, there is a power law distribution that can overwhelm any single machine.
Then there's the fun part that Twitter is so reliable these days that when something else on the internet breaks, we check Twitter for updates. Robustness at that scale is more than buying a pair of SSDs and setting up RAID.
One thing I've learnt from being a consulting engineer is to stop assuming other engineers are morons. Sometimes we do things out of ignorance, or out of fear, or out of necessity. But often, when you inquire, there is a solid non-obvious reason for non-trivial diversions from the simplest thing that could possibly work.
I always find it useful to give the benefit of the doubt.
But .. fulltext index done right, including per hashtag, doubles the storage requirements at most. And the ability to call up any tweet from history does not need to be as instantaneous is recent tweets (It's ok to wait 5 seconds for a tweet from 5 years ago).
> The fanout problem turns sharding into a lumpy problem because again, there is a power law distribution that can overwhelm any single machine.
Not really, in the case of twitter (where you only trace edges, not edges-of-edges and edges-of-edges-of-edges like Facebook and Linkedin do). In the simplified model I described, the web frontends all pull from the backend, but for the hot 10-million account followers -- of which there aren't that many --, you can just push them to the webservers so they don't even need to query)
> Robustness at that scale is more than buying a pair of SSDs and setting up RAID.
I most definitely agree. But it is cheaply doable when everything is so perfectly shardable as it is in Twitter.
> Then there's the fun part that Twitter is so reliable these days that when something else on the internet breaks, we check Twitter for updates.
That's possibly an illusion. Yes, Twitter is mostly reliable - but do you have any latency stats? e.g., would you know if, on a daily bases, 10% of tweets take 60 seconds until they appear on a viewer's refresh? The "hot" accounts would be cached and immediately updated, but the long tail might have minutes delay and almost no one would notice.
> But often, when you inquire, there is a solid non-obvious reason for non-trivial diversions from the simplest thing that could possibly work.
> I always find it useful to give the benefit of the doubt.
I do not assume other engineers are morons. In my experience as a consultant, though, those solid non-obvious reasons for non-trivial diversions are much more often than not "historical, do not apply any more", "we didn't have time or resources to do it the right way", "the guy who did the initial design did it wrong for whatever reason, and now we are stuck with it", "there's a legal reason that's not obvious why we can't do it that way", and a few others.
Twitter had the resources to do it better from very early on, and they didn't (I think it was 2012 or so before Twitter became reliable). For all I know, their system now could be the most efficient beast ever, and run on a ZX81 with 1K Ram with ultimate reliability, way better than I could ever hope to build.
I was just pointing out that the user facing side of twitter, technically speaking, is not very impressive. I've been doing it for a while on various technical forums - and not once did anyone offer a reason for why it's much harder than it would seem -- most of the responses were along the lines of "but mysql/pgsql/oracle can't take the firehose load, I tried!". Which is correct, but irrelevant.
By saying it can be done with $12,000 of hardware, you essentially claim it is trivial.
> everything is so perfectly shardable as it is in Twitter.
Shardability, or partitionability, depends on the underlying distribution. It's only "perfectly" partitionable if that distribution is perfectly uniform.
Twitter's traffic patterns are power-laws. Some tags vastly outnumber others. Some users have vastly more followers. Those distributions are unstable across time and space in surprisingly short order.
> Twitter had the resources to do it better from very early on, and they didn't (I think it was 2012 or so before Twitter became reliable).
Twitter wasn't interested in a contest to use the least hardware. They were in a race to keep up with surging load, while migrating from Rails to something that didn't exist yet. Meanwhile they have, under the hood, built infrastructure that could not be bought off the shelf from anyone at the time.
We can't actually come to a conclusion here, because you're arguing a counterfactual. It's always easy to have the ideal solution to a problem you never actually solved yourself, because the visible features of a problem are obvious and the many, many invisible problems are where the bulk of the effort are hiding.
If you want to prove that Twitter doesn't need such large and complex distributed systems, you can probably round up a few hundred thousand in angel funding and sell yourself to Twitter for a nice return.
Not at all. What I'm saying is that, with the proper software (which is not trivial to write), you can do it with very little hardware. I know stunt programmers making $500K/year, and they are in some senses infinitely (not just x10) more productive than bad and even average programmers - because they quickly produce working systems that others just can't.
Whether it makes sense to pay $50K/hardware and $500K/programmer or $2M/hardware and $100K/programmers depends on how you run your business, though - and in many cases, the $2M/hardware+$100K/programmers are the more economical choice (because you risk starving for stunt programmers choosing the first)
> Some users have vastly more followers. Those distributions are unstable across time and space in surprisingly short order.
And yet, as a system designer you actually get to choose what the distribution is of - and a good choice makes it uniform. The power laws may or may not favor sharding on uid specifically, but usually there's a simple way to shard uniformly. And if you can't find a way to programmatically shard uniformly, then shard using a lookup table on the userid - 8 billion user ids require all of 8GB of ram if you have 256 shards or less (and if you bundle users in groups of 256, 32MB is suddenly enough). Migrate around to keep balanced. This has been a solved problem for years. Really, even migration patterns. Look at "consistent hashing" literature -- (it's not the fundamental issue that consistent hashing solves, but the peripheral solutions are well known, common, and apply here: migration, redistribution, redirection).
> Twitter wasn't interested in a contest to use the least hardware.
> We can't actually come to a conclusion here, because you're arguing a counterfactual. It's always easy to have the ideal solution to a problem you never actually solved yourself,
First, I basically agree with you, if it wasn't clear. I know not what problems twitter were facing. I suspect that they weren't technical in nature, though -- because the technical issues have been solved before them. It might be management issues, it might be ego issues. I've seen more projects fail or stumble on those than on technical issues.
I actually did solve those problems myself, on a smaller scale (hence my interest, but with perfect sharding that would have scaled to any size, live migration and redistribution and all). But that startup folded because we sucked at getting traction - which only shows to go you that technical prowess matters not in these issues, or at least not much.
> many invisible problems are where the bulk of the effort are hiding.
Again, to be clear - I totally agree with you. I'm just disagreeing with the aura of engineering excellence that Twitter gets in this (and many other threads). They may have it, or may not - I haven't seen evidence that they do. They can keep twitter running well in the last 3-4 years, which means they are reasonably competent. That's all I have evidence for.
> If you want to prove that Twitter doesn't need such large and complex distributed systems, you can probably round up a few hundred thousand in angel funding and sell yourself to Twitter for a nice return.
As I have mentioned several times in other replies (and other discussions) - twitter's problem is not, in fact, engineering, and hasn't been since 2012 at least (it definitely was 2007-2009). They threw enough money/people at the problem, and solved it.
Right now, they are bringing in $2B/year, but spending $2.5B/year. If they are spending more than $200M/year on user facing hardware at this point, I'd be surprised. I'd even be surprised if they are spending $100M/year. But let's assume that I can save them $200M - that's nice, but won't actually save them - the company needs much more significant changes than that. And that part of the infrastructure only brings in users. I know not what systems they use to actually bring in money -- which is actually much more important to optimize.
And ... what I'm doing now is more profitable than what I can likely get from such a project (and I don't have to gamble or raise money). So, thanks, but I'll pass.
Let me ask you this, though: Look at the healthcare.gov debacle. I assume our discussion would have essentially been the same (including the counterfactuals) up until the point where a team rewrote it in a fraction of the time, with a fraction of the resources, and much much better - and I guess at that point we would be able to agree factuals.
Would you consider the discussion futile in either case? I don't care about agreement (factual or counterfactual), I'm trying to learn about the problems that are invisible to me. So far with Twitter, I've learned non so far over 8 years of the (more or less twice annual) discussion.
Again, I've mentioned several times: Whether or not they now run their infrastructure efficiently is very unlikely to matter -- unless they are horribly incompetent, which I assume they are not. I'm just trying to understand the aura of excellence (or alternatively, the depth of the problems) that Twitter deals with technically.
I assume the average tweet is closer to 80 bytes, and taking advantage of that can let you store twice as much tweets in memory.
[0] Does twitter count utf-8 bytes, code points, graphemes, or something else towards the 140-"char" limit? Regardless, the computation would not yield materially different results.
WhatsApp also doesn't provide a CDN, mobile analytics/monetization/debugging platform. WhatsApp also doesn't contribute to open source like Twitter does (Storm, Bootstrap, Aurora).
Describing twitter as a "140-char" mailing list is very reductionist. OTOH WhatsApp problem is very well served by Erlang/Mnesia an open source distributed messaging library developed by telecommunications company who were likely doing WhatApp-scale 10 yeas before WhatsApp came around. While WhatsApp engineers are great, there is very much a "standing on the shoulders of giants" - giants who have well researched the problem of building a system that lets two or three individuals communicate. There wasn't anyone doing Twitter 10 years ago (well, maybe Google).
I don't mean to imply that Twitter is an efficient well oiled machine, but the comparison to WhatsApp is daft.
[1] http://highscalability.com/blog/2014/3/31/how-whatsapp-grew-...
[2] https://www.erlang-factory.com/upload/presentations/601/EUC2...
In any case, my main point isn't that that WhatsApp is simple. When I said the parent posted was being reductionist, I was referring to Twitter's product offering. AFAIK (I could be wrong) WhatsApp is a messaging backend and multiple client frontends. The "50" engineers is deceiving becaus, despite the fact that Mnesia could or couldn't have scaled, there is a very non-trivial amount of engineers who worked on it and Mnesia is a good & proven system. In a similar light, Instagram was 13 engineers (or was it employees) and hundreds of django instances - and that maybe all you need for a photo feed, however Instagram didn't have a trending feature until earlier this year and how knows how many engineers that took.
If WhatApp hired another 200 engineers, likely 195 would either be twiddling their thumbs or working on new features outside what was needed in that acquisition to Facebook. Now we can make the argument that Twitter is bloated because they have too many products (Fabric, Vine, Crashlytics, and their numerous other acquisitions) - and the scaling back of those engineers may mean sunsetting those products - and thats fair. To imply they have 4000 engineers building a glorified 140-char mailing list is not.