Back of the Envelope Calculations at Twitter 2.0
cohesive.so
cohesive.so
https://d18rn0p25nwr6d.cloudfront.net/CIK-0001418091/947c0c3...
Twitters cost of revenue was 1.7B in FY2021. That includes their data centers, operations salaries, depreciation of assets, and AWS costs (plus other categories, it’s defined in the financials). There’s also a delta of 400M from FY2020 and over half of it is salaries.
So from these pieces of coarse information and the suggestion in this piece that storage costs alone are around 1B, I can surmise that these envelope calculations are off by a lot.
0.54 petabytes of “tweet” data a day, no multimedia, seems extremely high. This also assumes no compression or LUTs, which would make it multiple orders of magnitude off.
The multimedia as well leaves me with questions, how much of it is original content? What’s the rate of deduplication (because many people post the same non-original content)?
I’m sure they work with orders of magnitude of data larger than most of us are working on, but I highly doubt that Twitter alone contributes to 5-10% of Amazon AWS total revenue.
Not to mention, the author also completely ignores read traffic (which is the majority of traffic for a social network), replication/redundancy/HA, so many other things. It's just a garbage article. It appears the company hosting this blog pays randos to submit posts on any topic? https://www.cohesive.so/write-for-cohesive
The same mistake is made for video. But one paragraph later the author uses the per five years value as if it's per day.
> 1% of tweets contain videos of about 100MB each
I can't imagine the average video on twitter is taking up 100MB of storage.
W/o bulk discounts, ~$0.4 gets you 32GB of memory and 16 vCPU on AWS.
So this is only 400k blades. ~20 blades per rack. ~20k racks.
That's like 1 large data center in US, EU, Asia, & South America. Sure, Twitter should be getting a lot more for that because they have a larger scale and AWS has like >50% margins or whatever.
But before anyone says this is out of control for a company the size of Twitter, it's really not.
Another order of magnitude is you're forgetting the cost of 3x storage & buffer for 99.99% uptime.
Another order of magnitude is that data transfer is REALLY expensive. They're not paying just for compute & storage.
If you're Twitter, this is also true to a lesser extent.
Netflix & Twitter have more server expenses than just querying databases.
I don’t know that it makes sense to say a company is profitable when they make a net loss, regardless of gross profit. In 2021 Twitter made a net loss of over $220 million.
They only managed to report net income in 2018 and 2019.
So looking at the overall picture I still argue that saying Twitter is a profitable business to be very very optimistic and not really true.
* Profitable in 2018
* Profitable in 2019
* Would have been profitable in 2021, if not for the shareholder lawsuit
* Profitable in H1 2022, despite having an operating loss, due to making a $655m profit on the sale of MoPub
I'm not saying this is a great business. But some folks are making it sound like Twitter was about to collapse and go broke on its own, which is ridiculous. They still had $2.7B cash + $3.4B short-term investments on hand at the end of Q2!
Due to a legal settlement payout of $800M. Without that one-off, they'd have been $580M in profit.
My assumption would be that a new image is uploaded per 1000 tweets a day. As retweeting an image is more popular, that doesn't occupy space.
And 100 MB videos are pretty rare. Most are short, few seconds.
It seems pretty unlikely that twitter is ingesting 15M videos a day. That would mean that they are ingesting (15 million * max video length of 2.3 minutes) / (1440 minutes per day) = 400 hours of video every minute.
Given Twitter's video length limit it suggests that more people are uploading videos to Twitter than are to YouTube which seems unlikely.
The number is derived solely from the MAU/DAU number, but bots are not counted there. One active user may be running several bots that automatically post content on a schedule or as replies to keywords.
I'm not sure how to estimate the number of bots, but by their nature they can produce a lot more tweets than a person.
Fagpacket maths is great for specific problems, Can I make a mobile bridge, What price should I sell product x at. But all these examples have hard and agreed upon constraints.
The example fails because the author doesn't appear to have worked with systems of a big enough scale to make good assumptions. Once you start hitting petabytes its way way way cheaper to do on prem. The issue becomes one of backup and retrieval speed (which they touch on)
The reason why this is important is because for this fagpacket calc to work, you need to know the rough price per meg for storage of each asset class(and its purpose, ie, for machine learning, serving the CDN, for logs etc etc).
Tweets will be stored differently from pictures, and pictures will be different from video.
Archiving is impractical for an always on system (ie all old tweets and media need to load within 150ms)
The other big issue here is that caching is never really mentioned here. CDNs will be one of the bigger costs for twitter. The more efficient caching is, the less your raw storage costs are.
TLDR: Fagpacket maths is great, but you need to know your constraints and parameters first. Otherwise you'll endup with wildly incorrect numbers.
OTOH, perhaps that omission makes this a more accurate representation of how Elmo makes decisions. :P
Not much of a surprise there, but necessary to significantly reduce these operational costs despite all the deranged and exaggerated scare stories of the immediate and complete collapse of the blue bird site.
Any time the numbers are that far off, I’d question the conclusion being advanced by the person who isn’t legally liable for errors.