I can’t speak to how it’s actually solved, but I can give a sketch from my experience. The game is generally to (1) consider denormalizing data and (2) rethink when you have the computer do work.
Denormalizing data refers to moving away from a model where you store one and only one copy of each ‘tweet’ in something like a relational database, to a model where you might actually store a separate copy of each tweet for each follower. That is kind of an extreme example of denormalization, but it’s a good way to illustrate how you could make it near-instant for any user to load their twitter homepage. If you stored the interesting tweets of every person I follow in a table just for me, you would make ‘reads’ (loading my homepage) incredibly cheap, but writes (someone with a lot of followers tweeting) very expensive.
Those trade offs exist everywhere in a system like this, which is (2), you get to decide when you do work. If celebrities with a million followers tweet an average of once a second, but people load their feeds a hundred thousand times a second, it is entirely acceptable to do 10000x more work for a celebrity tweet posting. This is actually the same concept behind using indexes in relational databases, but done more explicitly.
My personal answer to this question would probably start somewhat space-innefficient. I would take each tweet and put it into an event processing queue which writes a reference to it into each followers feed. This sounds nasty, but it scales with the number of writes, not reads, and we have less writes, and it scales linearly with the follower count of the writer. I would then improve efficiency by thinking about dormant accounts, caching, and maybe doing a bit more work on read.