When you say "@foo Hi, foo!" You're sending a message to some person named foo. That's a push. If you start thinking of it as a pull, then you get in to the sort of trouble you're mired in.
In a distributed system the process goes like this:
1) Sender queries name server: WHOIS foo? 2) Nameserver responds foo is http://www.foo.com/mytwitterfeed. 3) Sender pings recipient http//www.foo.com/mytwitterfeed?ping=messageid&origin=bar 4) Recipient queries name server: WHOIS bar? 5) Name server responds http://bar.com/mytwitterfeed 6) Recipient requests message text, http://bar.com/mytwitterfeed?id=messageid and verifies it actually is addressed to him.
From then on, you have to worry about spam but that's a solved problem (killfiles, Bayesian filtering, etc.) No caching is involved. Recipients store messages addressed to them, just like any other messaging system. This isn't at all a hard problem.
Yes, people can send you false pings, which is why you need to go back through the name server and do a reverse lookup, and once you fetch the message you need to parse it and make sure it within bounds (less than 160 characters or whatever) and addressed to you (has @foo in the message text.) and is not spam (the name isn't in your killfile, is in your whitelist, akismet says it isn't spam, etc.) Rate limiting helps too, real people aren't going to send you 100 messages a second.
Personally I would go with:
User Adam sends message to Bob: 1) Adam sends message to Twitter service. "Bob>Hi." 2) Twitter check to see if Bob knows Adam. 3) Twitter sends "Adam>Hi." to Bob.
(With optimal storage so Bob and Adam can see their old conversations)
Your system let's random spammers people find out people's address which IMO is bad.
PS: Or Adam could send a message to "All>I like this soup." and Twitter then sends the message to everyone that cares about Adam including Bob.
Edit: It looks like Twitter is sending around 2million Tweets a day.
Also, all these proposed messaging architectures are somewhat flawed. Twitter isn't really "sending" messages to other people. It's more like a person's message history is bound to their account, and appropriate privileges are applied. Then, when people try to "read," privileges are obeyed and information is produced...
Am I the only one who believes the solution to Twitter does not involve a massive distributed system? Twitter is inherently centralized... people don't see that.
PS: Text is cheep EX: Slashdot is running off of ~4 computers. And bandwidth is not a scaling issue until you need a single system to handle more than ~1GB of bandwidth per second. (Saturating an OC-3 line costs a lot of money but you don't need to change your architecture to do so.)
Aka - a scoble broadcast out to 20,000 different endpoints causes challenges - whether it's pushed or pulled..