Twitter Announces Fire Hose Marketplace: Up to 10k Keyword Filters for 30 Cents
readwriteweb.com
readwriteweb.com
which is ((24x7x52)/12)x$0.3 = $218.40 a month
140 million tweets per day[1] / 24hr / $0.30 = 19.5million tweets per dollar.
_Surely_ there's some valuable data to be gleaned out of 20million tweets?
[1] wildly assuming Techcruch's numbers are connected with reality - http://techcrunch.com/2011/03/14/new-twitter-stats-140m-twee...
Thank God I use different pseudonyms on the internet. That just gives me the willies.
---
On another note, this may come at a very opportune time when the Obama 2012 is kicking into gear.
Yeah, people like you and I are still open to having any anonymity data-mined out of us through aggregate manipulation -- but at least it's a simple layer of abstraction.
In the meantime, if someone can get rich using the wealth of public information that every vapid college girl posting a thousand twitpics a night from her cell at the club puts up online, then more power to them.
Seems like data hackers are being taken care of.
Also, if you want a lot of social media cheaply, check out the sample web app for Google Buzz that runs on AppEngine. I ran it last summer with some of my own filters. I could run it about 5 hours a day before I hit the limit of a free AppEngine account - so it would not cost too much to pay to keep a derivative of this example program running 24x7.
Sure, they could do something else instead of selling data, but until end users pony up for the service, then they're the product.
I'm not pretending this is all it takes to run Twitter, but I'd be surprised if storing a few TB a year is a major cost center. (Serving up so many concurrent users seems like a much bigger and more expensive problem — that's an average of 1600 tweets per second, to say nothing of readers, and I suspect tweet rates are very lumpy.)
Clearly storing just the 'tweet' contents alone would be unhelpful because what about the username or any of the other 40+ metadata point a tweet carries.
What about keeping the mechanisms needed to store, sort, search, send those tweets, etc etc. I could go on.
Also, what were you expecting - Twitter to their business at-cost?
That's a huge underestimation. A tweet isn't 140 characters - it's 140 characters plus a huge chunk of surrounding metadata and indexes (who tweeted, when they tweeted, where they tweeted from, was it a reply, did it mention anyone, did it include any hash tags, did it link to anything, was it a retweet, its unique ID, how many users was it delivered to...) - all massively denormalised for performance reasons. See http://www.scribd.com/doc/30146338/map-of-a-tweet for an idea of the data involved.
Then there's the fact that a reference to each tweet has to be written in to the "inbox" of every user that receives it - so if Tim O'Reilly says something a reference to that tweet gets written 1,452,801 times, once for each of his followers.
On top of that, there's all of the associated stats collection, including link click tracking and a ton of data around who is doing what in the Twitter interface.
This article from last year suggests that Twitter were storing 8TB/day back in October, and it's only going to have gone up since then: http://techcrunch.com/2010/09/17/twitter-seeing-6-billion-ap...
Exactly. It should come as no surprise when any free site that accumulates a massive amount of data turns around and starts selling that data -- even if users feel like their privacy is being violated.
Twitter really has nothing to sell but data. Same for Facebook and others. They can sell that data indirectly (by allowing targeted advertising) or directly (by selling massive blocks of data for $0.30 an hour), but nobody should be surprised when it happens; it's all that they have to sell.