Twitter Starts Rolling Out Option To Download Your Twitter Archive In One File
techcrunch.com
techcrunch.com
When you click the "get my tweets" button, it kicks off a process in Twitter land somewhere. A few minutes later, you get an email with a link to a zip file.
The file contains a full archive of every one of your public tweets, including @replies you've made, but not DM's or replies to you or follower/following information, etc. It's just your public tweets. The tweets themselves are stored as CSV and JSON. Which is actually pretty cool because it means you can build your own apps, or archive apps like Thinkup can ingest your tweets, if you're so inclined.
As the article states, you can explore your archive via a web app that all works client side in a browser, based on Bootstrap, natch. The app works quite well. You can search your archive quickly and easily. It points to the canonical URL of the tweet on twitter.com. There's some pretty basic visualization of your tweet archive.
The javascript that runs the app is all minified so it's kind of hard to explore. The application is named "Grailbird" which I thought was kinda clever ( http://en.wikipedia.org/wiki/Ivory-billed_Woodpecker ).
On a personal note, it was pretty great (though often cringeworthy) to be able to roll back through 6 years of tweets. I found the very first tweet that the woman I'd end up marrying every replied to. The tweets that led to friendships and career changes.
It's a solid first step. Nice work, Twitter.
Would you be willing to share an example of the JSON structure for a tweet?
Grailbird.data.tweets_2006_12 =
[{
"source" : "web",
"entities" : {
"user_mentions" : [ ],
"media" : [ ],
"hashtags" : [ ],
"urls" : [ ]
},
"geo" : {
},
"id_str" : "547413",
"text" : "counting down the seconds until 5",
"id" : 547413,
"created_at" : "Sat Dec 02 00:57:17 +0000 2006",
"user" : {
"name" : "Jim Ray",
"screen_name" : "jimray",
"protected" : false,
"id_str" : "35623",
"profile_image_url_https" : "https://si0.twimg.com/profile_images/1234214846/avatar_normal.jpg,
"id" : 35623,
"verified" : false
} ]
The CSV data is much more basic 547413,2006-12-02 00:57:17 +0000,counting down the seconds until 5,10000000000000001 === 10000000000000000 True
Wow I would hate to work at Twitter.
Plus, CEO's comments sound like it's been set up as a side project. The kind of management style books like Mythical Man Month and Peopleware warned us about.
http://engineering.twitter.com/2012/04/mysql-at-twitter.html
http://highscalability.com/blog/2011/12/19/how-twitter-store...
http://www.percona.com/live/mysql-conference-2012/sessions/g...
http://www.slideshare.net/yousukehara/introduction-of-twitte...
http://engineering.twitter.com/2010/05/introducing-flockdb.h...
http://www.infoq.com/news/2009/06/Twitter-Architecture
http://engineering.twitter.com/2010/05/introducing-flockdb.h...
"It looks like Twitter has started rolling out the option to let users download all their tweets"