My popular Twitter Analytics app has reached technical limits. Need help
tphq.tumblr.com
tphq.tumblr.com
* Technically: I had a well-optimized PostgreSQL database which had a few parts: Follower graph schema with a revision id (which node is following which), it got cleaned out every N revisions; a delta schema which took the last two revision ids and diff'd them; an aggregate schema which did a bunch of queries and summarized the results every T interval; metadata schema which stored cached information about each node (updated every time that user object was fetched).
I'm pretty obsessive about query and schema optimization, and I had a comprehensive benchmarking suite which helped me consistently improve my performance with bulk insertion as well as aggregate queries as well as user displayed queries. Each job was broken down into small efficient pieces that were executed in dependency order by my custom task scheduler, Turnip (open source at https://github.com/shazow/turnip).
I don't remember the exact numbers but I was approaching 100M rows on a single 512mb Linode.
Redis would have worked too but I would have needed much more RAM or more moving pieces to move things in and out of RAM for processing. None of my queries were slow enough to worry about this.
* Pricing: As others mentioned, higher prices make it easier to scale. I charged based on the size of the account and how many accounts you wanted to monitor (basically proxies to how many API calls you'll cost me). A small account cost something like $6/mo, 5 medium accounts $14/mo, 25 bigger accounts at $50/mo, 100 large accounts (1M+ followers) at $125/mo. I had modest revenue but I can't say my pricing scheme was perfect. I was actively messing with it towards the end.
* I had a legacy Twitter whitelisted account which gave me 10K api hits per hour. This helped me a lot. At the same time, I was careful to not become too dependent on that account in case I lost it. I was well within the boundary of normal user limits the entire time and only really used my whitelisted account to experiment or backfill new data. I made sure to always make the most efficient API calls to avoid wasting them. I too had issues with timeouts but it more came in waves when Twitter was having infrastructure issues rather than consistently. It wouldn't surprise me if this has gotten worse.
Also, I used and stitched all three Twitter APIs: REST, Search, and Streaming. It was painful.
* Diversify. Twitter is becoming an increasingly developer-unfriendly platform to build on, and your business should not be dependent on it. I added Facebook support to SocialGrapple, and I was going to add Google+ support too. Today, I'd also add app.net support. That said, the majority of my business was still Twitter, and that sucked. This was a big factor in my decision to sell out and shut it down—I didn't see the developer ecosystem as a place where you can have a sustainable business, let alone a thriving one.
I actually had several conversations/negotiations with Twitter about how they'd interpret their terms of service wrt my product. It helped to know people at the company to get a favourable ruling, but I still felt like it could be reversed—err, "provided with guidance" at any moment.
For what it's worth, I found it more rewarding to build an analytics product that was super useful for a smaller group of people than a little helpful for a lot of people (I'd say tweepsect.com is the latter). Think about where on the spectrum you want to be as this makes decisions, like pricing, easier.
Best of luck! Shoot me an email (in my profile) if you'd like more details.
again, thanks for this, i am going through it later again, just responding to over 60 other e-mails with help and support, just fascinating to see this. if one thing, we can hopefully make it clear that betting on someone's platform will provide tremendous opportunity but also introduce a considerable uncertainty if it takes off.
Don't build services based on other people's APIs.
It's sad but true. For all the talk about mashups, it's rare to find a demo that's actually cool and much rarer to find a real application because of these problems with API limits. Sure, you might be able to build something that can handle 99% of twitter users, but the interesting and profitable 1% will blow out the API limit and then you're hosed.
I guess it is a good solution, but it would require a lot of fundamental physics work so it might not match the time-frame he needs.
My point, and I apologise for the snark, is that when people have a problem and ask for help, they are not asking for judgement on things they should have done in the past, they are looking for ideas for how to move forward. If you feel strongly that people who build on APIs do not deserve this help, then perhaps consider making that point on any of the frequent "X API sucks" threads the pop up from time to time.
So I scale back up to 100 follower requests for the next call. It goes through. Next 100, fails. I scale down again …
So, that sounds like a cache is being primed on the first request. If this is a consistent pattern, you should be able to issue a request for one record, followed by a request for 100 and get it through most of the time (E.g. unless you run into a garbage collection cycle). If you code against this assumption, you should be able to utilise 50% of the theoretical limit of a 100/hour.
Is there something I'm missing?
but sometimes it's just random (depending on twitter overall traffic I'd assume).
plus, some records are faulty and can not be fetched. so this causes other issues as well and drops the api call.
1. Can you modify your system to return results based on a subset of their followers? If this is still of value to the user, it looks like the way to go, providing results based on a subset followed by another set of "final" results based on the full follower graph.
2. If you have enough celebrity-scale accounts using your service, are you able to share their followers details and cut down on time used to pull them?
3. On the business side, the service sounds dangerously like something that would be killed by Twitter if it becomes successful, by some definition of successful because you are pulling out their follower graphs. Look at Tumblr and Instagram. While it is good to milk it while you can and build it with ambition, it is also wise to look further and prepare for the day if it gets shut down and revenue goes to 0. If you have nothing to lose, go ahead, but be aware.
2) I can get IDs faster, and then theoretically use my own DB to see if I need details for that ID (i.e. follower details) from Twitter or have them stored. Problem is: I am duplicating Twitter's database, also, I had this before and the database grew to immense size and kept crashing all the time despite efforts to avoid it. Also: Celebrities have a large distribution of followers, so unlikely to save much time by seeing who's details I already got elsewhere.
3. Yes, although Twitter announced their quadrants recently, of services that will be supported by them, one of them is social analytics. This is what I do. So that should be good. But you have a point, I can't touch any money for up to 1 year in case Twitter shuts it down, I want to pay back the remaining unused time to my users. So it's frozen money.
but as said in other comments, i'm not sure twitter is happy with someone duplicating their entire user database over time.
I can imagine they are ok with that if your app provides value to their ecosystem.
i tried signing up through one of their partner thing forms but as expected, no response.
First, just check which of the the top 100 (1000) most-followed users follow the celebrity - this will likely eliminate the heaviest hitters immediately.
For accounts with normal amount of followers, do as you already do.
For the in-betweeners,
Pick your celebrity.
For a random 1% (.1%) of his/her followers:
Retrieve list of people followed by this follower
For each user followed by a follower of the celebrity:
Increase number of people following this user by 1
For each user in the top 1% (10%, .1%) of the above table of users:
Retrieve number of followers
Report user with highest number of followers from the above
This is, of course, based on the idea that often-followed users who follow a celebrity will also have many followers among the followers of the celebrity. Results will become more accurate as you poll more followers, of course.(I don't use Twitter or their API, and the above may be completely wrong.)
Second, maybe you could use the streaming API: you could get part of the data that way, and have more credits. If users follow back, you could use the sitestream, although it's quite different to work with then the REST API.
Thirdly, if I read correct, if Alice and Bob are both your clients, and Fred follows both, you now collect Fred's data twice, right? I would put a cache in between that. Riak, cluster of redis, or even S3 or DynamoDB. If I can help more, send me an email (in profile)
Lastly, if you have twitter investors in your userbase, ask for an intro to talk to twitter. They see the value of your service.
The streaming API idea is a good one to get the most out of all of Twitter's data sources... someone else suggested this as well. Seems like it's worth a try.
Your Alice/Bob/Fred assumption is correct. I had a cache of sorts, through a large mysql table which went way over a couple of gigabytes and kept crashing all the time and restoring took half a day.
The twitter investors haven't had a chance to see any of the service since they are still waiting for results :) tough one to ask for intros ;)
That seems to be the quickest win to reduce the number of requests, but it will depend on the overlap of your usersets how much this will help.
I'd build it with two stages. First, a cache that holds user information (sort of replicating profiles), that is, follower count and etc. This would be shared among all users, and can go into a dedicated Redis instance (and why not, also replicated to a MySQL-InnoDB for convenience).
Then, the "graph" DB (follower list), that I'd put into Redis. With some scripting and Redis magic, you can keep automatically sorted (server-side) users by their follower count. You'll just need a lot of RAM (get a dedicated server, look at ovh or others, cloud is usually more expensive and less reliable when it comes to RAM).
You can collect profile information before they go on to the 1.1 (which forces auth), to populate the global DB. Then, you'd only have to fetch users' follower IDs (using the 1.1: followers/ids), which I believe is way more reliable (and progressively, pooling queries, populate the profile database in batches of 500 or 250 users, using followers lists with user details).
This means that data can be queried dynamically without killing the server (or the servers, there should be more than one), therefore allowing for "partial results" (1M followers -> info about the first 10,000 just after signing up, for example).
also the system started making sense after a while when I had user ids that I had already cached (you are correct, the IDs I get through followers/ids which is much more well-though out function in terms of limits).
but then mysql constantly crashed and reparing/backing up a multiple-gigabyte-table exceeded my technical abilities and i gave up. so I split up everything into per-user-sqlite databases that I backup to S3. i lose the ability to access a cache of users though since I can't query other user's sqlite databases in a sane way to see if they have meta data for that user id.
major problem is that I believe twitter will eventually shut me down if I duplicate/replicate their user database (and I constantly need to refresh since user data will eventually be outdated).
By the way, which fields (nickname, avatar, follower count, following count…?) are you storing for follower representation?
this is the culprit call: https://dev.twitter.com/docs/api/1.1/get/users/lookup
i store anything that might provide sensible statistics later on from that response (so no profile colors or photos)
Because it's not a very good fit for this job?
What I do not understand is why these major players do not want to introduce API for money? I am ready to pay for it and I know a lot of people who also are ready to pay for it. But please, remove these limitations and make your APIs more stable.
And all of them (1M+ followers) have to wait up to 60 days or more to get past the login page due to a bug and limits of Twitter.
I feel there's a smart way to work around this and I have always managed to do so in the past, but now, I've hit my technical limits and need help.
I am willing to split upcoming PRO account payments 50/50 with anyone able to help me code moving forward / solve this issue.
and just to add: I think my particular problem is, and that's sort of the selling point, my analysis and reports is not just growth data (which is easy), but my calculations require all of your follower details to work correctly (to provide correct results). Twitter is sending me random follower details, so having a partial set means little reliability (for up to 60 days).
It does occur to me that the OP is effectively trying to replicate large chunks of the twitter datastore & that's going to be very difficult to manage! It's not like twitter themselves were particularly reliably to start with after all.
and regarding replicating, originally that's what I did. I had up to 20M of Twitter's 140M records cached almost - but that probably wasn't cool with them on the long run and i was unable to maintain a database with one table having multiple gigabytes of data.
Perhaps for some of the accounts that do not use much credit to search for followers, you could also search for those who it follows. Then on your backend, you can see if this subset exists within a previous list of followers of a big user.
I haven't actually checked out your service because i'm not really a twitter person so maybe you do this already, but could you provide statistics based on the amount of data you've got so far? (i know they will be inaccurate), but you can sort of give them as a "moving target", based on x number of your followers type thing. That way the user gets a little bit a value right away.
Also try enabling/disabling gzip compression for API calls.
simple http connections to parse/spider the follower records from public pages is a no-go since twitter blocks the IP then, and scaling this out will eventually not end in a good way.
I'm sure twitter must account for this, but how do they? You don't need to provide much information to get an API key.
I would be interested in knowing other solutions that work for you.
also keeping track of when the failures occur is important, and potentially valuable information.
This solution is only partially ugly if you are only after the follower count of a username.
I'm sure before the end of today someone here would help you out. Better still it might help trying to contact a few peeps directly.
i feel sort of bad asking for advice here, since many have better things to do and i am the one making money with this, but i've reached a point where I don't know how to continue. and this is very odd.
I am having inaccurate results. I can confirm I have a verified account and more than 0 retweets weekly.
I am guessing this inaccuracy is as a result of the challenge you are having with indexing.