Twitter spam wave
twitter.com
twitter.com
> Twitter believes that your account may have been compromised by a website or service not associated with Twitter. We've reset your password to prevent others from accessing your account.
The spam tweet posted on my behalf was automatically deleted.
(Edit: For reference, the spammy link pointed to a domain called apaloreto dot info, but led to a 404 in my case)
If Twitter used the same app token for iPhone and Nexus twitter apps, and your Nexus twitter auth cred was stolen, they could use an implementation of the twitter client API with the iPhone user-agent, then post with your creds. I have absolutely no idea if any of that is accurate, but it might explain the access.
However, I wonder, shouldn't Twitter be able to pick these messages up automatically fairly fast, after (I assume) hundreds if not thousands of users have flagged them?
Also, the spammers can't have unlimited IP's. Twitters anti spam kinda seems to lag back behind E-Mail (subjectively).
Is there a reason the same techniques used in E-Mail aren't applicable to Twitter?
Twitter relies on very low latency - ie, once you tweet something, if it a whole minute to appear in your friends' timelines, it could already have lost much of its value.
Lots of spam reduction techniques introduce latency to levels that are unacceptable to Twitter's use case.
I'm not sure why it didn't catch these, but I can imagine why the same techniques aren't applicable in general.
Source: I wrote a system that can do it in less than 10ms at Grooveshark.
You can learn a little about how we think about securing login here: http://vimeo.com/80460475.
Although it operates at a much smaller scale than most of the things linked to here it does work pretty well.
I think the real problem is not latency, but simply they don't have the signals needed to differentiate spam from not.
Appears to have started on March 31st and has affected > 100K Tweets. It also appears to run between the hours of 2PM and 10PM PST, peaking at 7PM each day.
> shouldn't Twitter be able to pick these messages up automatically fairly fast
Theoretically, sure. As a human looking at an attack, it's usually pretty easy to pick out "obvious" attributes that they should have been able to catch. But when you're operating at a scale like us or Twitter, even stuff that looks like it's obviously-indicative-of-badness often has false-positives (posts flagged as spam that are not). The long tail of weird stuff that a billion users do can be pretty crazy.
At the same time, the "obvious" attributes of an attack are often very cheap for an attacker to change. Instead, we try to go after more expensive resources (domains, source IPs, etc).
> after (I assume) hundreds if not thousands of users have flagged them
Sadly, looking at flags of content is not a silver bullet. The signal is very sparse (a given spam post is rarely flagged), and nonspam posts are frequently flagged (religious and political speech are great examples - and they are the worst kind of false positive if you delete them as spam). These problems can be somewhat mitigated if you aggregate flags over a dimension that's expensive for the attacker (domain-posted, IP that posted the content, text shingles), but even then the recall isn't necessarily great and you could still catch e.g. controversial political domains.
> the spammers can't have unlimited IPs
True, though you can rent space on a botnet that has many, geographically-diverse, real-user IPs. Also, I imagine a significant chunk of posts to Twitter come from apps, many of which each use a single IP to post tons of content.
> Is there a reason the same techniques used in E-Mail aren't applicable to Twitter?
There's definitely some overlap. I'm not an expert at email anti-spam, but in general it's a relatively different problem. "Traditional" email spam is sent from some random email address on / via a compromised machine or open relay, and seems to be a relatively-well-solved. But it sounds like this twitter attack was caused by compromised accounts. At least anecdotally, it seems that email vendors are also not great at detecting this kind of attack. For example, my gmail account (with arguably the best spam protection in the industry?) gets a message every few weeks from some compromised friend's account. (i.e. someone had their email password stolen and the attacker is using it to "legitimately" send mail after authenticating to that email service with the correct password).
For sites operating at a smaller scale, this could be a good way to surface content for manual review though.
I imagine that while spam is annoying, it probably doesn't impact your bottom line in a big way.
Have you guys looked into sharing likelihoods of affected users? In other words I rarely share stuff I click on. If a higher than normal number of high view - low sharers like myself are sharing its either extremely popular or its spam. (Worth flagging for a manual check)