Twint – Twitter scraping tool written in Python
github.com
github.com
https://github.com/pauldotknopf/twitter-dump/
You can use it to download every tweet from a user, not just the last 3000 that their API supports. It uses the same query syntax that the web search uses.
Check ```twitter-dump auth``` for instructions on how to use your web cookies with the command.
When the command queries Twitter, it will look as if it is coming from you.
Unhandled exception. System.ArgumentOutOfRangeException: Length cannot be less than zero. (Parameter 'length') at System.String.Substring(Int32 startIndex, Int32 length) at TwitterDump.Program.ParseCurlCommand(String curlCommand, Dictionary`2& headers) in /home/pknopf/git/twitter-dump/src/TwitterDump/Program.cs:line 150 at TwitterDump.Program.Auth(AuthOptions options) in /home/pknopf/git/twitter-dump/src/TwitterDump/Program.cs:line 113 at TwitterDump.Program.<>c.<Main>b__0_1(AuthOptions opts) in /home/pknopf/git/twitter-dump/src/TwitterDump/Program.cs:line 33 at CommandLine.ParserResultExtensions.MapResult[T1,T2,TResult](ParserResult`1 result, Func`2 parsedFunc1, Func`2 parsedFunc2, Func`2 notParsedFunc) at TwitterDump.Program.Main(String[] args) in /home/pknopf/git/twitter-dump/src/TwitterDump/Program.cs:line 30
What'm I doing wrong?
Oh god... no wonder they suspended me
Not sure how to do that with Twint off the top of my head, though.
We tried to use Twint in production but it didn’t work for us. I ended up writing one that works very well. Let me know if I can help you get your tweets.
Seems that Twint does all it's work unauthenticated so it may not be possible.
I suppose I could try using the API but I don't think I care that much.
> PS C:\Users\Jesse Hattabaugh\Documents\GitHub\twint> twint -u jessehattabaugh > CRITICAL:root:twint.get:User:'NoneType' object is not subscriptable
It requires extra hardware at the cash register (if using bluetooth) or modification to the vendors terminal to be able to display a qr code.
Sadly the app is extremely slow and cumbersome to use at a cash register compared to other options such as NFC payments on android or just straight tap and pay credit cards.
It would be a great system for payments online as scanning a code is very easy compared to entering all your CC information but becauses the fees are so high (compare to credit cards) the largest online electronics retailer (digitec) in Switzerland dropped them after the intro rates expired.
I wrote the Splunk integration with Twint for crawling Twitter timelines: https://github.com/twintproject/twint-splunk
Feel free to hit me up if there are any questions about that part of Twint.
Minor stuff: Printing instead of logging. Would prefer a package that only does the retrieval and nothing else. Hardcoded SQL(ite?) storage.
For Twitter in specific, isn't HTML scraping vastly preferable to using their official APIs? Otherwise you run into pretty arbitrary usage limits and missing features.
There's a small list of services where I think I prefer HTML scraping and browser piloting for a 3rd-party client: Twitter, Patreon, Facebook, LinkedIn, a few others. Services where the official APIs are underdeveloped or crippled to the point of almost uselessness.
Then it turned out Twitter refused half of the attendees the API key.. (maybe they thought it was spam coming from the same wifi, same time).
So then I just gave out my API key to the rest of the class, and in a few minutes it was blocked..
For a service that has a history of empowering users to protest and to spread news in crisis situations, it's a shame their API is so locked down.
The hardest part when working with data is often not manipulating data per se, but spending time on crap like this.
Took me 20 minutes to install twitter-dump, and despite a successful twitter-dump auth, I end up with 'Request exception Forbidden {"errors":[{"code":200,"message":"Forbidden."}]}'.
Not going to spend time to fix that, I'll use the dirty solution that works.
Maybe they could combine efforts? Maybe they could look at the code, and if the licenses allow, port things to theirs.
It kind of comes down to how well you can defend yourself from it being called a DOS attack (follow politeness standards and robots.txt), from violating their copyright (generally not problematic if you don't distribute the data), and from violating their terms of service (this is key in the case of twitter and reddit, carefully read their TOS).
However, the scraping of public information like in the case of tweets or reddit posts is the less problematic part. It's when you distribute the data or aggregations of the data that it could be problematic to use scraped public information.
As an example, here are two key features that are missing from the Twitter API at the moment:
- Bookmarks. You can privately bookmark tweets on the Twitter website and apps. There is no way to access the list of tweets you have bookmarked in the API.
- Threads. The concept of threads - where a tweet has replies from the same author get special treatment in terms of display - is key to how Twitter is used today. The official API doesn't support them, in that there is no way to look at a tweet and see that there exists a threaded tweet reply.
There is no good commercial reason for excluding either of these features from the public API, other than that Twitter have made a strategic decision not to invest resources in expanding the API to keep up with new features they are adding to the platform.
Given that, is it any surprise that people are resorting to scraping?
Which isn't a bad reason! But it's not a good argument for people not to scrape their own data.
Why would you want to scrape your own data if you an already request all your data and get a whole archive? https://help.twitter.com/en/managing-your-account/how-to-dow...
The Twint README calls out reasons for going beyond the API at the start - things like Twitter's increasingly strict rate limits and the limit of only 3,200 historic tweets for a user account.
Believe me, I would LOVE if the official API supported threads. I have tried several times to make it work, but the official APIs are stuck in a circa-2012 idea of how Twitter works. Replies just aren't a thing to it.
Most scrapers including this one lack these characteristics.
Not my problem. You make your http resources accessible, I'll download them however I want.