Show HN: Instagram-scraper – Scrape instagram photos by tags, without API
github.com
github.com
as somebody that has fielded numerous emails from friends asking me to remove tagged photos of them from flickr, i sort of wonder about the ethics of harvesting these sorts of images from instagram, a community whose norms sort of revolve around semi-public sharing of photos. I don't doubt that there's some rationale for harvesting the images from ig, but aside from thumbing your nose at their TOS, it feels like it's a greater violation of trust to harvest your friends and strangers photos for an ML project without their informed consent.
at the very least, it's worth considering pointing your app's gaze at a set of images licensed for any purpose whatsoever rather than ones that are explicitly licensed All Rights Reserved by their respective photographers.
Query Facebook for a list of location IDs near your location -> use those location IDs to get photos tagged with that location on Instagram -> wait for response for all of those photos to come back and then sort by recently taken. It ends up taking fairly long.
I still made the app anyway: https://itunes.apple.com/ca/app/feels-see-what-it-feels-like... but I am going to transition to using public Snapchat stories.
Even made an issue on the repo when I ran into an issue setting it up on AWS and the maintainer was fast to respond.
Used it to make a bot to scrap soup special menus from a sandwich place near my work: http://blog.matthewbrunelle.com/projects/2018/05/07/Soup-Bot...
The rate limit by instagram is a bit tough, though, but as i only for archiving a few of my close friends, as it supports private accounts, that's OK.
In terms of ethics I typically apply a sort of "try to be considerate" test. If I am doing personal, non-commercial scraping, only scraping public data, not releasing the data, using timer delays, respecting 429 codes, and not doing too much to mask my identity, and the servers I'm hitting are massive services handling multi-million user concurrency and I'm using a DigitalOcean droplet, then I'm not really a problem. And I trust their well paid sysadmins to block me or contact me if I am the problem. I do check robots.txt in advance. If there is one, it doesn't stop me, but it generally means I take extra precautions to avoid causing trouble for the service in question.
On the other hand, I once scraped the DPRK's English language press office as part of some research, and about two weeks later it became inaccessible from any country other than Japan, so I'm pretty sure I almost caused a diplomatic incident. Oops.
In this case, it's 40 lines of python code using a single threaded requests request, only scraping a single page, not doing any funny business with user agent spoofing, etc. I think you're right, it's probably not something Instagram wants to happen and theoretically I'm sure they could send a takedown, but I guess I have a pretty laissez-faire attitude about this: is it really a good use of their time to stop this guy?
I understand what you're getting at, but if someone tells you explicitly not to do something (e.g. in the terms of service of their website), doing that thing anyway doesn't seem very considerate.
That said, however, there's no straightforward way to work with Instagram. It's original APIs are both limited and locked down. The new Facebookified APIs are limited and next to impossible to work with (they are geared exclusively to ads/marketing). So ¯\_(ツ)_/¯
In a side project I use a library that effectively reverse-engineers Instagram's private API, pretends it's a user using a browser etc.
If you search for Instagram Private API, you’ll find implementations for almost any programming language (with aforementioned PHP being the first, probably)
Not saying this is right, but if the service provider wants to play cat and mouse then I’m happy to take part.
My point is, you can’t just put content online and expect to put restrictions on how I send the HTTP requests for it on my side. And if you think you can well I’ll do my best to prove you wrong.
Edit: I totally agree that the data they’re giving me is their property and I merely have a license to use it. I’m not advocating for copyright infringement or anything like that. But if the license allows me to use the content for X purpose, then the way I request the content (whether through a browser or a scraper) shouldn’t matter.
https://en.wikipedia.org/wiki/Browse_wrap#Summary
Remember, anyone can sue you for anything and it is then your burden to defend yourself if they are persistent and wealthy enough.
just take every search engine.
Want to show me an ad? Have an actual human being manually pick the ad in real time and deliver it to me!
The physical world has plenty of laws and/or common sense/customs about what behavior is not okay even though it is fully possible. Why should the net be any different?
FB and LinkedIn both scrape contacts (and who knows what else), from their users email accounts (and surely lots of other darker sources as well). Pretty much all of Google's search engine's content is scraped from other websites.
I agree with you, but technically that argument is true
And as for Google, in theory a robots.txt would block their scraper.
There was an earlier case (I don’t recall the details) about scraping violating copyright, which might be a valid defense - that you’re engaged in unauthorized “copying” the (protected) source code (of the webpage), albeit momentarily, into your computer’s memory while your script processes it to extract the relevant contents. So even though it might be fine to copy the image/data/whatever else, you have no right to copy source code - which is certainly protected by copyright (your browser processing the same page is something the copyright owner explicitly allows, by virtue of making it available on the web). By that same token, it can also be considered illegal to save webpages on your harddisk. But do refer to the last sentence in my earlier paragraph.
The LinkedIn vs HiQ case is currently being appealed in the (court of appeals for the) 9th Circuit. Regardless of which way it rules, the decision of the 9th circuit only applies in its the states under its jurisdiction (maybe as a precedent in others, but it won’t be binding). It can still go to Supreme Court, depending on how determined the parties are.
There was a ruling in a different (but somewhat related case) in Jan 2018 that violating the ToS is not a crime https://www.eff.org/deeplinks/2018/01/ninth-circuit-doubles-... To quote: “[T]aking data using a method prohibited by the applicable terms of use”— i.e., scraping — “when the taking itself generally is permitted, does not violate” the state computer crime laws”.
There was a very long discussion on HN about scraping over a year back: https://news.ycombinator.com/item?id=13884357
Long story short I don't care about the "legality" of it. If it's publicly posted it's fair game. I don't care about tos or copyrights either. Everything is on the table.
I reversed engineered the API myself a couple of weeks ago which was great fun - especially figuring out Instagram's rate limits on interactions such as comments and likes per day/hr.
re.compile('(?:#)([A-Za-z0-9_](?:(?:[A-Za-z0-9_]|(?:\.(?!\.))){0,28}(?:[A-Za-z0-9_]))?)')
Thanks!