Show HN: Distributed Scraper
stdlib.com
stdlib.com
The question we asked ourselves was --- if I'm a developer today, and I want to release services for people to easily use and compose, take advantage of new "serverless" technology, where do I put it and how do I distribute it? How do I keep my services organized? Can I build a business around an API? What does that look like?
The name "Standard Library" was a natural fit --- I consider it a huge amount of serendipity that the domain was available. We've also been lucky to have a huge amount of support from our developer community --- we got a few questions like this to begin with, which naturally caused some unease, but far more positive reactions in total. I'm very happy we've been able to build something people like. The name choice has not negatively impacted our SEO or anything like that, thankfully! :)
If you happen to play with it at all, let me know what you think!
(Humorous aside; it is more frequent for newer developers to ask us about "sexually transmitted diseases" than it is for them to get the C reference - which is totally fine, don't mind sharing knowledge! But when I told my non-CS mother about the business we were building, definitely had to really quickly explain the background of the name.)
Hmm. web-devs are known for NIH, this will only encourage this reputation..
But if the content is pretty neatly and easily scraped then the other recommendations make sense.
Though a bit of Javascript can now take you a long way ;)
AMA!
I haven't had issues hitting a wall with getting caught doing any scraping. But then again, I haven't done it at a 10k/pages/sec rate or anything like that.
I have wanted to implement something like this one - ie: Lambda doing the downloading of the page itself.
I have wondered how it would work with very strict sites like Yelp - limits similar to what you would get in their API (so doesn't make sense not to use their API).
What are your stats like if you don't mind? How much people are using it and much are getting blocked (404 or 500 after 1000 requests, etc.)?
Edit: Is it possible to use my own credentials for AWS?
Based on StdLib's dashboards – a bunch of folks have been using it per month with a steady pace of a few 100 scrapes a day type of thing.
We've been using it internally for quite a while now.
And as far as I know, StdLib doesn't allow you to use your own AWS credentials. They have their own gateway and a bunch of stuff on top of Lambda that makes the whole experience a lot easier and more powerful (e.g. 128MB limit on payload vs 5MB for Lambda)