HNHacker News
TopNewBestAskShowJobs

stummjr

215 karma · joined December 7, 2009

submissionscomments
stummjr··on Ask HN: Website listing/explaining common Python pitfalls&tricks?
This repo may be helpful to you: https://github.com/satwikkansal/wtfpython

It’s a collection of snippets to explain behaviors that may be considered unexpected.

stummjr··on A CS degree is better than teaching yourself how to code?
I disagree with this article in so many ways. I have a CS degree and I work with many people who don’t, and they are just as good as (or even better than) me at all the points raised by the article.

They meet deadlines, they are incredibly good at communication and collaboration and they have pretty good networking. Most of these traits come from the fact that they needed to develop them in order to succeed in learning by themselves.

It is a pretty limited view of the world to think that only college can bring you this. Immersing yourself in a coding bootcamp for some people means leaving the jobs they need to survive in order to have a better job in the future. I can’t imagine how being on college can teach more about meeting deadlines, teamwork, communication and perseverance than that.

I wish this article provided more facts to back its beliefs up.

stummjr··on Ngrok: Secure tunnels to localhost
Ngrok is just awesome! A huge shout out to the developers!
stummjr··on How to Crawl the Web Politely with Scrapy
Crawl-delay is not in the standard robots.txt protocol, and according to Wikipedia, some bots have different interpretations for this value. That's why maybe many websites don't even bother defining the rate limits in robots.txt.
stummjr··on How to Crawl the Web Politely with Scrapy
That's kind of what Scrapy's AUTO_THROTTLE middleware does.
stummjr··on How to Crawl the Web Politely with Scrapy
Yeah, but that's not just because of web scraping. Plagiarism has been an issue for centuries.
stummjr··on How to Crawl the Web Politely with Scrapy
Scrapy is asynchronous, but it provides many settings that you can use to avoid DDoS a website, such as limiting the amount of simultaneous requests for each domain or IP address.

And yes, crawling politely requires a bit of effort from both ends: the crawler and the website.

stummjr··on Ask HN: If you don't permit telecommuting, why?
Scrapinghub is 100% remote from day zero. Nowadays there are ~140 people spread around the world, covering almost all timezones.
stummjr··on Show HN: Portia2Code — Turn Portia Spiders into Scrapy Spiders
Hey! I work for Scrapinghub. Feel free to ask any questions.
stummjr··on Introducing Scrapy Cloud 2.0
Hey, Valdir from Scrapinghub here! Feel free to ask any questions you might have about the platform.
stummjr··on A (not So) Short Story on Getting Decent Internet Access
This is such an inspiring story! I guess a lot of people today go the opposite way, renting or buying places where they know there is a decent internet access. I'm guilty about it. :)
stummjr··on Scrapy Tips from the Pros: April 2016 Edition
Hey, I'm the author. Feel free to ask any questions.
stummjr··on Scrapy Tips from the Pros: February 2016 Edition
Hey, of course. We are glad that you are interested in testing Kumo.

Please email us (help at scrapinghub . com) your user id, organization ids and the project ids you want to migrate to Kumo. Then we'll get back to you, giving early access to Kumo beta and documentation.

Just keep in mind that it is still an experimental platform.

stummjr··on Scrapy Tips from the Pros: February 2016 Edition
Hey, I'm the author of this post. Feel free to ask any questions or to suggest topics for the next month's post on the "Scrapy tips from the pros" series. :)
stummjr··on Scrapy Tips from the Pros
All sorts of careers, for example:

- developers who want to develop some data-based product (a travel agency website, who finds the best deals from airline companies);

- lawyers can use it to structure the data from Judgments and Laws (so that they are able to query the data for things like: which judges have interpreted this law in their judgments) (more on this: http://blog.scrapinghub.com/2016/01/13/vizlegal-rise-of-mach...)

- (data-)journalists who work on investigative data-based articles (they use it to gather the data to build visualizations, infographics, and also to support their arguments).

- real state agencies can use it to grab the prices of their competitors, or to get a map of what people are selling, what are the areas where there is more demand.

- large companies that want to track their online reputation can scrape forums, blogs, etc, for further analysis.

- online retailers that want to keep their prices balanced with their competitors can scrape the competitors websites collecting prices from them.

More on Quora: https://www.quora.com/What-are-examples-of-how-real-business...

stummjr··on Scrapy Tips from the Pros
Maybe you should give a try at Portia (http://scrapinghub.com/portia/). It does exactly what you mean.

You may also be interested in this library: https://github.com/scrapy/scrapely

stummjr··on Scrapy Tips from the Pros
Spider Contracts can help you: http://doc.scrapy.org/en/latest/topics/contracts.html
stummjr··on Scrapy Tips from the Pros
Scrapy can do the job, for sure. We use it to crawl more than 2 billion pages a month. :)
stummjr··on Scrapy Tips from the Pros
What do you mean by "data like that"? Metadata in Microdata format?

Btw, nice to hear your own experience here. :)

stummjr··on Scrapy Tips from the Pros
You're not alone on that! :)
stummjr··on Scrapy Tips from the Pros
There is an idea to make Scrapy support Spiders written in other languages[1]. It has been featured on the Google Summer of Code 2015[2].

[1] https://github.com/scrapy/scrapy/issues/1125

[2] http://gsoc2015.scrapinghub.com/ideas/#other-languages

stummjr··on Scrapy Tips from the Pros
1) detect pages that had changed since the last crawl, to avoid recrawling pages that hadn't changed?

You could use the deltafetch[1] middleware. It ignores requests to pages with items extracted in previous crawls.

2) detect pages that have changed their structure, breaking down the Spider that crawl it.

This is a tough one, since most of the spiders are heavily based on the HTML structure. You could use Spidermon [2] to monitor your spiders. It's available as an addon in the Scrapy Cloud platform [3], and there are plans to open source it in the near future. Also, dealing automatically with pages that change their structure is in the roadmap for Portia [4].

[1] https://github.com/scrapinghub/scrapylib/blob/master/scrapyl...

[2] http://doc.scrapinghub.com/addons.html?highlight=monitoring#...

[3] http://scrapinghub.com/scrapy-cloud/

[4] http://scrapinghub.com/portia/

stummjr··on Scrapy Tips from the Pros
Hey, not sure if I understood what you mean. Did you mean:

1) detect pages that had changed since the last crawl, to avoid recrawling pages that hadn't changed? 2) detect pages that have changed their structure, breaking down the Spider that crawl it.

stummjr··on Scrapy Tips from the Pros
Hey, author here! Feel free to ask any questions you have.
stummjr··on Ask HN: A martial art for a programmer
I think Aikido is a very good option, not only to programmers, but to all the people. It's a martial art which is meant to preserve the integrity of both players. Both mental and body health are the main concerns of this martial art.