HNHacker News
TopNewBestAskShowJobs

jancurn

748 karma · joined October 31, 2011

Jan Čurn

Software engineer | Founder & CEO @ Apify

https://x.com/jancurn

https://apify.com/jancurn

https://www.linkedin.com/in/jancurn

submissionscomments
jancurn··on Ask HN: Who is hiring? (June 2018)
I forgot to add, we're also looking for junior web developers
jancurn··on Ask HN: Who is hiring? (June 2018)
Apify | Full-stack software engineer | Onsite - Prague, Czech Republic | Full-time

Apify is a web scraping and automation platform that enables people to turn any website into an API. In the fall of 2015, we participated in the Y Combinator Fellowship in Mountain View, CA, where we publicly launched our service. Now we are 15 people and we are looking for a senior software engineer to join our team. We build a serverless computing platform with features like no other, build advanced HTTP proxy infrastructure, handle large-scale network and computing operations, publish open-source projects (https://github.com/apifytech) and have fun while doing that.

URL: https://www.apify.com

Technologies: Node.js, AWS (EC2, S3, DynamoDB, SQS, ECS, ...), Docker, Meteor.js, React, Linux, ...

More info: https://www.apify.com/jobs

Contact: jobs@apify.com

jancurn··on Open-sourcing gVisor, a sandboxed container runtime
Amazing work guys! I'm only wondering, is this ready for production environments?
jancurn··on Show HN: Headless Chrome Crawler
You can also use page.authenticate() for that - see a note at the bottom of the article. Also see https://github.com/GoogleChrome/puppeteer/pull/1732
jancurn··on Show HN: Headless Chrome Crawler
There's a workaround - https://blog.apify.com/how-to-make-headless-chrome-and-puppe...
jancurn··on Show HN: Web scraping page analyzer
Yes we are! Please see https://www.apify.com/jobs
jancurn··on Show HN: Web scraping page analyzer
You can use https://www.htbridge.com/ssl/ for that
jancurn··on Ask HN: What are best tools for web scraping?
Apify (https://www.apify.com) is a web scraping and automation platform where you can extract data from any website using a few simple lines of JavaScript. It's using headless browsers, so that people can extract data from pages that have complex structure, dynamic content or employ pagination.

Recently the platform added support for headless Chrome and Puppeteer, you can even run jobs written in Scrapy or any other library as long as it can be packaged as Docker container.

Disclaimer: I'm a co-founder of Apify

jancurn··on [dead]
Hey guys, sorry for the bother, we certainly don't want to abuse HN. All the posts had entirely different topics and we believe there are people who might genuinely benefit from them. If that's not the case, people won't upvote the posts and they will quickly fall into oblivion.
jancurn··on Show HN: Ready-to-use API to convert any web page to PDF using headless Chrome
I believe it's possible, by adding something like "scale: 1.5" to "pdfOptions" you might render an accessible PDF
jancurn··on Show HN: Ready-to-use API to convert any web page to PDF using headless Chrome
Not yet, but we can quite easily add these features. Just let me know at jan@apify.com what would you need
jancurn··on Show HN: Ready-to-use API to convert any web page to PDF using headless Chrome
If you'd like to add specific features, please let me know at jan@apify.com, I'm sure we'll figure it out
jancurn··on Show HN: Ready-to-use API to convert any web page to PDF using headless Chrome
We can easily add this feature. So you'd like to pass a single HTML file?
jancurn··on Show HN: Ready-to-use API to convert any web page to PDF using headless Chrome
Here we go - another Apify act that uses pdf2htmlEX to convert PDF to HTML:

https://www.apify.com/jancurn/pdf-to-html

jancurn··on Show HN: Ready-to-use API to convert any web page to PDF using headless Chrome
Now there is - I've just added the "sleepMillis" input option.
jancurn··on Show HN: Ready-to-use API to convert any web page to PDF using headless Chrome
Actually we're simply using Puppeteer - see the source code at the bottom of https://www.apify.com/jancurn/url-to-pdf
jancurn··on Show HN: Ready-to-use API to convert any web page to PDF using headless Chrome
There's another Apify act that extracts text from a PDF using the pdf-text-extract NPM package - see https://www.apify.com/juansgaitan/pdf-scraping If there's any library or tool that can convert PDF to HTML, it will only take a few minutes to setup such an API on Apify.
jancurn··on Show HN: Ready-to-use API to convert any web page to PDF using headless Chrome
You can call the API from anywhere, from resource-constrained servers, Docker containers that cannot run headless Chrome, JavaScript on a website etc. Also, we'll keep adding new features to this act to make it worthwhile to use it, e.g. retries on failures, posting of the file to some URL etc.
jancurn··on Show HN: Apify – Turn any website into an API
The crawler uses PhantomJS, but you can also create jobs that use a plain HTML downloader or even cheerio. The platform is very flexible.
jancurn··on Show HN: Apify – Turn any website into an API
Yes, every user has to agree with the Terms of use when signing up.
jancurn··on Show HN: Apify – Turn any website into an API
We consider ourselves only as service providers. We provide technical means to perform the crawling but the responsibility for the actual crawling is with our users. It's similar to Amazon AWS - they only provide infrastructure and it's your responsibility not to use it for anything illeagal.
jancurn··on Show HN: Apify – Turn any website into an API
Our crawler (https://www.apify.com/docs/crawler) currently runs on PhantomJS, while Actor (https://www.apify.com/docs/actor) can run arbitrary jobs, including jobs that use headless Chrome and/or Puppeteer - for example, see https://www.apify.com/docs/actor#examples-puppeteer

BTW we're working on migration of Crawler to headless Chrome.

If you could provide more details about your use case, I'm sure we will figure out how to do it with Chrome.

jancurn··on Show HN: Apify – Turn any website into an API
Thank you :)
jancurn··on Show HN: Apify – Turn any website into an API
Hey guys, this is Jan, co-founder of Apify.

Two years ago we showed HN Apifier - a hosted web crawler for developers - during our time in the Y Combinator Fellowship.

Today we're launching the largest upgrade of Apifier to date - a new product called Actor and a complete redesign of our website. Also, we changed our name to Apify.

Actor is a new serverless computing platform that enables execution of arbitrary pieces of code in the Apify cloud (we call them "acts"). For example, you can have an act to run a web automation job in headless Chrome with Puppeteer, to post-process data from the crawler in order to remove duplicates or to upload new contacts into your CRM. The possibilities are unlimited.

We also launched a library (https://www.apify.com/library), where you can find acts and crawlers built by other people and share yours to support the web scraping community. There are already several acts and people are adding new ones almost every day.

We're really looking forward to hearing what you think about Actor and to seeing what you can built with it. If you have any ideas, questions or feedback, just let us know on support@apify.com or here.

jancurn··on Introduction to web scraping with Python
With a headless browser the web scraping script can be even simpler. For example, have a look at the same scraper for datawhatnow.com at https://www.apify.com/jancurn/YP4Xg-api-datawhatnow-com
jancurn··on Does anyone know what's on this page?
Yeah, but check https://www.example.com/another vs https://www.example.com/anything-else
jancurn··on Is SQS down?
same here, our systems running in US East are behaving weird
jancurn··on Show HN: Jam API, turn any site into a JSON api using CSS selectors
Well, https://www.apifier.com does essentially the same thing, plus it supports JavaScript, can crawl through the whole website etc.

Disclaimer: I'm a cofounder there

jancurn··on Our Experience of the Inaugural Y Combinator Fellowship
Thank you, we didn't draw it ourselves, but our friend is a great artist :)

Yes, we got other advice too, primarily about funding and dealing with investors, but there's only so much you can squeeze into 1-2 hours.

jancurn··on Our Experience of the Inaugural Y Combinator Fellowship
I can't speak for YC in any way, but I can quote from http://fellowship.ycombinator.com/faq/: "We'll consider anything we can imagine becoming a very significant company."
← PreviousPage 2 of 3Next →