HNHacker News
TopNewBestAskShowJobs

changmin

62 karma · joined April 5, 2017

submissionscomments
changmin··on Quick and simple methods for opening numerous tabs at once [video]
1) CTRL + Click

2) Linkclump (Chrome extension)

3) Open Multiple URLs (Chrome extension)

4) Copy All Urls (Chrome extension)

changmin··on Show HN: Web Scraping with just one click. (Algorithm-based)
Thanks. It works on any webpages. The result looks nice to export to Excel. :)
changmin··on Web scraping with one-click. (Algorithm-based)
thanks, I submitted again with 'Show HN' as your suggestion.
changmin··on HTML to Excel data extraction
I appreciate your hard testing and feedback. Your suggestion is very good to me.

Actually, any (partial or full) HTML source code is available; <div></div>, <p></p>, <span></span>,<html></html>, and etc. Following to your advice, I changed the placeholder description to "any HTML Source Code".

Secondly, my server returns 500 error only if there is nothing to extract such as your code. I will fix it soon. Thank you.

changmin··on HTML to Excel data extraction
I tested now on Amazon product search, https://www.amazon.com/s/ref=nb_sb_noss_2/130-9531298-529675... . It works well though the result comes out slow (about 45 seconds). In result page, you can find the product list with the number of 28. For better speed, I agree to publish API or chrome extension.

In addition, it also works well in seconds with Craiglist apts / housing page. http://seoul.craigslist.co.kr/search/apa

Sorry for being slow. This is my private work. I could not predict a lot of new visitors, I need to scale up and out the server.

changmin··on HTML to Excel data extraction
I think Chrome extension is the best way for end-users, too.
changmin··on HTML to Excel data extraction
I have used Google Spreadsheet to extract <TABLE> or <UL> content. It works very well with them.

Compared to it, listly.io works well with all types of tag if there are repeated structures.

In my experiment, it works well with hunderds kinds of web sites.

e.g. Google/Bing search result, Amazon/Walmart/Ebay product list, Twitter/Facebook/Tumblr posts, Twitch list, Bloomberg finance info, Threads of a forum, Instagram comments, and etc.

changmin··on HTML to Excel data extraction
Thanks for the idea. It makes me think how to build API. I need to take a look at graphQL.
changmin··on HTML to Excel data extraction
The difference:

Import.io needs user's click to determine what to extract, thus, the user has to repeat it whenever the web page changes.

Listly.io needs URL or HTML codes. It always works even if the web page chages.

changmin··on HTML to Excel data extraction
Thank you, all.

Listly.io is my private work built days ago. I hope to hear opinions if it is useful for you... or not.

Listly.io turn HTML to Excel in seconds without coding. It finds the pattern of repeated structure and extracts all of image links and texts. It does find not tags (table, ul ...), but the structure.

Ideally for developers, I think API would be the best way to adapt this extractor to other scraper or your own scraper.