2) found html tags and incorrect whitespace when exporting to TXT this page: https://www.packtpub.com/packt/offers/free-learning
I want to try this tool with pages that Instapaper fails to grab.
2) found html tags and incorrect whitespace when exporting to TXT this page: https://www.packtpub.com/packt/offers/free-learning
I want to try this tool with pages that Instapaper fails to grab.
For the Instataper pages that fail if you could send us some domain name you wish : hn at documentcyborg.com and we will test the parser against it to make sure it works.
1. Externally-provided content is dangerous. You might use a hash of the domain name, but I'd avoid files named after the sources.
2. Metadata such as a date would be useful.
3. Despite 1, a highly-sanitised hostname could be informative. An iconv to 8-bit ASCII [-a-zA-Z0-9_], and not allowing the first character to be '-' might be a start. Put a length limit on that as well.