Show HN: Transform any web page into a document
documentcyborg.com
documentcyborg.com
Why does it return .zip that needs to be unpacked? To save bandwidth your could just use gzip 'Content-Encoding' end return the format requested by the user, which would be unpacked by the browser.
Returned file name is [documentcyborg.com].zip, it would be nicer if the domain of the requested document was used instead.
Here's a freebie name that's (as of writing this) unregistered: page2doc.com
2) found html tags and incorrect whitespace when exporting to TXT this page: https://www.packtpub.com/packt/offers/free-learning
I want to try this tool with pages that Instapaper fails to grab.
For the Instataper pages that fail if you could send us some domain name you wish : hn at documentcyborg.com and we will test the parser against it to make sure it works.
1. Externally-provided content is dangerous. You might use a hash of the domain name, but I'd avoid files named after the sources.
2. Metadata such as a date would be useful.
3. Despite 1, a highly-sanitised hostname could be informative. An iconv to 8-bit ASCII [-a-zA-Z0-9_], and not allowing the first character to be '-' might be a start. Put a length limit on that as well.
other wise nice tool, thx
https://www.reddit.com/r/DotA2/comments/50neqc/will_future_u...
>We couldn't find any text for creating the document. Please send us the problematic link : https://documentcyborg.com/ via our contactus form.
This was the first page I thought to try it on, so you might want to consider adding more text to your landing page so that it will work.
I don't understand what is the advantage of server side processing. I have always depended on Evernote Clearly and Redability mobilizer to turn any web page into nice text which can be copied and pasted on MS Word. Then I can use Prince (http://www.princexml.com/) to batch convert docx into whatever I want. If someone finds it slow then they can enable auto capture clipboard. press ctrl+` to enable Clearly, ctrl+A to select all text, ctrl+C to send it to clipboard.
I wish if someone can work on this to make it smother for a desktop users.
PS: I also make use of reading view mode of Firefox.
In some ways, Oh By[1] performs the exact opposite role. Which is to say, Oh By allows you to transform any document into a web page.
Well, any document 4096 characters or shorter ...
[1] https://0x.co
While you're at it, why not throw in the conversion to a few (popular) image formats as well? I can think of (at least) some scenarios where that would come in quite handy (e.g. posting long articles to Twitter).
All the best moving forward.
Not bad but the title of the article is missing.
The tweet is missing too but I can't decide if it's a good or a bad thing.
There should be options to remove pictures too I think.
Honestly I'd be interested by a standalone product like this. I don't like the fact that you know everything I store.
Later I tried a text export of the same page, some HTML remains (</div> elements).
Using the title of the page to name the zip would be nice too.
https://medium.com/@subes01/this-is-your-life-in-silicon-val...
It seems to search for the largest body of text and omit the rest.
Nice service. Only advice I can give is to tweak the typography a bit for the PDF output: I'd restrict the measure (line lengths) to around sixty characters, and boost the leading (space between lines) to about 1.5 times the line-height. Personally, I'd make the body text a bit smaller too, but the bigger text might be preferred by some.
Great for offlineish portfolios.
Interesting idea though.