Show HN: Doc Converter – Convert PDF docs to Word documents on your computer
docconverter.app
docconverter.app
I have been gradually improving it for the past five years. It is a part of my photo editor https://www.Photopea.com. I know really a lot about PDF, I wish I didn't know that much :D I am glad to see that there are others who try to "make sense" of PDF files instead of just rendering them :)
** fun fact: Often, a PDF contains text as an array of characters, each has its X and Y coordinate and a style (white characters omitted). It is up to you to "cluster" them into words, lines, paragraphs ...
** Often, PDF text is made uneditable (on purpose). You see a text "Hello", but in fact, there is a text "bsiin", and a font, which renders "b" with a shape that looks like a letter "H", "s" as "e", and so on. If you open that PDF in a PDF viewer, select "Hello" and copy-paste it elsewhere, you get "bsiin".
https://community.adobe.com/t5/photoshop-ecosystem-discussio...
But the fact that we as a society have accepted
> adobe has shut down the licensing service for the version I have on disc (CS3).
as something normal and acceptable is insane to me.
Photopea also helps when you're at a random computer and can help someone do an edit that would otherwise require access to a computer with software installed.
I think you meant raster
To edit it, I guess you could paint over the text with white, and add a new text on top of it.
If so, it's hardly noteworthy. If you've written your own PDF to DOCX converter, then you have an interesting technical story (or ten) to tell -- do tell.
Successfully installed PyMuPDF-1.21.1 fire-0.5.0 fonttools-4.38.0 lxml-4.9.2 numpy-1.24.1 opencv-python-4.7.0.68 pdf2docx-0.5.6 python-docx-0.8.11 six-1.16.0 termcolor-2.2.0So, it's a wrapper around not panddoc but pdf2docx,
https://github.com/dothinking/pdf2docx
which parses PDF via PyMuPDF,
https://github.com/pymupdf/PyMuPDF
which is a wrapper around MuPDF (which does the heavy lifting parsing PDF),
and writes DOCX via python-docx,
However, from an everyday user point of view, it does make it rather simple to convert pdf to word document. An everyday user won't be up for doing that via cli commands. And every alternative user friendly solution requires uploading your documents to servers (which could spark privacy concerns)
If the customer base is less technically adept, wouldn't most of them not care and just upload it to a cloud service? I ask sincerely - recently I've realized I don't have as firm of a grip on the 'average consumer' as I thought.
The Windows app is an unsigned executable - not planning on running it myself.
What could possibly be missing at the moment is a written instruction that documents where to locate the code base on the user's machine post-installation
it's not will it sell it's how many will it sell
Getting PDFs into the pandoc intermediate representation would probably work on such a small subset of PDFs, pandoc does not even bother trying.
https://support.microsoft.com/en-us/office/opening-pdfs-in-w...
https://support.microsoft.com/en-us/office/edit-a-pdf-b2d1d7...
> This works best with PDFs that are mostly text
I somehow doubt OPs converter does much better. I hope to be proven wrong though.
Of course I imagined a different kinds of integrations but maybe this is a case of great minds thinking alike!
Its license is also linked to the computer it was activated on. Change computers... too bad!
It's particularly egregious where there don't seem to be any substantive further improvements planned, and the underlying engine was not built by you.
> there don't seem to be any substantive further improvements planned
About that, the truth is, it often starts with a use-case as simple as "convert PDF to Word". Improvements usually would come from user feedback, and continuous maintenance. While I deliberately started out to keep the App simple, there's a good chance that features and functionality would expand when user feedback gets in the mix.
I'm saying this from my experience with maintaining IGdm (https://github.com/igdmapps/igdm)
This is not true, as far as I can tell.
> PDF has an edge here in practice because most use cases are read-only
I think many folks would be extremely happy to have full-fidelity read-only access to Microsoft formats without having to have Microsoft Office.
On the other side, it's _extremely_ common to both produce and consume PDF without over touching an Adobe product.