The Editable PDF Initiative
editablepdf.org
editablepdf.org
Perhaps it's that way because it's read-only? If PDF files had to be generated in such a way that they can later be edited, things would get a lot more complex and probably less reliable.
Also it feels like it's the wrong way to go about it, because no matter what PDF editing will never be as powerful as a proper text editor. So it would be the wrong tool to collaborate on a document (because as soon as you want to do something more advanced with layout, images, etc. you probably can't). Maybe it's good if you want to quickly amend a contract before sending it, but then you need to remember that your .doc is no longer the latest version.
Basically a PDF document shouldn't be the source of truth for document editing as that would lock you to the wrong format.
That's why it's used in marketing for things like lookbooks and elsewhere for things like contracts, that are read-only by design and should never be edited by anyone but the entity that wrote it in the first place.
It's a little more complicated but not difficult to edit charts.
Really the issue is that we need a non-MS-Word editable document format that includes hash/signature features to ensure the edit/publish state of the document.
Sometimes you want to send out a finalized document and want to make 99% of the users unable to edit them. That's what PDFs are for. Imagine lawyers needing to send out a finalized contract. Or a graphic designer sending out the finalized design. Or an electronic book that has gone through the work of the author, the editor, and the publisher and needs no more changes. PDFs give an air of permanency and stability when so many other digital formats are malleable.
Literally the only reason I make and send out PDFs is because they're effectively read-only. (It's not really perfectly so, but nobody should claim it's a tamper-proof document...)
[1] https://editablepdf.org/faq/
>But isn’t the whole point of PDF that you can’t edit it? No. The fact that standard PDFs are difficult to edit is more of an accident than a feature, as PDF’s roots are in printing, where only final-form documents needed to be transmitted. Many people believe PDF to be “impossible to edit,” but beware: minor edits in PDFs, such as swapping figures on an invoice, are trivial — therefore you need other technologies, such as digital signatures, to verify that your PDFs have not been tampered with. More extensive edits, however, are more difficult, as they require the document’s logical structure to be automatically detected, and this is an error-prone task.
[...]
>Are you sure we need such a editable PDF format? I believe one of the most important benefits of PDF is its concrete, solid state. The idea of Editable PDF stems from a real-world need to improve the efficiency in the way that we work with documents. Today, the only editable file formats are those native to the applications that generated documents, and none of these formats guarantees the layout to be preserved in the same way as PDF. Furthermore, despite improvements in compatibility, using a native file format still often requires the recipient to be using the same software (and often the same version) of the application, which may not be available.
PDF’s largest asset, its rock-solid visual presentation, will remain, and editable PDFs will be backwardly compatible with the current installed base of PDF viewers such as Adobe Reader and Preview.
I do not consider my objection addressed. If anything, that FAQ further emphasized the current difficulty of editing PDFs and they are only proposing to make things easier. So no one is really objecting to the fact that currently changing PDFs is hard. Whether or not that is desirable though, is a separate matter, one that the FAQ does a poor job explaining.
I disagree with this - I was an IT journalist when PDFs came out and all the blurb at the time was centred around the advantage of being able to create a document and know that it could be displayed by anyone, irrespective of machine, OS etc while retaining visual fidelity.
PDF only became a thing in the print production world quite a bit later. For many years, you were making sure your printer got the QuarkXpress Files and all the high-res asset files collected together into a single folder and zipped up.
Their main complaint seems to be structural metadata (this text is a heading, this text is the same font as that text on another page, so if you can change one the other should change, etc). I don't think at that point it's worth keeping PDF in the name, it'll confuse people. I certainly don't want to receive files built like that since printers tend to have memory issues with bloated files.
You can already do minor edits anyway with some knowledge (the spec is pretty easy to read) and some programming or a hex editor. The only issue I have is fonts, they are very complicated.
The PDF file is reliable in that aspect while other widely used document formats like .doc and .docx are basically a gamble, even if you open the document on the same machine with the same software. The same with presentations: You just need to open the file in a slightly different software or version or encounter one of the countless bugs and suddenly you have a picture overlapping with text.
By your argument, any programming language, no much how safe it purports to be, fails because they allow you to write bugs.
Few programming languages claim to provide bug free programs, so your example is irrelevant.
kccqzy's comment was clearly referring to security.
That is not what PDFs are for. PDF is, well, a Portable Document Format. It is not convenient to modify a PDF, but PDF is not securely resistant to modification (discounting its cryptographic features [0]). Its resistance to modification is a side-effect of its design, not a primary goal.
An attacker will be able to modify your PDF. This gets easier every year, as we'd expect. That doesn't matter, though, as an attacker can always recreate the document, with whichever changes they wish. (Again, neither of these attacks will work if you use cryptographic signing.)
If you want secure assurance of authenticity, you use cryptographic signing. No excuses. If you're a lawyer, I'd hope you aren't placing any stock at all in the inconvenience of modifying a PDF.
[0] https://acrobat.adobe.com/uk/en/sign/capabilities/digital-si...
This is not about how easy or how hard it is to modify a pdf, it's about the intended purpose. The fact that it's meant for publishing means we get to optimise it as such, both in terms of simplicity of the format itself, and in terms of the tools that interact with it. This makes consistent-ish rendering much easier. The features that would enable the format to be "editable" are also the sort of features that make consistency hard.
If reports that come out of systems need to be edited they should dump to excel or word and not pdf.
While the P in PDF means portable I think it’s better thought of as “published” as in “published document file”.
The user story here is that PDFs get sent around as forms to be filled out, and that poses a problem for non Mac users or users without sufficient technical skill.
And since you reference mp3 and jpg, you surely know that both formats can be modified in ways that many people will not recognize as modifications. It just pushes the skill level up a bit. But there's always a technically capable person available for hire to modify one of the "permanent" formats you mention.
Why is that a problem? I'm pretty sure there's a PDF reader app for every major platform.
I think of PostScript as an "ink on paper" format.
While you can take apart the PDF / PS format, dictionaries, etc. It's not a high level representation format, like a word processor. It's a way of specifying how to draw vector shapes onto "paper".
If editing PDFs is something you find yourself needing to do regularly, something is very wrong with the process that's leading to this. It may not be your fault, it may be an upstream party who should be providing you with the source material, but either way making PDFs easier to edit is not the correct solution.
This doesn't break the simplicity of the PDF, while making it easy to edit.
PDF is great for:
- Archiving. It’s self-contained and will work 20 years down the line.
- Math. Anything with equations.
- Printing.
I’ve tried various techniques to archive web pages with varying degrees of success. With PDFs I don’t need to think about it.
- ctrl-S - not self-contained, compatibility depends on which browser you use (Safari at least gets it right and puts it in a single file, but then other browsers can't open it. Firefox usually saves it correctly and portably, but now you have multiple files, and you can't rename them because they have references to each other, which makes it hard to organize. Then there's all the JavaScript that web pages sometimes have, which can break in the archive.)
If a PDF is available it's easily better than these alternatives. In general I save webpages I really want to refer to later by copy-pasting the text out and manually reformatting as Markdown, or sometimes with wget.
What I would like, is to make PDF's easy to markup with highlighting, circles, comments etc. Currently, even in Acrobat, it's not very intuitive.
Certainly, between me a mac user, and colleague using Linux I was able to provide feedback on their documentation in this way ...
Often, a PDF contains just a single raster bitmap with the whole content rasterized. Also, text is often converted to vector shapes, which also makes it non-editable (as text). But it can open / save PDFs from Google Docs and other editors quite well.
When the format was created, computers only had a few KBs of RAM. Yet the format should be capable of editing documents with thousand of pages. The format solves this issue by delegating the memory management to the user.
Also, the file was made with the assumption it was suppose to be printed, not shared. It is easier to hide parts of the document instead of removing the data.
A funny trivia. The PDF is suppose to be read from the end of file. That's why some documents need to fully downloaded before they can display the first page. Of course, nowadays most PDF are linearized and load, at least, the first page right away.
Over the years specification got so complex it became very hard to implement a minimal editor, viewer, parser or generator. If the format was simpler, it would be possible to make "save as PDF" more accessible.
I've other issues with the typesetting and the way color is handled (it has a printer first approach), but I think this post got too long already. I just want to point out the spec supports so many pointless features such drawing in 3D space, movies, audio, HTML support, etc.
Finally, I don't understand why most people are against a revision on the PDF format despite clearly having very little knowledge on how it works. I think multi person edition of the same entry with some version control can be useful. By the way, the format kinda let many people edit the document at once, as long as they are not working in the same part.
That's a good decision. Make the file format versatile and powerful. Don't constrain it by the limitations of contemporary hardware.
> Also, the file was made with the assumption it was supposed to be printed, not shared. It is easier to hide parts of the document instead of removing the data.
I agree it's made with the assumption of being printed, but that's part of the appeal—preserving visual fidelity of how the document looks. You can't send people a docx and expect them to see the exact same thing on their screen down to every detail.
And no it's not difficult to remove data. If you know exactly what to remove, it is quite easy to remove things. To remove text, find the Tj or TJ operators, remove them and their arguments. To remove an image, find the Do operator (occasionally BI, ID, EI) and remove it. You might have to perform decompression before doing that. For images, you might have to run another pass to delete the referenced object. But all these are all very easily automated.
> Over the years specification got so complex it became very hard to implement a minimal editor, viewer, parser or generator. If the format was simpler, it would be possible to make "save as PDF" more accessible.
The reason "save to PDF" is difficult to implement from scratch is not because of its complicated specification. Indeed parsers are quite easy to write. The real reason "save to PDF" is difficult to implement is because PDF wants visual fidelity; that comes at the price of specifying where exactly text should be placed, all the way from how paragraphs are flowed to how kerning of the letter is to be handled. Most applications do not care about these details. Most developers hardly have any interest in understanding line-breaking algorithms or interpreting font files to produce the right offsets and glyphs (think ligatures). These things are, rightfully, way beyond the business domain of typical applications and beyond the knowledge of typical developers.
I think you are severly underselling how difficult it is to do these things.
1. You need to be a specialist in how the PDF format works.
2. From experience, it's not trivial to have logic that correctly handles all possible formatting cases in a PDF. 2.
With the entry removal example, I was trying to show the format was not meant not to be shared. I know it is possible to remove data in other ways, and that probably every modern editor removes the data correctly. But it was not how the format itself deals with it. Of course, hiding entries with the flag had others uses such only print only the pages you currently working on without having to rescan the whole file.
I agree most devs don't have interest in learning how to do typesetting. But also, typesetting is quite complex by itself, specially when dealing with non western language. Luckily, projects such Harfbuzz (nowadays, hb is used even my emacs) makes it a lot easier.
Like I said in my original post, the format is anachronous. I don't think the format is intrinsically bad, I just think the format is not right for your time. I think we can do better nowadays.
PS: I've been thinking, it would be pretty cool to talk with the engineering team that worked on the first spec, and actually know what they were thinking back them and what they would change in it nowadays.
As it's a zip file it also needs to be read from the end although it can be linearised as well.
I don't think this is right. Postscript maybe ... but PDF in my experience came about in the 90s, when computers typically had between 4 and 16MB of RAM ...
No, exactly.
The PDF is a terrible format, yet if I'm sending an email with an attachment I want you to see exactly how it looks on my computer then I'm exporting to PDF.
However if your book is only available as a pdf I'm probably going to skip it.
PDF is good for short things, a contract maybe. The best use case is forms which this doesn't really talk about but seems to address, the web has basically solved it, but there are times you want to send people a form to fill out that you don't want the formatting to be go wacky on, but still need to be editable.
PDF can do this but isn't good at it, this seems to take that not good and make it good.
Wait, what format do you expext a book to be? I mostly skip any book that is not on pdf
One of the problems with including mathematical formulae in a reflowable document format is that the concept of reflowable mathematical notation simply doesn't seem to exist, so In practice you'll end up with something equivalent to a picture.
To preempt: I work within printing, yes there are tools to hack into PDFs and make certain alterations or fixes, but it's to get you out of a bind only, it's not a normal healthy workflow.
> Many people believe PDF to be “impossible to edit,” but beware: minor edits in PDFs, such as swapping figures on an invoice, are trivial — therefore you need other technologies, such as digital signatures, to verify that your PDFs have not been tampered with.
That's not really the point that it can't be edited, the value to me is that the sender has confidence that it will look the same to the receiver as the sender.
Different PDF reader software, and even sometimes the same reader software installed in different environments, can render PDFs, especially those containing any text differently.
If you want confidence that it will appear the same, pure image formats are a safer bet.
If I send you a PDF the point is that I don't want you to edit it. Otherwise I would have sent a docx file. An 'editable PDF' may as well get lumped in to the OpenDocument standard. It is already universally editable and I'm sure it has support for adding application-specific metadata.
Having a PDF editor isn't some violation of sacred principles; but making changes to the standard to make it easy to edit is not improving the situation. I want to send a format that is hard to edit.
It's an online service that lets you upload PDFs, then edit fields, add text, upload and paste images like your signature, etc. Perfect for filling out tedious paper application forms without having to deal with printing & scanning.
I have no connection other than as a satisfied user, and in fact I have no idea how they make money, since the free mode features suffice for every use case I've had.
It also does PDF editing perfectly. I really hope there will be some open source version of it at some point. Or that someone's working on one.
A PDF renderer basically needs to be able to rasterize fonts and paint glyphs on a page/screen – that's it. Layout, spacing and even kerning are left to the producing application.
The project mentions the lack of robustness inherent to web-based document formats, but I'm afraid that any alternative would either be severely limited in the range of achievable output documents or would end up reinventing the wheel.
As an analogy: SVG has been around for a while, and yet we still use PNGs and I don't see them going away anytime soon.
Maybe what we really need is just more widespread support of ePub, and maybe some extensions for more "document-like" (instead of book-like) functionality in editors for it, and potentially support for an embedded rendered PDF for layout stability?
In Polar we have taken the perspective that immutability is an advantage and is going to be the basis for our group collaboration around documents.
We ended up building out annotations on top of PDF including text highlights and area highlights which can then be commented on:
https://getpolarized.io/docs/annotation-sidebar.html
Some of our users keep asking for editable documentation and I think the main win here could just be using markdown which I'm thinking about adding.
The biggest thing that's needed though, for scientific use, is latex. Fortunately, there are plenty of markdown implementations with latex support.
PDF is amazingly good for printing documents but honestly 90% of the complex printing requirements aren't needed for regular use.
Although it's an interesting idea, I suspect it will never work in practice because word processing is just too complex. There are just too many complex features that people expect to have available. Different implementations will never be sufficiently compatible. Perhaps the solution is to bundle your document with a WebAssembly binary of a particular version of LibreOffice? OK, maybe you could separate the rendering functionality from the UI stuff, but it's hard to see how in practice you could get documents to be editable and rendered in the same way everywhere except by having everyone run the same binary to do the rendering, and there will inevitable be a hundred versions of that binary in use as new features get added.
It would be a lot of effort to create a document format with the kind of richness that PDF supports. I am dubious it would be worth it.
I think most people do not need an editable PDF in the first place, so this is a minority problem. If you do want this, for most people there is already a working solution... just store Open document format within PDFs.
While editing PDFs on Linux for me was always connected with pain I also had no joy using a plain macos for this. While the preview app is able to do some things it cannot do others that matter.
I wanted to copy some text just yesterday - while I could select and copy it I could not insert it as a text again in the same application.
To have to use some extremely overpriced adobe product for sometimes doing tasks like this is overkill and really unnecessary.
To all the people who like PDF because you cannot edit it like you want: This is the "obfuscation argument" because anyone who has the right tool or googles for 10 min. can somehow edit PDF - it is just a real pain to do so most of the time and the result may look like the patched overhead transparencies we saw back in school in the earlier days.
This quote was supposed to be an absurd hypothetical. But I guess we'll live to see it in reality.
Pdf's aren't promoted as a portable editable format, but a portable, sharable, and archival format.
Why promote PDF over ODF? Is the issues of document reflow, of an editable document such an issue that they need to develop a new set of tools, and change the structure of PDF to resolve the issue, if that is the case, it seems they could contribute to resolving the issue in ODF?
Being read-only is the thing why we have PDF in the first place. If you want to do changes, go back to the program where it came from. It's simple as that! :)
Inserting a scanned signature is also not a problem at all these days, and even fits the PDF model quite well.
Yes, but the person creating the PDF must know this, and must know how to do it with the software they have. In practice, many forms I've seen came out of situations where this was not the case.
> Inserting a scanned signature is also not a problem at all these days, and even fits the PDF model quite well.
It's not a problem to open a one-page PDF in Inkscape or the Gimp and paste a signature in there. That's what I said. It gets tedious with multi-page PDFs. Do you have a better solution for this?
I used to write print production software. I'm no stranger to PDF.
I recently had to fill out a PDF form and send it back. It took me way too long to figure out the "form elements" were just images. I kept trying to use different clients, thinking the content creator must have used some poorly supported corner case of the PDF spec.
So I printed the frikkin PDF, wrote on it, scanned it, and sent it back.
What could be easier?
You also get enhanced accessibility. (I often need to reflow the pdf when reading from mobile device)
It would be nice to have an open standard for 3D food printing though.
> It doesn’t adapt to screen sizes
That is a feature. I expect my PDFs to display with pixel-perfect consistency everywhere.
There are other formats that adapt to screen sizes. HTML is good for that, if we ignore how people break that with styling.
I worked with archivists on a few projects and never appreciated the dumpster fire that electronic documents presented.
PDF is an amazing thing as you get an expressive format that preserves look, feel and content and will likely do so for the foreseeable future. Just the fact that the US Federal courts standardized on PDF for most filings will ensure that it is a viable format for decades or more.
It's a great WORM format. Every added feature makes it worse.
There are great systems for those already.
When I want a PDF, it's because I want a format that I know is always going to look the same.
A PDF is a great archive format. It's perfect for a scan of a document, or a printout.
I never want my viewer to add anything to it, I never want it to detect anything, I never want it to adjust anything.
Just render it exactly the same way, every time.
I worked on a project where we were digitizing and cataloging various records. It was less challenging to do this with papers from the British colonial administration from the late 1700s, than to decipher certain 1980s documents written with a defunct word processor. PDF is a compromise that helps address that issue.
I would not recommend maintaining your general ledger in a PDF. But an annual report that may be referenced for decades is a great example of why a PDF is a useful format.
PDFs are often display focused and difficult to parse, but it’s certainly possible to do so.
It’s success in the market as compared to a edit focused format like ODF underlined how important display consistency is.
It’s fine for that purpose but it’s terrible for eBooks, manuals, science papers and a lot of other stuff it’s used for. Some HTML with everything in one file would be much better in my opinion. Something like CHM maybe which MS used to use for help files.
You just described the ePub format. It’s a ZIP file with HTML docs in it.
In short, I don't think we could have typeset these grammars without PDF.
We also did a grammar of Dhivehi, which is the only language in the world that uses the Thaana script. Thaana can be typeset quite easily--if you happen to have the right font. Most people don't. I guess the same thing holds if you happen to be publishing grammars of languages that used cuneiform--not many computer systems have cuneiform fonts!
But in general, if you generate the PDF with an authoring tool like LaTeX or InDesign, or if you print to PDF from a webpage or document, it's going to be selectable in a sensible way.
> Why not use web standards, such as HTML/CSS?
> We do use the relevant parts of HTML and CSS, where appropriate. But web standards do not provide for specification of the layout of the document in a robust way, which is guaranteed not to reflow when opened on other systems. Furthermore, browser technologies are a moving target, with implementations changing very rapidly. Therefore, they do not provide a suitable basis for archival documents.
If you want to open your documents in a decade or more, redesigning a “cleaner” PDF would only make that less likely. If you want something cleaner than PDF, then XPS is already here. I don’t understand what scenario we’d have where designing a completely new format would give us better software support. So, the reason I’d see for designing a new format is if neither XPS nor PDF are good enough for some application.
So we should kill a format because it is used for something it was never intended for?
How is the lack of a widely compatible, self-contained markup-plus-resources format PDFs fault?