Dangerzone: Convert potentially dangerous PDFs, documents, or images to safe PDF
github.com
github.com
I for one have been looking a lot into PDF/A for security. PDF/A is really meant for archival, but as a side effect has disallowed an awful lot of weird PDF features which are a security nightmare and pdf readers tend to implement badly/buggily. PDF/A-1 for example, the strictest level, disallows JPEG2000, TIFF, JavaScript, PostScript, embedded files... (PDF/A-3, FWIW is essentially useless from this angle, because they decided to allow arbitrary embedded files, so a valid PDF/A-3 could have pretty much anything in it).
There now exists a good PDF/A validator (https://verapdf.org/) which can be used to ensure PDFs conform to the standard, but of course, won't fix them if they're not.
PDF/A has an interesting implementation detail however - compliant PDF readers are supposed to automatically "turn off" non-PDF/A features when they encounter a PDF which declares itself as a particular PDF/A variant (even if it then goes on to attempt to use non-compliant features), which would hopefully prevent dangerous sections from being decoded and avoid exploitation). Another interesting feature of PDF is its appendable nature, which might raise the possibility of being able to "declare" an arbitrary PDF as PDF/A by simply appending an extra section to it, hopefully rendering it less harmful (though possibly at the expense of it appearing to have missing content when rendered).
If the reader can open a PDF/A-1 file and ignore the bad parts, can't it open a PDF file as PDF/A-1 and remove the bad parts, before saving it again?
Then you could use the "re-rendering" technique to extract images:
1. Create a stripped PDF/A-1 file from the original PDF.
2. In a VM, render the two PDF files to two high-resolution image sets.
3. Use some CV algorithm to find the differences. For example, gaussian blur, subtract, threshold, find islands.
4. Use this to come up with areas which use complicated PDF features and/or images. Say this returns that there is an image on page 9 in the rectangle ((17, 338), (400, 300)).
5. Crop out page 9 in rectangle ((17, 338), (400, 300)) from the original PDF. Use some CV algorithm to detect the DPI and whether it's best to encode as JPEG or PNG. Encode it and add to list of images.
6. Add sanitized images back from list, mark PDF/A-1 file with images as PDF/A-2.
Of course, you could do this for links or whatever as well. Spit out a list of rectangles and link targets in the PDF, and then put them back in.
Fully sanitizing the PDF yields better guarantees of security at the cost of lost functionality.
Yes, note my original emphasised use of the term "supposed to".
> Fully sanitizing the PDF yields better guarantees of security at the cost of lost functionality.
Indeed it's a tradeoff. But if you're willing to throw away the features which this extreme sanitization would trample across and have any ability to design PDF out of your system, you're probably better off not using PDF at all in favour of some straightforward image format.
As jrowley said, if you trust the reader to sanitize it safely, why not trust it to open normal PDF files safely?
Well, they are. You try using a thousand page manual that doesn't have a section outline, or try sending around PDFs whose authenticity is legally important (to people who aren't capable of using gpg). This is only scratching the surface - there are many features which are important to people with e.g. accessibility needs which aren't directly visible.
> This has some minor advantages, but it's also a large attack surface.
These are not a large attack surface. TIFF, JavaScript/PostScript, 3D content and video (yes - PDFs can contain video!) are a "large attack surface".
> As jrowley said, if you trust the reader to sanitize it safely, why not trust it to open normal PDF files safely?
Well, firstly, I emphasised how "PDF readers are supposed to automatically 'turn off' non-PDF/A features", so I already acknowledge the caveat. And as far as "trusting them" goes, disabling the decoding of certain features involves very few LOC. Completely implementing the whole specification for the more exotic PDF features is an incomparably huge number of LOC, which are proportionately more likely to have flaws.
I ... think there’s another technique that relies a bit on trusting the printing drivers to do the right thing, where you can tell Ghostscript to print your document, and target another PDF. This should at least remove interactive components in a PDF
https://bugs.chromium.org/p/project-zero/issues/detail?id=16...
Which is why when you print a PDF from Firefox, it doesn't look very nice. But it's safer than sending unsanitized PDFs to printers.
(1) https://news.ycombinator.com/item?id=19344146
Edited: Added in the comments of that post there is a reference to pocorgtfo16.pdf: is valid as a PDF document, a ZIP archive, and a Bash script that runs a Python webserver which hosts Kaitai Struct’s WebIDE which, allows you to view the file’s own annotated bytes. The zip archive has further resources to insane reversing deep dives, code to study and more.
I don't open any files on my PC from people I don't personally know -- use webviewers.
Click-through rate for a technically legitimate "you are party to a lawsuit" email must be sky high.
edit: Thank you for both answers. I thought it had to do with sandbox rationale, but couldn't mentally get past the fact that sandbox could potentially be escaped too. Eh, I think it is time for sleep.
Safer: definitely. Given that the collective amount of PDF attacks is some number, now this particular PDF needs to attack PDF and the webviewer. Assuming that 1% of all PDFs do that, I'd say it's 100 times safer than not using a webviewer.
If you still think that 1% of all potential PDF attacks is too unsafe, then that's a different discussion.
If you think my 1% is off, then that's a different discussion too. All I'm saying is that it's safer.
What's the most popular, Chrome browser I'd have thought?
The Preview.app sandbox is not quite as secure as what is used by browsers such as Firefox or Chrome on macOS for web content, so there is probably still a benefit to viewing PDFs in the browser, depending on whether the browser or PDF viewer is more hardened against these attacks.
See: https://theinvisiblethings.blogspot.com/2013/02/converting-u...
Then over time Adobe added a number of interactive (forms), multimedia and rich media (embedded JS) features, leading to even more vectors.
Note the attempts at the link to sanitize image formats that don't have over the top complexity.
If your question is why an electronic document format has support for images and interactivity, I don't know what to tell you.
What you would end up with is an image that looks like a link but would not be clickable.
I'm not completely sure, but wouldn't this parse links and make them accessible again, possibly even clickable?
Fortunately smart forms require Adobe Viewer, and there's an approval step (similar to agreeing to Excel macros), but after that it can do whatever the hell it likes.
Download Dangerzone
So instead of opening a PDF from the internet on my machine, I now run code from the internet on my machine?