Minimal PDF
brendanzagaeski.appspot.com
brendanzagaeski.appspot.com
Some top tips; if you decompress the streams first, you'll get something you can read and edit with a text editor
mutool clean -d -i in.pdf out.pdf
If you hand mess with the PDF, you can run it through mutool again to fix up the object positions.Text isn't flowed / layed out like HTML. Every glyph is more or less manually positioned.
Text is generally done with subset fonts. As a result characters end up being mapped to \1, \2 etc. So you can't normally just search for strings but you can often - though not always easily find the characters from the Unicode map.
I think a lot of people in the dev/power uesr community would mind paying $1 for a Kindle ebook where you note all your findings.
There have been so many instances where I wanted to do stuff with pdfs but ended up deflated.
> subset fonts
So you mean if a font has been embedded with three glyphs, 0x41=A, 0x61=a, 0x62=b, then string Aba would be \1\3\2?
I think you mean "a handcrafted pdf"?
To answer your question, subsetting a font just means taking a portion of its glyphs and it doesn't imply remapping. In fact for almost sane PDF files you will find ASCII characters mapped to themselves, making text search within decompressed PDF possible. My dirty watermark remover script basically uses qpdf to decompress the thing and then use regular expressions to search for Tj or TJ right after the specified string.
http://wwwimages.adobe.com/content/dam/Adobe/en/devnet/pdf/p...
This is a long document but it is very well written, if you read it on the bus or while you're waiting for your compiler to finish, you will get to understand it.
You can get a good overview of the state of the fonts in your PDF using:
pdffonts file.pdf
There's a column which tells you if there's s Unicode map available for the font. That's important. Because PDF is just rendering glyphs at positions, it doesn't even know what the character names are. To allow you to copy and paste, most fonts in most pdfs will have a Unicode map from the glyph id to the Unicode symbol.If that's not available, in some cases you can rebuild it yourself by looking at the character encodings and substitutions.
On the book, do you have any examples? I'll probably never get around to writing anything down, but if it looks easy enough it's probably worth having a stab at.
Also, large caveat, I'm not a PDF or font expert. I've probably decimated the terminology here but hopefully it gives you a rough idea.
Wish I still had a copy but it was a while back.
For historical interest, older versions (going back to 1.3) are here: http://www.adobe.com/devnet/pdf/pdf_reference_archive.html
1.2, 1.1, and 1.0 can be found elsewhere on the Internet.
A long time ago when I only had access to an extremely slow 2G network but I had to send a large-ish PDF file, I used qpdf to decompress the whole file as much as possible and then using xz -9 to compress it. Way better compression ratio.
If you need more, the "free" (trade for your email) e-book from Syncfusion PDF Succinctly demonstrates manipulation barely one level of abstraction higher (not calculating any offsets manually): https://www.syncfusion.com/resources/techportal/details/eboo...
"With the help of a utility program called pdftk[1] from PDF Labs, we’ll build a PDF document from scratch, learning how to position elements, select fonts, draw vector graphics, and create interactive tables of contents along the way."
Just a note that Google lets you get a new email address that isn't spammable, through security through obscurity:
-> You can use your gmail address, add a + after it, and add a keyword. So if you are jsmith@gmail.com you can give out jsmith+syncfusionpdfsuccintly@gmail.com and then later if that starts getting spammed you can redirect it.
NOTE:
This is an incorrect solution (Google, please fix this) because anyone can run a regex removing the + part.
Instead the correct solution is that if you have gmail open, in a single click you should be able to generate a high-entropy gmail address (that does not deplete the namespace) and link it on your end with "syncfusionpdfsuccintly".
If I already have gmail open, 7 seconds to create a new gmail address as follows:
1. Click something to start the process
2. Type "syncfusionpdfsuccintly" to tag it on my end
3. Click something to copy a resulting high-entropy gmail name into the clipboard.
I should then be able to paste it into a form, get it delivered straight into my inbox (never spam), and redirect it to spam if it starts getting spammed.This would allow people to contact us without ever getting into spam, while entirely removing their ability to contact us if this email address starts getting spammed. There are no downsides.
I believe Google's engineers are smart enough to move from security through obscurity (relying on the knowledge that no spammer can ever invent and run the exact regex s/\+[.]+@/@/g to remove the security through obscurity, as this would entirely break this security, exposing the underlying "protected" email addresses) to something that works.
Until that day comes you can rely on the security through obscurity to give out a secure email address that can't be spammed. Just add a + and a tag!
Please.
Google: I believe you are smart enough to understand this comment and implement this solution, which can be prototyped in 30 minutes and solves the spam problem forever. You can do it! I believe in you. You're 99.999% there and your security through obscurity works very well for me. I use it.
I hope you will go above and beyond and solve the remaining 0.001%. It would just make me feel better to know that a 13-character regex couldn't defeat your solution.
Right now it just silently drops expired addresses. But it is so satisfying to think about bouncing (but stuff like bouncing behavior is something you have to consider when running a mail service).
I was thinking of turning it into a service, but I'd have to read up on how to scale it. Running an SMTP server takes a lot of care. I've found that just using nearlyfreespeech.net's mail forwarding is most reliable to receive emails. So I do that for now, since it is on a small scale.
I just got so frustrated with how this problem has such an obvious technical solution. At least for us users, anyway. It's not a solution for the marketers.
I strongly suspect that Google is very reluctant to do anything to make the email landscape unstable. I think that if Google started offering this, it would shake up so much of their business.
Running an inbound SMTP server is much easier than running outbound smtp for laypeople. The software (postfix, exim, etc.) is rock solid (you have to REALLY mess up to lose emails) and the protocol is very forgiving (all serious senders have good retry policies). I encourage you to try!
This a reasonably priced paid service that does already exist, it's not from Google and no, I won't name it neither publicly nor privately because I used it for years and I don't want their domain to be banned by those requiring your email address for everything. If you know how to search you will probably find it. Basically you sign with them using one valid email address, then on their interface you can create as many addresses as you need (IIRC there is a limit but I used even dozens at a time without problems) and add a keyword to them. All of those addresses will be redirected to the email you signed with, but the From: field will also contain the keyword you specified so that if you create an address for each service you sign up for, you instantly recognize who is spamming you when they use that address. This is very effective and I filtered out a lot of spammers. I'm surprised there are no more services like this one around, or probably there are many but they keep a low profile to avoid being banned. That's why I'm not going to name that service, sorry. But it does exist indeed and is technically easy to implement.
I'm not asking for the service to "exist". I'm asking Google to take twenty minutes and fix their solution, which already works but is security through obscurity.
Google seems to have built their brand intentionally to be the opposite of what you're asking for though; and absolutely they could be blacklisted with a simple "GMAIL ADDRESSES NO LONGER ACCEPTED HERE".
>which already works but is security through obscurity.
I'm not sure which one you are saying is security through obscurity here... blah+real.id@gmail.com... or the high entropy mkKAjgsdf788hf87hf@gmail.com, both are obscure, but its a stretch of imagination to start labelling this a security issue.
mkKAjgsdf788hf87hf is not the only possible high-entropy format, it could be if the type that gfycat uses such as "uncommongrimyladybug". That is quite hard to blacklist.
Nobody is ever going to stop accepting gmail addresses, that suggestion is pretty ridiculous. Especially since I suggest that these addresses should be delivered straight to your real inbox (unless they start getting spammed). There's no reason people should stop accepting them.
> Google seems to have built their brand intentionally to be the opposite of what you're asking for though; and absolutely they could be blacklisted with a simple "GMAIL ADDRESSES NO LONGER ACCEPTED HERE".
I think "you can't block GMail" here is meant in the sense that "you can't block the Google crawler". It's certainly technically trivial to do so, but the opportunity cost from lost users will be, for most businesses, unacceptably high.
Excellent interpretation. Gmail = Google crawler. I've made a note of this now.
What needs to happen next is a deep discussion between yourself and logicallee, in the context of Google crawler as well as how to make gmail come further out of the dark ages with high entropy and no security obscurity.
1. Register foo@gmail.com 2. Give out your email address to friends and family as foo+bar@gmail.com 3. Give out your email address to services as foo+{service name}@gmail.com 4. Reject anything coming directly to foo@gmail.com
Maybe there is a way to whitelist keywords and only deliver tags I add one at a time via filters, but it is not the usual interface.
The underlying library is just fine, but a decent front end is lacking.
#include "pdftk/com/lowagie/text/pdf/PdfReader.h"
...
if( input_pdf_p->m_password.empty() ) {
reader= new itext::PdfReader( JvNewStringUTF( input_pdf_p->m_filename.c_str() ) );
PdfReader is actually java/pdftk/com/lowagie/text/pdf/PdfReader.java in the pdftk source distribution. Yes, this is a C++ program that's instantiating a Java class. As far as I can tell, what's actually going on is that all the Java code is compiled to C++-ABI-compatible .o files using GCJ and pdftk.cc links against them, giving a native program that is nonetheless mostly written in Java. Yikes!Perhaps unsurprisingly, GCJ didn't get a huge amount of traction, and it has been deleted from the GCC tree entirely. Good riddance, maybe, but it makes it rather difficult to compile pdftk.
A lot of the scanned PDF ebooks on archive.org use JPEG2000+JBIG2, and the filesize vs. quality difference compared to more traditional formats like JPEG is quite apparent. They do take a noticeably longer time to render, however...
That's mostly due to distinct lack of good JPEG2000 decoding libraries. We're building a PDF renderer library and JPEG2000 is a constant pain int he ass due to it - JPEG decompression is hardware accelerated on many platforms and also has a bunch of SIMD optimized libraries. For JPEG2000 there's practically nothing and due to complexity of the format we count decoding times in seconds for some images even on fast mobile phones.
That is not at all something that would have to be true.
On the other hand, it is possible to make a completely valid PDF and bootable ISO. The first 32KB of an ISO is officially "unused", which is probably why GRUB decided to put itself there, but that can be relocated somewhere else --- the El Torito boot descriptor will need to be updated to point to it --- and the PDF signature (which can be a valid one) and as many objects as will fit can be put in that area, with the rest anywhere else. The xref table can be moved to the very end and the offsets updated to point to the objects.
They are a real pain because they render fine in Adobe Acrobat, but most other PDF renderers (including browser built-in ones) can't render them. Instead they render a blob of interstitial "loading..." text that is also embedded in the PDF (which the XFA rendering would then overwrite). It was a pain to me personally because I had to figure out a way to do programmatic form-filling of some fillable form XFAs, and most PDF libraries don't work with them (they expect traditional AcroForms fillable forms).
But in reading the XFA specification I found it interesting it had its own JavaScript interpreter (including supporting XHR requests as part of some internet-integrated form-filling feature) and another proprietary scripting language called FormCalc. I guess it opened my eyes to PDFs being a container format and the kinds of things they allow you to embed.
chris_st's recommendations are how I learned it.
I remember going though HP and Epson printer manuals, writing down their control escape codes into a xBase table so that our Clipper application could talk to the printers and do the respective formatiing.
Having access to a PS printer would have been a much more positive experience.
When I was actively play with Sudoku programs, I wrote a bit of code that generated sudoku images in svg, (e)ps and and a few other formats. It was a bit fiddly, but not really complicated.
This is not better than paper and pencil, in terms of accessibility. And we need to do better somehow.
My home-brew accounting software generates HTML invoices; I need them in PDF also, though.
It would be good if a way was available for any application to take a full window screenshot, rather than the viewport screenshot.
It's nice to have the web page's text searchable without having to OCR the PNG.
It’s one of the things I miss most when using an other OS.
The level of integration is not even close to the same & it’s disingenuous to pretend it is currently, let alone from 2000.
Although, they say it doesn't show up in every application.
So, as I said, not the same level of integration.
Nothing wrong with those choices (they give the end user more flexibility & control for example) but it is a trade off
It's underused these days, but still available to apps, and they can interchange data in that format. Linux support for PDF isn't anywhere near as integrated.
This is a common misconception. Display PostScript was never present in any released version of macOS. It was replaced by the Quartz renderer, which is rather different.
Quartz can display and output to PDFs, but it does not use PDF as an internal format.
Of course, PDF is intentionally so weird: it was a move by Adobe because other companies were getting too good at handling postscript.
Embedding custom compression inside your format is seldom worth it: .ps.gz is usually smaller than pdf.
it's used everywhere because you can do everything with it.
This also leads to the problem where you can do anything with it.
so each industry is kind of coming up with their own subset of pdf that applies some restrictions in the hopes of making them verifiable.
the downside is that these subsets slowly start bloating until they allow everything anyway.
i'm looking at you PDFa. grr.
It’s simply the most reliable print-ready format. It is a Portable Document Format in every way that matters to the end user
There’s a reason one of the main uses of LaTeX is outputting to PDF
That solution just happens to be latex, especially since virtually all computers will have some way of viewing and printing pdfs by default.
I hate it so much, if it wasn’t for its excellent abilities in specific areas (hyphenation, etc) I’d much prefer CSS.
Yes, TeX makes me prefer CSS for layout. That’s how painful I find it.