PDF Forgeries Are Surprisingly Rare (2022)
gwern.net
gwern.net
PDFs are meant to be a presentation document format, not an editable document format. I have some experience editing PDFs, and it's difficult to make significant modifications. Editing small sections of text or replacing a date is fine, but once you try to start editing entire paragraphs or changing the layout, it quickly becomes unmanageable. The kerning looks weird, the spacing between words looks off, and when you look into it you realize the the sentence you're trying to edit is actually split across a dozen text objects. For these it becomes almost impossible to edit without destroying the spacing, so you just give up.
And the approved way to edit these PDFs is via Adobe Acrobat Pro, which costs $20/month (ugh), which I'm not willing to pay for.
So I think this is why we just don't see that many PDF forgeries. The editing tools just aren't very good. Maybe it's different now we can just ask Claude to edit our PDFs for us.
Inkscape is a pretty good tool to edit .pdfs
when you import a regular .pdf, it needs to convert it once, and once saved, you can edit the pdf and work on it as if it was a regular .svg
if you choose to embed the fonts when saving, .pdf files becomes easy to work on, you don't even need to export them.
Even generating them on the fly can be a real pain. I had a workflow when I was out of work to auto-generate a CV based upon re-writing sections to match the spec. So I'd have pre-written blocks and slot them in. And yet, somehow, even then it would break formatting all over the place.
[1] https://www.fincen.gov/mortgage-loan-fraud
[2] https://www.justice.gov/usao-mdfl/pr/licensed-mortgage-loan-...
and all those things listed wouldn’t be prosecuted, as long as you don’t touch mortgages
I would say KYC in general is a joke and nobody cares, as long as FinCEN and the DOJ believe KYC is working. But as far as the peddling of forged PDF’s go, very common, just not at the point of where the more damaging crime occurs ….. except in mortgage fraud, where that crime explicitly considers the fake employment letter to be the crime
If you want to forge a PDF, you do it because you have a private target in mind. Financial fraud, corporate espionage... There are other settings but people don't just share PDFs around for you to see a deluge of fake ones.
1. General public don't read papers.
2. General public have lowered trust to them.
3. It's not impossible to produce and publish genuine paper with forged data.
Students giving out free Microsoft licenses they could access through their University to people they knew was a thing, because Microsoft activation hacks had all kinds of problems.
Not like altering lines or anything material (I can already imagine the breathless replies being typed about how it's "antisocial" if someone's deck doesn't meet setbacks by 1ft or something) but changing text to remove references that would indicate other survey work was done that they don't wanna submit to public record, changing a labeled distance by a couple feet to make something that already exists technically compliant.
Like with everything else it's the little guys sticking their necks out. The big guys can afford good surveyors who "get it" and don't need to be told when a difference within the bounds of measurement error and will have four or five figure implications or that a survey result is so unfavorable that they should just call the client, bill whatever they would have charged as a consulting fee and not screw their client with the liability of a paper trail.
Some software rewrap the page into say additional XObject, some add unique elements, some change one IDs but not the others.
I had a great blog post about it but can't find it right now. Will post if find the link.
Seen as (unsigned) PDFs are basically images of documents, it's wild to claim that there are no forgeries, anyone doing a good ole fake document, fake diploma, fake wire transfer, fake articles of incorporation, etc... is using a pdf for faking. A nigerian prince scam sending an attachment mentioning 15.000.000$ locked up in inheritance would constitute a PDF forgery.
> There is plenty of incompetence, fraud, and malice online, often in PDFs… but only new PDFs. I can’t think of a single fraud accomplished by editing a real PDF & just uploading it for Google Scholar etc. or where I’ve been burned by even mislabeling.
If someone were to share a preprint PDF that I don't believe, my instinct would be to double-check it against the institution associated with its publication. If it was never published then it's basically just a screenshot of a blog post.
Ohhh okay the opening makes more sense now. The blog above is publishing papers in form of blog posts as a protest against scientific publishing. I love the energy, but ruminating on how weird it is that so few people take advantage of the thing that you're weird for doing in the first place feels... obtuse?
Most people just never understand how powerful can be TeX and fear this path, because it takes a long training and the investment in time is huge. But you don't need to master all the details to do simple text replacements or filling a form with your name and address. That part can be learned in a day, and is totally free.
> (In the rare cases anyone tampers with PDFs, it quickly turns into a technical morass and he-said-she-said, and requires several orders of magnitude more effort to prove than to do; consider the hoops epxperts had to jump through in the Craig Wright cases, or even just a landlord editing a contract—where the contract was done through a digital document timestamping service!)
Specifically "epxperts"
The Dieselgate is a good analogy. Incentives, how easy it is to get to profit.
> "Whereas, if you were so epistemically careless with images on, say, Facebook, you would wind up with a folder stuffed full of lying images which have been Photoshopped, claimed to be things other than what they are, ‘deep faked’, etc"
Maybe because a paper is a comprehending medium, expecting the reader to understand its source and interpret the content on a more intellectual level, whereas effective images/videos are usually framed to be understood immediately.
i.e. if I want to convince people that lizard-people live among us and I want to manufacture proof, my demographic probably responds more to images with a lizard than a reinterpreted academic paper...
Can this actually be done without an ISBN or DOI number?
This summer I downloaded a 2014 ACM conference paper from the Internet Archive, opened the PDF, changed the author's name on the title page, changed the email address, changed a GitHub URL in a footnote, rewrote the Author field in the file metadata, and put the altered file on the internet.
The paper is Vanessa Freudenberg's SqueakJS paper (DLS '14). She died in 2025. In 2024 the same paper won the SIGPLAN DLS Most Notable Paper Award, credited to Vanessa. Every public copy of the PDF still had the name she no longer used: ACM, dblp, Semantic Scholar, and every Wayback snapshot of the file on her own site. Same bytes. No corrected edition had ever existed.
She had already said, on this site, what she wanted. In 2021 she replied to me about Dan Ingalls's HOPL Smalltalk paper:
https://news.ycombinator.com/item?id=29125515
>Dan published an updated version of that paper here: https://smalltalkzoo.thechm.org/papers/EvolutionOfSmalltalk.... Would be great if you could cite that one next time. The main improvement for me is not being deadnamed. There are other corrections as well.
Dan had quietly put out a corrected PDF of that paper while she was alive, outside the publisher's version of record, because the official copy was too slow to say her name. I did the same thing for the one paper he could not: hers.
The coding agent (Fable, in Cursor) did the search and proved the negative. It proposed the two institutional paths: petition ACM under the name-change policy, or ask her co-authors to re-typeset from LaTeX. It would not propose opening the published PDF and changing the bytes. I had to tell it to. Two trained rules collided. Do not tamper with a historical document. Do not leave a deadname sitting there. The first one won until a human who knew her pushed.
I then wrote down why that push was legitimate, and put the constraints into a reusable skill so the next agent can mention the option with the rules attached, instead of hiding it. All five, or do not do it:
The subject asked, in their own words, or someone with standing acts on that record. The original file stays bit-for-bit, hashed, in the same directory. Every change is listed in a README and in a Note inside the PDF. The filename says what the file is. The publisher correction is pursued in parallel; the edited file is a bridge, not a replacement of the record.
Drop any one of those, and the same edit is just a silent substitution.
Doing it right is the part Gwern is pointing at when he says there is no everyday Photoshop for PDFs. TeX does not store a name as a string. It stores a TJ array with kern numbers between syllables, so ordinary grep never sees it.
The Times subset on the byline happened to contain every glyph for "Vanessa"; we still had to recompute the V-a kern and shift the Td so the line stayed centered over her affiliation. The sans subset used for the email ended at glyph 117, one code point short of the letter v. That font literally could not spell her name. We set the line in Helvetica, sized to match, and resized the link annotation so the clickable box fit the new URL.
Fable did not stop at the byte edit. It rendered page 1 with Ghostscript, cropped the header, and looked at the pixels with machine vision. The first pass was beautiful and wrong: the byline said Vanessa, the email underneath still said the old address. The next pass put the new email in Times, which rendered clean and was the wrong family. It swapped in Helvetica, recentered, and zoomed to 220 dpi next to the neighboring authors until the line sat quietly with theirs. That loop is why the file does not look like vandalism.
The original and the memorial edition sit next to each other. ACM's copy is still uncorrected. Option 3 of their name-change policy is the canonical fix; her co-authors are the right petitioners.
Her comment:
https://news.ycombinator.com/item?id=29125515
The 2021 thread, preserved:
https://github.com/SimHacker/moollm/blob/main/designs/presto...
Dan's corrected HOPL paper, the one she asked us to cite:
https://smalltalkzoo.thechm.org/papers/EvolutionOfSmalltalk....
The unaltered 2014 PDF (Wayback; her site is down):
https://web.archive.org/web/20250119071632/https://freudenbe...
Memorial edition:
https://github.com/SimHacker/moollm/blob/main/designs/presto...
Original, hashed, beside it:
https://github.com/SimHacker/moollm/blob/main/designs/presto...
Every edit enumerated:
https://github.com/SimHacker/moollm/blob/main/designs/presto...
The case, including the nudge:
https://github.com/SimHacker/moollm/blob/main/designs/presto...
Why the agent would not propose it, and the five conditions:
https://github.com/SimHacker/moollm/blob/main/designs/presto...
How the bytes were changed (Ghostscript and vision at minutes 68-84):
https://github.com/SimHacker/moollm/blob/main/designs/presto...
Renders and kerning:
https://github.com/SimHacker/moollm/blob/main/designs/presto...
The change-name skill (scan, discuss, edit, verify, publish):
https://github.com/SimHacker/moollm/blob/main/skills/change-...
PDF playbook:
https://github.com/SimHacker/moollm/blob/main/skills/change-...
ACM version of record (still the old name):
https://doi.org/10.1145/2661088.2661100
ACM name-change policy:
https://www.acm.org/publications/policies/author-name-change...
SqueakJS, where she moved the repo:
https://github.com/codefrau/SqueakJS
Remembering Vanessa:
https://github.com/SimHacker/WillWrightShowForFood/blob/main...