A case study in PDF forensics: The Epstein PDFs
pdfa.org
pdfa.org
There are also other documents that appear to simulate a scanned document but completely lack the “real-world noise” expected with physical paper-based workflows. The much crisper images appear almost perfect without random artifacts or background noise, and with the exact same amount of image skew across multiple pages. Thanks to the borders around each page of text, page skew can easily be measured, such as with VOL00007\IMAGES\0001\EFTA00009229.pdf. It is highly likely these PDFs were created by rendering original content (from a digital document) to an image (e.g., via print to image or save to image functionality) and then applying image processing such as skew, downscaling, and color reduction.
https://www.justice.gov/epstein/files/DataSet%207/EFTA000092...
Is that remotely plausible? I can't imaging faking a scan being easier than just walking down the hall to the copier room.
I mean even in this thread you got what are essentially one-liners to do it.
Definitely less hassle then doing it irl
Yeah they might have used some web converter, but that on the other hand would have been extremely incompetent handling of the secret data.
If they were faking the documents rather than the delivery method they definitely could have invested some time in flawless looks.
Don't attribute to malice what can be attributed to laziness, these are government workers
ROTATION=$(shuf -n 1 -e '-' '')$(shuf -n 1 -e $(seq 0.05 .5))
for pdf in "$@";
do magick -density 150 $pdf \
-linear-stretch '1.5%x2%' \
-rotate 0.4 \
-attenuate '0.01' \
+noise Multiplicative \
-colorspace 'gray' \
"${pdf%.*}-fakescan.${pdf##*.}"
doneNote that you can get random numbers straight from bash with $RANDOM. It's 15 bit (0 to 32767) but good enough here; this would get between 0.05 and 0.5: $(printf "0.%.4d\n" $((500 + RANDOM % 4501)))
But yea, this will work as long as you have imagemagick and Nautilus installed.
Sign a blank paper, scan it, paste the original doc on it. Then keep the scan for future docs.
The only reason I can think of for why someone would want to do this is to pass off fraudulent or AI generated images as real.
Other investigations into the files have found oddities like redaction of the word "don't" indicating a haphazard find-&-replace approach to redaction, possibly LLM-aided.
The DOJ/Akamai online hosted search feature is also incomplete - potentially due to some of these "digitally scanned" files not being subject to OCR.
Possibly but I don't find it compelling, if only because a significant portion of the media reportage on the files has made claims that are entirely baseless - if there were a narrative to be sold one would expect such reportage to be actively leveraging such fraudulent images.
But if you want to do it to 2000 documents...
Was the motivation for this benign (an employee skirting regulations) or malicious?
Who can say what effect it had on the world, but a presidential candidate reposting himself personified as Pepe the frog was still weird back then, and at least a nod to the trolls doing so much work on his behalf
https://medium.com/tryangle-magazine/meme-magic-is-real-you-... (dismissable login wall)
Summary: Trump used memes not in the sense of pepes but in the original (Dawkins') sense of "earworm" soundbites, along with a torrent of scandals, each making the previous seem like old news, to exploit a public tired of the "status quo" into voting for a zany wildcard pushing for reactionary policy
Looking back on it, I wonder if this was priming.
I didn't fall for it. They are still losers, but the encyclopedia dramatica with swastikas looks way way way less funny in 2026 than it did in 2008.
That many serious commentators didn't see this was itself very funny as anything with lots of attention on the internet does become influential! It is funny to a troll to see people pay serious attention to them "I am just a clown and they think I'm serious!". But don't think that they were actual comedians, lol, they are as serious as HN users.
In the dawkins sense of the word: the "meme" wants to spread and grow and the mechanism for it's virality was the immune response to it.
On another angle, the responses also gave the target an identity. Groups get defined as groups from outside more than from within. And it's always a wrong characterisation which also helps define the in group in relation. "You guys are all toxic Linux dude bros" inside: "but some of us love macs and windows, and some of us are girls, they sure dont understand our ways"
I'm just saying, it's a symptom. The crazy found critical mass, broke containment. From there it was laundered in millions of Facebook groups and here we are.
https://www.justice.gov/epstein/files/DataSet%2010/EFTA01992...
If you radicalise the 0.01% of people who are prolific meme creators, you radicalise the masses.
* I did say old...
Of course in 2026 it is apparently fine to break into homes without a warrant and execute protesters. The same people are able to "believe" two literally opposite concepts.
There would be times when you would go to the r/all and half the page would be posts from them.
Not to mention a lot of the organized harassment a lot of the mods/power users of that sub caused in the years after. It was a mess.
Hey quick question, around January 2021, what would happened that caused Trump to be deplatformed? Anything stick out in your mind?
This is a good book about it.
The reason I don’t agree is that moot banned any Gamergate discussion and those people then went to 8chan, a site which moot had no control over.
And it was Gamergate that put some fuel on the fire which (IMHO) increased support for Trump. The 8chan site grew a great deal from it, then continued from that first initial “win”.
- moot was fundraising for his VC backed startup during the years the emails are in, and he was likely connected via mutuals in USV or other firms. These meetings were clearly around him trying to solicit investment in his canv.as project.
- /pol/ was /new/ being returned; the ethos of the board had already existed for a long time and the decision to undo the deletion of /new/ was entirely unsurprising for denizens at the time, and was consistent with a concerted push moot was making for more transparency in the enforcement of rules on the site and fairness towards users who followed the rules. /pol/ didn't start a culture war at this time any more than /new/ had previously - it just existed as a relatively content-unmoderated platform for people to discuss earnestly what would get them banned elsewhere.
We're just not going to talk about that one I suppose?
https://news.ycombinator.com/item?id=33755016
You can also unironically spot most types of AI writing this way. The approaches based on training another transformer to spot "AI generated" content are wrong.
I have no idea if specialized tools can reliably detect AI writing but, as someone whose writing on forums like HN has been accused a couple of times of being AI, I can say that humans aren't very good at it. So far, my limited experience with being falsely accused is it seems to partly just be a bias against being a decent writer with a good vocabulary who sometimes writes longer posts.
As for the reliability of specialized tools in detecting AI writing, I'm skeptical at a conceptual level because an LLM can be reinforcement trained with feedback from such a tool (RLTF instead of RLHF). While they may be somewhat reliable at the moment, it seems unlikely they'll stay that way.
Unfortunately, since there are already companies marketing 'AI detectors' to academic institutions, they won't stop marketing them as their reliability continues to get worse. Which will probably result in an increasing shit show of false accusations against students.
You're assuming the people making accusations of posts being written by AI are from humans (which I agree are not good at making this determination). However, computers analyzing massive datasets are likely to be much better at it , and this can also be a Werewolf/Mafia/Killers-type situation where AI frequently accuses posters it believes are human, of being AI, to diminish the severity of accusations and blend in better.
I'm on Reddit too much and a few times there were memes or whatever that were later on pointed out to be AI. And that's the ones that had tells, more and more (and as price goes down / effort/expenditure increases) it will become harder to impossible to tell.
And I have mixed feelings. I don't mind so much for memes, there's little difference between low-effort image editing and low-effort image generation IMO. There's the "advice" / "story" posts which for a long time now have been more of a creative writing effort than true stories, it's a race to the bottom already and AI will only try and accellerate it. But sometimes it's entertaining.
But "fake news" is the dangerous one, and I'm disappointed that combating this seemed to be a passing fad now that the big tech companies and their leaders / shareholders have bent the knee to regimes that are very interested in spreading disinformation/propaganda to push their agenda under people's skins subtly. I'm surprised it's not more egregious tbh, but maybe it's because my internet bubbles are aligned with my own opinions/morals/etc at the moment.
even when people deliberately try to feign some aspects (e.g. switching writing styles for different pseudonyms), they will almost always slip up and revert to their most comfortable style over time. which is great, because if they aren't also regularly changing pseudonyms (which are also subject to limited stylometry, so pseudonym creation should be somewhat randomized in name, location, etc.), you only need to catch them slipping once to get the whole history of that pseudonym (and potentially others, once that one is confirmed).
but, those changes are usually pretty gradual and relatively small. thats why when attempting to identify someone via writing, you look at several aspects of the writing and not just word choice (grammar, use of specific slang, sentence length, paragraph structure, punctuation, etc.). it is highly unlikely that all aspects of someones writing changes at the same time. simply removing "ha" is inconsequential to identification if not much else changed.
additionally, this data is typically combined with other data/patterns (posting times, username (themes, length, etc.), writing that displays certain types of expertise, and more) to increase the confidence level of correct identification.
But on a serious note, what did "la" mean in your context? I've never seen this.
https://news.ycombinator.com/item?id=33755016
It turns out stylometry is actually a pretty well-developed field. It makes me wanna write an AI browser assistant that can take my comments and stylize them randomly to make it harder to use these sorts of forensics against me
The old trick years ago was to translate from English to different language and back (possibly repeating). I'd be curious how helpful it is against stylometry detection?
If you want to be grouped with foreigners who don't know English, it might work well, although word choices may still be distinctive enough to differentiate even when translated.
That said, best to assume that the various government agencies have tools like this, and better - if you're trying to hide your identity online, don't just change users or go through VPNS/proxies/TOR but change your writing style too.
(Also I'm convinced most VPNs/ proxies / TOR nodes / public access points are honeypots)
Either people on that level rarely write anything on their own and have completely forgotten how to construct proper sentences or maybe that just how they communicate. Sort of language internal to the group.
Some people postes conversations, and comments, but I don't feel like they actually grasp what's being discused and they just latches on to key words.
Why not? Clear motive, matching timeline, mentions of that reddit account in the released FBI documents of her case
Other mods knew them personally and were still in contact. The user claims they heard of the rumor and decided not to reactivate for the lulz.
I am not familiar with the mod side of reddit - couldn't fellow mods audit her mod action logs to find more juicy details we would have heard about by now?
If Maxwell is indeed a spy and doing what she is claimed to do, it is highly unlikely that she'd put her last name and a reference to her specific family's property in her username. This would be a glaringly arrogant choice for someone who had been groomed from an early age for spycraft, and who had any degree of oversight.
If she were part of a spy network, they would be highly remiss not to commandeer the account at the time of her arrest to avoid suspicion unless they were completely incompetent.
I am mostly familiar with cold war espionage so it just doesn't sound like the general MO to me. Unless Opsec or whatever has badly decayed since then. That's not impossible.
The mentions of the account in the files are from anonymous tips, some of which are highly absurd. They vetted a lot of tips, and I saw no information in the new releases indicating they thought it held water. We've seen the subpoena and IP tracking for the Epstein prison guard whistleblower, but no such thing on this topic.
hopefully someone is independently archiving all documents
my understanding is that some are being removed
Take with a grain of salt, obviously.
The author of gnus, Lars Ingebrigtsen, wrote a blog post explaining this. His post was on the HN front page today.
Who paid him?
Who did get paid?
Term, not plural. There was one (1) interceding administration following Epstein's death.
Trump promised that the Epstein files would be released if he was reelected, and then withheld files. Congress passed a bill remediating this, hence the newer tranche of files: https://en.wikipedia.org/wiki/Epstein_Files_Transparency_Act
>and then withheld files.
So did he sign that willingly in the end? Did he have to sign it? Did he cave because he said publicly he would?
Stop making this a partisan issue. It’s not, and nobody that’s not completely biased beyond any rationality will ever see it as such.
Seems sensible...
So, yes that is exactly what happened.
You people need help. Nobody can be sane and that biased.
No, they did it to protect wealthy and influential people, regardless of party.
It happens that such people are disproportionately Republican aligned, there are fewer places to hide this behavior in the Democratic tent - really just at the top - and the current POTUS seems to be very close to the center of it all.
Independently, we are learning that extremely wealthy and influential men often commit sex crimes through shared fixers like Epstein.
This is about huge wealth and power imbalances, no accountability for the wealthy and powerful, and the behavior they get away with as a result.
If there are any "good guys" here, it's Massie and Khanna for shaming Congress into forcing DOJ to release something, even while DOJ does everything it can to avoid/minimize it.
Democrats engaged in the same cover-up and lies and sexual abuse as the Republicans wrt Epstein. Democrats supported ICE and the murder of immigrants and citizens. Democrats supported American imperialism, oligarchy and genocide.
The parties aren't the same. Would we have the same open chaos, violence and instability under a Harris regime? Probably not. Would they have released any of the Epstein files on their own? Also probably not. Voting for the lesser (or more restrained) evil is valid when no good option exists but make no mistake the Democrats are not really a principled opposition party. It's mostly kayfabe.
OCR is so bad of course that decoding the Base64 seems futile without a lot of effort.
Example: https://www.justice.gov/epstein/files/DataSet%2011/EFTA02609...
(More mentioned here: https://old.reddit.com/r/Epstein/comments/1qu9az2/theres_unr...)
Maybe I'm underestimating the issue at full, but isn't this a very lightweight problem to solve? Is converting the images to lower DPI formats/versions really any easier than just stripping the metadata? Surely the DOJ and similar justice agencies have been aware of and doing this for decades at this point, right?
Another guess is that perhaps the step is a part of a multi-step sanitation process, and the last step(s) perform the bitmap operation.
But isn't it a contiguous sequence of data whose length is determined by the container format?
But I agree, presumably the image data part of the file is well and exhaustively defined. I would be very interested in counterexamples that have practical consequences.
Note that there will still be concerns about stenography and fingerprinting which would warrant such a disclaimer from the creator of a tool aimed at a nontechnical audience.
Some of the gathered data is shown here, right? Probably not all.
Now ... that's static information though. That's not really an analysis, most definitely not an independent (open ended) analysis. And it will only show a very incomplete part of the full picture.
This is why I think the "release the files" movement, as good as they are, seems incomplete. I'd rather know a lot more about how they operate their networks, getting away involving underage women. How about secret services of other countries? Should that not also be highly important? So why is there not really a larger investigation as well as independent analysis? Those .pdf files alone can not tell the whole picture. That can just be the tip of the iceberg; and it evidently involves other countries too, with Prince Andrew being the most famous here (aka, the UK, but we already saw that other countries also have similar issues where people suddenly had to step away from politics when it was found out they visited the party-locations of Mr. Epstein).
That said, in my opinion they are using "pages" as the metric because it makes the number sound huge.
(But seriously, great work here!)
I personally understand a year in the submission as a warning that the article may not be up to date.
I'm not used to typing it yet, either.
One thing that is telling about the Epstein case study is how long it has stayed in public view. Pizzagate, which involved more powerful people, was shut down faster than I've ever seen for anything else. I still remember and have archived the more extreme content it's sick.
Is the scope at least limited somehow? Generally I favor transparency, but of course probably the most important parts are withheld.
https://en.wikipedia.org/wiki/Epstein_Files_Transparency_Act
An act of congress, for one.
Also, AFAIK, federal privacy generally ends at death, as does criminal liability; so releasing government files from a federal investigation after death of the subject is generally within the realm of acceptable conduct.
It seems unlikely you lose all rights when you die or it would be chaos - imagine all the secrets people die with that affect everyone they know. An integral part of every estate plan would be incinerating records. Wills do have real power.
OTOH, there's a 2004 case, National Archives & Records Administration v. Favish[1], which establishes the surviving family's right of privacy to death scene photos, but that's technically not privacy of the deceased.
[1] https://www.justice.gov/archives/oip/blog/foia-post-2004-sup...
(It also surprises me that this passed anyway, given that both sides of the aisle seem to have people with clear reason to keep it covered up... ?)
(Also, Maxwell is specifically named, and is still alive... ?)