Elsevier embeds a hash in the PDF metadata that is unique for each download
twitter.com
twitter.com
1. Download the content with N accounts, preferably from different networks.
2. Run your watermark removal tool on each downloaded data independently.
3. Check if the processed outputs are bit-for-bit identical.
Have fun writing watermark removal tools.
Your tool can be as crude as running the document through pdftotext and only keeping the text output. It can just throw away all these possible side channels that are not relevant to the actual content.
Others mentioned other more sophisticated side channels, as rewording sentences, but I'm sure authors would not welcome that.
So if you have a watermark that associate files with the time they have been downloaded, and download it from multiple networks etc. at the same time, and remove the diff from these versions, you'll still have the time-based watermark in the resulting file, so you'll leak when you downloaded it.
They might still contain data that groups those multiple accounts together though. Eg everyone in X university gets the same watermark.
I had the case a few months ago very frustrating. The PDF was encrypted, I couldn't remove the watermark.
From what I understand Adobe Acrobat might have done the job, but it's an expensive software and I refuse to run cracks anymore.
It depends on the encryption being used. If it is coming from Abode Digital Edition, then DeDRM plugin need your credential to decrypt it and convert the file without encryption. It needs your credential since ADE ties the encryption to your credential, so it need that to get the key to decrypt it.
First setup will be a while but once it is properly setup, it will takes seconds to decrypt the encrypted/DRMed PDFs. Even you can print the file without worrying if it will be in garbled mess.
I wouldn't mind them protecting their work if it didn't harm readability and if the report was worth the money I paid (but it's another problem).
Probably not. I bet they're just "allowing" you to have the file.
Though that's my moral standpoint, from a legal standpoint the situation may well be FUBAR, at least in the U.S.
https://gist.github.com/sneakers-the-rat/172e8679b824a3871de...
So when they asked my help I told them, you don't have a copyright problem, you have a SEO problem - since people looking for their publications could find them easily online with the stolen copies sometimes appearing higher on SERP than their own. They said they would address that and they somehow came to terms with the fact that the past documents are lost, but they would like to avoid this situation in the future.
I checked some of their publications. They were moderately priced and the content was quite interesting. Many of their customers were very supportive - actually they were sending reports when they noticed the pirated copies. And they weren't saying "Why should I pay for your books when I can get them online?" but rather "Please take care of protecting your content since we want you to survive and publish more books". So, to answer your question, I was very happy to work for them, and I would do it again.
[0] I was told they changed their approach later and started to collaborate with publishers.
Ah, Scribd, the scummiest company YC has ever had any connection with.
If my email address embedded in the PDF enables that then I think that's reasonably balanced ethics arithmetic.
(edit - I consider public-funded scholarly research to be a different matter to private purchases of commercial books such as fiction or trade textbooks)
Do that randomly, combine the products, and you might get enough entropy to create unique fingerprints for each download. Randomly do that, combine the products, and you might get enough entropy to create unique fingerprints for each download.
(This silly example can create 4 unique fingerprints)
In the most extreme case, imagine if this was a book of poetry.
Those can be "removed" by rendering to text and regenerating a PDF, though. Or even with print + scan + OCR.
Neither are trivial, but doable.
But in this particular point it wasn't even necessary. After a couple of months it turned out that the person who had been uploading the unprotected versions made a mistake and was located as a 20-something living with his parents in a small house in the East Coast. It was enough to notify them and the malicious activity stopped. The company wasn't interested in extracting every penny from the the kid (or his poor parents), they just wanted him to stop, and one letter from a lawyer was enough. If they wanted to go full steam, they would have involved the police and I'm sure they would have found quite a lot of incriminating evidence on his computer, but they were clear ruining someone else's life was not their aim.
Our ebook market is so fragmented that there's no DRM solution that all (or even most) e-readers work with. If you add in smartphones and car stereos (for audiobooks), the situation gets even worse. Therefore, most publishers use watermarking instead of DRM, usually giving you Epub, PDF and Mobi, which you can read on any device you want.
The most common form of watermarking, at least when epub is concerned, is a 1px by 1px div containing a nonsensical hex or base64-encoded string. There are rumors of watermarks in cover images and even the text itself, though. Apparently, sometimes spaces get removed or extraneous spaces are added, lines get split up a little differently, or some common spelling or typing mistakes are made. Considering the huge number of alteration points you have, let's say 10000 per book, even doing two alterations lets you uniquely watermark 10000^2 copies.
Similar things are done to audiobooks, whether by modifying the audio itself in imperceptible ways, or by modifying the internal structure of the mp3 files. From what I've heard, messing with how frames are laid out and what's in-between them is a common tactic.
If you were serious about piracy and wanted to release books en masse, you'd probably use stolen credit cards, stolen accounts or something of that sort. I don't think that's the goal here, though, those watermarks are mostly for deterring casual piracy, sharing books with friends and so on.
While you can apparently just strip it from the metadata properly like suggested on twitter, maybe a "low level" approach like comparing the files on a binary level and setting any bytes that differ to 0 would be more robust. It would still work if they move the hash out of the meta data into the document itself. Only downside is that this requires the hash to be fixed size.
Why can't the algorithm just keep whatever is common between the two files?
Or just take the shorter side of the diff and zero it out?
Such a tool would need a lot of smarts to work reliably enough. At this point, I feel like a metadata stripper that understands the various watermarking methods may be easier to write.
Elsevier could easily make this completely non-trivial. For a very silly example, imagine how this would work if each user got a PDF using a different font. How do you automatically normalize the difference between Comic Sans and Times New Roman? Sure, it's trivial to write a tool that understands what "fonts" are and does this (especially for PDF), but you can't do it with a simple binary tool.
And of course Elsevier can do something entirely more complex.
Bad bad bad bad, again bad idea. That changes the structure and the meaning of the content. A single word replacement can change the context of the entire sentence which can change the content of the entire paper. Changing it can create unintended effects which can make Elsevier to thrash their reputation and universities will move on to different scientific/academic journals site.
If Elsevier tries this method with peer-reviewed papers, then it have to go through the reviews again to ensure that the original and the revision have the equivalent expression which is difficult to do. Authors chose those words and structure to convey their expression in those papers. They chose it for a reason and Elsevier is not going to risk their reputation to change the authors papers without affecting the content of the paper.
I recommend having a tool that works on a single source, then verify that it produces the same output from multiple sources.
Also when downloading multiple times, try to do that from different public IPs and accounts.
If you can donate to SciHub, and if you're publishing look for open alternatives to Elsevier's claw.
Copyright doesn't really expire in some places, or after a very long time. I believe this is way more damaging to progress. Especially for publication, since science works better in a tight feedback loop (and it doesn't work as well if... authors die before others can reply to their papers).
When I think about it doing it well might mean journalists or scientists get caught doing something I morally agree with
yeah, I confused these two
Instead of a single per-user unique value, I could use several values that track different groups of users. The set of values together would uniquely identify a user, but for any 2 PDFs there would be at least one shared group value that would exist in both.
Using your method, leaking a single PDF would identify a group containing the 2 users of the PDFs you compared. If the groups are randomized for each new article, every PDF you leak would further identify you as the common member of the leaking groups.
* anything military
* Elsevier
Hope I'm not giving any ideas to Elsevier and all the other greedy publishers with this ;-)
As, along the same lines, could the original be "scrubbed" for storage so there's no "paper trail" of you having received it?
I m considering that elbakyan has enough credibility to start her own series of open source journals
Could make a nice research project :)