Getting a full PDF from a DRM-encumbered online textbook
vgel.me
vgel.me
A few years ago I was involved in a project that required capturing content from an online document viewer site that used a SWF for each page, and our product basically did a vector-vector conversion from SWF to PDF. One of our competitor's product did the "render each page as an image, and combine the images into a PDF" and the output difference was amazing. IIRC it was around 2 orders of magnitude in size, and 4 in speed of generation. (They used Java, and we used C, which might account for some of that too, so it's not a totally fair comparison.)
I've done similar (though usually less effort) textbook trickery a few times. The Adobe Inept hack is very handy. Oh, and a recent one was stupidly easy: you could view the ebook in your browser, and save excerpts as a pdf, but only 100 pages in total per book. Problem was it stored how many pages you had saved in a cookie, so "Clear the last 5 minutes of browsing history" and you could get another 100 pages, rinse and repeat for all the book and then staple the files together with pdftk.
PDFs typically are full of Postscript (except if they're just scanned images), which is just a text rendering language. As long as you keep the Postscript format valid, you could remove the watermark by just deleting that text.
I didn't know about PDFtk, but Ghostscript can take a PDF and turn it into text Postscript, and it can reverse the process.
But for normal PDFs PDFtk is incredibly useful!
Has anyone made an index of which colleges require DRM textbook purchases in their courses?
Once I got into the CS courses, most if not all of my professors just provided PDFs of either their own material or some open source textbook they were contributing to.
So then most US colleges? Like spectralblu said, this is common in lower level classes—especially math classes for some reason.
For me, one of the wonderful things about copyright is that works always end up available for free to the general public. A DRMed work will never be free in that sense, and should then not be covered by the regular legal protections.
I say a maximum of 25 years free copyright (i'd rather see something like 5-10 years), and then progressively increasing fees that start becoming crazy after something like ten years.
Then use that money to finance culture.
I think it would be hard to prove that he cracked any DRM.
He didn't download the book 10 pages at a time, and he didn't use an image editor to remove the watermarks.
He wrote a script that simulated navigating through the book with a mouse and keyboard and a browser, and generated a bitmap image of every page.
Source: https://archive.org/stream/GuerillaOpenAccessManifesto/Goamj... (https://archive.org/details/GuerillaOpenAccessManifesto)
There is safety in numbers. The more people do this, the less any particular one will be targeted.
Consider it was not so long ago that almost everyone talked about pirating something, usually with torrents or some other P2P, and nothing happened to the overwhelming majority of them.
That doesn't seem particularly difficult for them to do. I doubt anyone would actually bother, but still, it's not tricky.