A motivated publisher could embed codes by altering in subtle ways the differences in distances or color between adjacent characters, so that they would survive most color or grey scale conversions; a seemingly innocuous frame drawn around a photo could be either larger or smaller by say one millimeter, representing de facto a bit, therefore using enough pages they could identify a book among billions.
Unfortunately there's no way to be 100% sure that a complex document doesn't contain some form of embedded code.
You could try and break this by adding something like some random noise or jitter maybe, or slightly transforming the proportions of the pages perhaps, or shifting colors in a stochastic way would probably complicate their efforts. The frame around the photo will no longer be exactly 1.1231 mm and will throw off their embedded code reading systems. The colors won't be the same hex codes they are expecting and won't be shifted evenly. Spacing is now all off between the characters.
Good information hiding and watermarking doesn't get affected by common transformations. Most changes will be relative to other content, so noise and resizing shouldn't impact it, especially if there's redundancy in the fingerprint codes. It's not "frame is 1.1231 mm == 1", but rather "frame is slightly wider than average of other pages == push 1 into FEC".
How would be able to get past stochastic transformations? "frame is slightly wider than average of other pages == push 1 into FEC" could be stymied by making pages randomly wider or narrower, so now the average is different and the frame you are expecting to be slightly wider may even be slightly narrower than the average now, garbling your encoding.
Sure, but you're taking two things for granted: you know this is the approach used, and it's the only approach used. If we assume those, you can work around any watermark.
Reminds me of printergate. Fair. Ok, so what about using an OCR tool to convert to text, then converting that back to PDF?