I'm not a big proponent of LLM proliferation, but I was thinking that mass review of tons of scanned documents might be exactly the sort of thing they're really useful for. Given an AI that hasn't been ruthlessly tuned to be as politically neutral as possible, you could have a huge database and query it in plain English like "were there any documents that made overt reference to extremely corrupt behavior?"