Google PDF Search: “not for public release”
google.com
google.com
I encounter this problem pretty frequently with System-on-Chip spec sheets. They’re confidential while under development; and still confidential while being floated for sale by private negotiation to private device makers; but once they’re in a production device, those devices will have their own public PR pages describing exactly the properties of the SoC, so the SoC can’t really hide its specs any more at that point.
So, once the SoC manufacturer makes a high-profile sale, they will usually “declassify” the spec... in the sense of allowing the release of the spec sheet. But they don’t usually bother to ever publish a copy of the de-watermarked spec sheet at any canonical URL; they only begin sending the de-watermarked sheet in place of the watermarked one to new customers.
So, when you find one of these spec sheets online, they’re just as likely to be a copy that a device maker got from the SoC manufacturer from before the specs were declassified, which they released once they were told it was okay to do so. And such documents still have the “confidential” / “not for public release” watermark.
Here’s a database that I always keep at hand whenever I need ideas for my projects [1].
INAL. Regarding ToS. In the US, scraping public data is a fair use exemption under the first.
I'm not sure that just because some data is in Google search it becomes any sort of 'public'.
If you add 2019 to the search term you'll get a bit more relevant content. You can also use doc, docx, etc for the filetype.
Be careful. The fact that it's possible to download something doesn't necessarily give you the legal right to.
And if it's sensitive stuff (not classified, secret, top secret, or scif) you can still be forbidden to even provide the link other than to your approved security contact.
It's really better not to know if you have any sort of clearance.
Me neither.
That would have been very wrong.
I was just asking is all.
"top secret" filetype:pdf
The Google summary for my first hit starts with this:
The attached material contains TOP SECRET information which bears directly upon the effectiveness of our national defense or the conduct.
TOP SECRET DISCLOSURE RECORD
;)
It’s just an empty template — https://i.imgur.com/H9u5sYi.png
So maybe this isn't as big a deal as folks want to make it out to be? 𐑴𐑝𐑻𐑤𐑰 𐑛𐑮𐑩𐑥𐑨𐑑𐑦𐑒, 𐑥𐑱𐑚𐑰.
http://www.defence.gov.au/foi/docs/disclosures/054_1213_Docu...
It's both marked as "Declassified" and "UNCLASSIFIED - Document not for public release". Hilarious!
That said, there's some great commentary:
> The potential for the leaked documents to adversely affect the safety of deployed forces, or operational security more broadly, is being assessed by the Defence task force. It appears unlikely at this stage that ADF or coalition forces will be directly endangered by the leaks. The information examined so far is tactical information that is now sufficiently aged that it poses minimal threat
It's interesting to see how despite the apparent "minimal threat" Julian Assange is still under quite a lot of heat from strong political forces over this (& other releases).
EDIT: formatting
These kinds of things are absolutely in the hands of the organizations who are authoring and publishing the documents, not the vendors like myself who provide the tools, but it is still a relief.
If you can detect “publishing smells” like this with a Google search without access to the tools, you also ought to be able to detect and raise warnings about them in tools; quite a lot of tooling exists in software development workflows to catch the programming equivalents of this kind of oversoght and prevent them from getting into deployed code. (And arguably even more similar, there are a wide array of tools designed to identify and protect against internal information being exfiltrated, intentionally or accidentally, via email, etc.)
So while responsibility for the results of what gets published is always the responsibility of the tool users and not the tool vendors, I don't think it's accurate to say that this is completely outside the hands of tool vendors. And, I'm pretty sure customers (especially in the enterprise space) would see support from publishing toolchains for preventing this as a competitive advantage.
Downloading confidential documents, even when they are publicly available through google, can be a legal liability.
There is at least one case in France where a security researcher/blogger faced legal issues after downloading confidential documents that where publicly available when using the correct search terms (see Bluetouff vs French justice).
I don’t know about the situation in the US for that kind of things though.
Or for example when Dropbox disclosed a brief moment when you could sign into anyone else’s account. What if all their links had just broken and had auth removed at that time? If you were to exploit that, are you hacking - or was it the fault of Dropbox? Are you wrong for taking publicly available files? They were publicly available when you downloaded them - but you knew you shouldn’t have.
How do you draw the line?
By analogy, suppose you receive a letter in the mail; you reasonably believe the contents are yours, so you open it. Oops, it was actually mailed to you in error; the contents belong to another person. Obviously you did not commit a crime, but you are obligated now to discard the contents.
For github/gitlab/etc links, they probably also could inform project owners about the problem.
1. Declassified/Released documents.
Many historical documents are no longer classified, yet some had its original classification markings. The same goes for commercial documents. For example, the Trump-Russia dossier was a previously classified document of a private company.
2. Whistleblowing documents.
These documents were released by whistleblowers, thus no longer private and classified. You can find many NSA documents leaked by Edward Snowden hosted on Amazon AWS owned by news websites.
3. Boilerplate classification.
Many commercial documents have boilerplate classification, regardless of whether they are truly sensitive or not. It's very common to find technical documents and datasheets with "Confidential" markings, which should have been removed long ago, probably related to (1).
In my experience, these documents make up 85% of the search result, only less than 15% of the documents in the search result are truly security breaches and confidential.