If they can't tell, someone may now be sitting on a lot of very juicy data, far beyond what may be left in these caches.
If they can't tell, someone may now be sitting on a lot of very juicy data, far beyond what may be left in these caches.
I mean, how could CloudFlare, or anyone, possibly differentiate this from normal scraping/polling/ manual F5 refresh behavior? This sounds like a PhD thesis.
I guess you are asking CloudFlare to quantify the amount of distinct bytes of unauthorized data sent to any particular user agent? But then, any sophisticated attacker would rotate IPs, UA identifiers, and probably even between vulnerable websites, if they had known about this vulnerability.
I don't think it's reasonably possible to rule this out, even with a massive dedication of investigative resources. Like the other commenter said, it's wisest to assume it happened.
If you have perfect information about what resources were requested when, you can look for a spike in queries for vulnerable resources. Once you see that, you know there was an intentional exploit and can start to look at who drove that spike, what was leaked, etc.
The problem is that we're talking about l huge amounts of data. I'm skeptical that CF has lots of sufficient length and detail to conduct this analysis, but have no real knowledge about their forensic capabilities.
Right now they'll be querying for political blackmail and insider material, then they will turn to Joe Schmoe who only gets his news from the television, where no doubt this will never be discussed. They will effortlessly and listlessly categorize the data largely as an automated process.
If we expect anything less, it's possible we are underestimating our adversaries. Better to overestimate them and feel embarrassed later when you later find out you were just being paranoid.
As you say, in the presence of uncertainty it's most prudent to assume that this actually happened.
The reason why I consider them dubious is that anyone simply searching the name of some HTTP headers in Google et all could have stumbled into this. I don't find it at all unlikely to happen in a timespan of 5 months.
So the key question really isn't "how likely was someone to find this", but "how likely is it that Project Zero was the first". I think it's hard to estimate odds, but I'd be surprised if it was even as high as 50%; there's too many teams, individuals, freelancers, state actors, etc. actively engaged in looking for this kind of thing.
The data it reveals isn't guaranteed to be obviously private and exploitable. It can just look like a valid but useless response, or a invalid and corrupted response depending on what you were looking for in the first place.