The very last people in the universe Cloudflare should be criticizing right now are the Google security team.
The very last people in the universe Cloudflare should be criticizing right now are the Google security team.
The length of time this went for was already disastrous. From the Cloudflare blog post, "The earliest date memory could have leaked is 2016-09-22".
As an example of how destructive it might have been beyond leaked information, note that Internet Archive have spent considerable time archiving the election and its aftermath.
donaldjtrump.com is served via Cloudflare.
Chances are it wasn't a domain that was leaking information - though we don't know as there's no publicly accessible list and no way to tell if they had the buffer overrun features active.
If it was however the Internet Archive now have a horrible choice on their hands - wipe data from a domain that will be of historical interest for posterity or try to sanitize the data by removing leaked Cloudflare PII details.
History will already look back at the early digital age with despair - most content will be lost or locked in arcane digital formats. Imagine having to explain that historical context was made worse as humanity deleted it after accidentally mixing PII into random web pages on the internet -_-
> Chances are it wasn't a domain that was leaking information - though we don't know as there's no publicly accessible list and no way to tell if they had the buffer overrun features active.
As far as I my understanding goes domains that had the features on were leaking the data of other domains. So it's near impossible to tell who was affected.
And seeing as cloudflare has some of the biggest pipes on the net, it would be easy to saturate a pipe to gather as much data as possible from any sites you can find that are affected.
The complication comes with the edge cases they'd need to face and @eastdakota's call to "get the crawl team to prioritize the clearing of their caches"[1]. This is also work they now need to do due to Cloudflare's blunder.
For Internet Archive and Common Crawl, these aren't caches, they're historical information. You can't just blow that away - but you also can't serve it if it has PII in it. Either they need to find all the needles and filter/tombstone them - which we'd expect to be very difficult given leaked information is still sitting in Google and Bing's cache now - or wipe/prevent access to the affected domains.
Wiping the donaldjtrump.com would be historically painful but even temporarily blocking access to the domain would be problematic.
Finally, and most importantly, the fact that non-profit projects need to worry about how to exfiltrate such information is ridiculous anyway. Exfiltrating the information is non-trivial as well and may well be destructive with even a minor bug in processing. Having looked at much of the web, it can be hard to tell what's rubbish and what isn't :)
(Fun example: I had an otherwise high quality Java library for HTML processing crash during a MapReduce job as someone used a phone number as a port (i.e. url:port). The internet is weird. The fact it works at all a mystery ;))
Polluter pays.
Why the help was not 100% appreciated...
They mention using the default mod_security rules. I know mod_security can look at the response body, but it is not a default setting.
They seem to care a lot more about marketing than anything else.
(On edit: which also means they have a fucking cheek whining about disclosure timings.)
Why do humans politicize so much? That's one thing I'll never understand, and one of the reasons why I refuse to become a manager.
But what does it really matter?
Tavis Ormandy linking to leaked data to settle some personal groll is a pretty good example how much they really care about data leaks. I guess when you've already leaked all your customers data to intelligence agencies around the world while moving it between data centers it's hard to keep any form of standards.
Want to actually argue your idea with any sort of competition I suggest going outside HN.
Cloudflare has a much stronger incentive to handle the situation in a way that limits the damage to their reputation. As you note, researchers also have an incentive to maximize the benefit to their reputation. But the incentives for P0 are much more closely aligned to the interests of the public here than are the incentives for Cloudflare. The information revealed by the bug is indeed really bad, and there's likely no possible way to tell which of Cloudflare's customers were affected, much less those customers' users. No company wants to acknowledge "yeah, a bunch of our customers' data was compromised, and we have no possible way to tell whether your data was among the data compromised". Their incentives push them to downplay the impact; to accurately describe the potential impact to their customers would probably be disastrous for them as a business. In contrast, downplaying the impact would be contrary to P0's incentive to benefit their own reputation.