Quantifying the Impact of “Cloudbleed”
blog.cloudflare.com
blog.cloudflare.com
I'm not surprised at this approach as the CEO has a background in law. I think that, combined with his hubris, shaped the tone of this post from an objective post mortem to a goal of minimizing the damage to Cloudflare itself.
I met Mr. Prince once at a tech conference. I mentioned that I followed him on Twitter, to which he freaked out a little bit and said "Oh I hate that when I run across people who tell me that". When I heard him tell his name to someone else there, he said his last name was "Prince, like King". I suspect that his attitude is prevalent at Cloudflare, and may have shaped not only the source of this issue, but the response to it as well, since CEO's generally influence the culture.
I don't quite follow from your description of his reaction. What sort of attitude were you attempting to describe/how'd you perceive his attitude to be?
That is how I understood it anyways.
Downplaying how bad it was that customer cookie's were leaked is what I found particularly egregious in the original response. A session cookie is almost as bad as a password leak. Usually it'll let you everything except change your password (as most sites require your current one as an extra validation).
Plus rotating the master session key (to force the issue of all users being reset) requires knowing you should do it. By downplaying the issue, CloudFlare is sending the message that customers don't have to do anything.
This statement is making an assertion about the behavior of vastly different software backends and the skill-set of the people administering the systems. Things that may appear trivial to you may not be for the sysadmin in charge of servers running healthcare applications that have an opaque internal state.
I don't feel like many people are actually concerned of the implication of having an internet that isn't an internet anymore, but merely a handful of big companies hosting everyone.
Or maybe it's me who don't understand.
Exactly how do you figure this? There are lots of alternatives to S3:
1. All the cloud stuff, which is the new fangled hotness (3 to 4 companies)
2. Non-cloud hosting providers (hundreds of them, from shared hosting to VPSs and above, and usually cheaper if you don't have high traffic)
3. Hosting your own stuff on your own hardware, which is the cheapest if you really need it, and can be done in a data center just about anywhere in the world with good internet.
So, what planet is this that you live on where 3 to 4 companies handle all hosting and we have to use them?
It's practically impossible to defend yourself well in a lawsuit against any respectably-sized company without a few million lying around, first of all. That affects everything, and big companies use that fact to bully upstarts and other small innovators into shutting down all the time.
My own business was effectively shuttered (had to stop selling our primary product) by a C&D from a Fortune 100. It would've been 5-10 years and ~$5 million to see that case all the way through, and under current precedent, it's very likely I would have lost.
Aside from that, industries frequently get laws and rules put in place with ostensibly-reasonable rationale, while the actual intent is to make it virtually impossible for disruptive competitors to enter the marketplace.
The CFAA is the piece of legislation that primarily enshrines entrenched players in the online space. We also need to reform copyright law and clarify some matters regarding the applicability of EULAs, especially with regard to clickwrap and browsewrap.
Once that's done, the flood of innovators that have been held back by big companies dispatching their law firms will finally be able to contribute, and the internet's competitve landscape will truly be back in the hands of the users. It will shift from "Who has my data? I have to go with them" to "I can use any interface I want to access that data", effectively resolving the chicken-and-egg effect that imperils any potential competing social network (not even Google could compete with Facebook on this!).
Large corporations keeping a bunch of lawyers on their payroll are the equivalent of large nation states stockpiling nuclear weapons.
"To get a flavor of how thoroughly the federal government managed competition throughout the economy in the 1960s, consider the case of Brown Shoe Co., Inc. v. United States, in which the Supreme Court blocked a merger that would have given a single distributor a mere 2 percent share of the national shoe market."
I have limited experience with AWS, but they always seem to be having problems that ruin everything that they don't acknowledge. That's what I get from people at the office who have to use it to support a particular customer. Nothing we do ever seems to run properly on it, even though our software is fine everywhere else. We're trying to push that customer off to using their own hardware, and they agreed with us that it's necessary but for different reasons.
Cloudfare, I don't even understand why anyone would use it. I understand the benefits it claims to provide, but you can get those without having an external MITM proxy that can spew information all over the Internet with one bug.
You're looking at the past through rose-tinted glasses. We learn all the time, and we will probably also learn to build some resilience into our systems against these issues. But as someone who's been through a thing or two (see above) on the Internet of yesterday, I like the one we have today.
Slashdotting was mostly a problem caused by Apache's incredibly inefficient design. It consumed huge amounts of memory per connection at a time when most of us had very slow connections. A link from Slashdot was, in effect, a Slowloris attack on your server.
The big change was moving from a fork/thread-based webserver (Apache) to an event-based webserver (nginx), which was made even more efficient by kernel features like epoll.
I still see this issue happening daily here on HN.
The problem with "Slashdotting" was the number of concurrent connections. Heck a fair portion of the time it was the database that keeled over first, not Apache.
Slowloris attacks send purposefully incomplete requests and hold them open with additional headers. Even with dial-up modems, connections were never slow enough for this to be a problem with actual requests, which are lightweight.
Responses are heavy and can tie up slow connections, especially if they have to go get stuff out of the database. But in that case it's no longer a Slowloris type attack. It's just too many concurrent connections.
The Slashdot effect was solved with static HTML caching, simply because caches are faster and don't touch the DB. Cloudflare is a simple, free example of such a cache, although certainly not the only one.
I didn't say it was a Slowloris attack. I said Slashdotting was "in effect" the same thing. Which it is, both problems are one of exhausting limited concurrency.
> The problem with "Slashdotting" was the number of concurrent connections.
Exactly the same problem a Slowloris attack exploits.
> Responses are heavy and can tie up slow connections...
Yes, responses tie up the limited number of available httpd processes.
The problem was that Apache couldn't even serve static files to many clients because of its heavy weight httpd processes and the fact that clients were so slow.
If your web server can only handle 200 concurrent connections, and you want to serve a 500 KB screenshot of your 1337 Linux desktop to clients that download at 3.5 kbyte/s, you can handle like ~1.4 req/s. Doesn't take much to get Slashdotted.
Whereas, event-based webservers could handle at least 10x more connections on the same hardware even before epoll existed.
I had this problem in 1998 and fixed it with select/poll based servers, and then eventually other epoll-based servers before nginx existed.
I guess I wrote what I did because the comparison to Slowloris seemed to over-emphasize the importance of handling high numbers of concurrent connections, since that is the only mitigation for Slowloris.
But, for a flood of real traffic, concurrent connections and throughput are related. The faster your web server can serve responses, the fewer concurrent connections it will need to handle. And as the percentage of dynamic DB-backed sites has increased over time, so has the value of caching. Basic page caching can speed up a Wordpress blog by hundreds of times for unauth'd users, for example. For most little sites, implementing caching will get them more than installing nginx.
And really, what good are valid concurrent connections if the throughput isn't there? For most users, a site that waits 5 minutes on a blank page is no better than a server that's down.
KeepAlive OffS3 going offline for five hours probably caused orders of magnitude more damage. The reason is simple: the stakes are higher today than they were 20 years ago. Back then the overwhelming majority of the business in most companies was conducted on paper and in person. Losing a website for a few days, or even a few weeks, wasn't a big deal. Today though? There are so many companies that host business-critical operations and infrastructure in the cloud that it's hard to fathom how they are coping with being taken offline for several hours in the middle of the week.
Even though two particular problems affected a great many sites, they didn't affect every site or even a majority of sites. That's because we don't actually have a "handful of companies hosting everyone."
Maybe I'm not a typical user, but the only site that was affected by S3 that I use was HN. The only three sites affected by Cloudbleed that I use were HN, Discord, and another smaller site elsewhere. The rest of the internet worked perfectly fine for me.
I'm pretty sure 90% of the sites typical users were unaffected also. I mean, facebook didn’t go down, did it? People could still search on The Google. I don't know of anyone whose job was affected by this aside from people who run websites. Neither of these were as bad as the Dyn outage, which wasn't the apocalypse either.
I agree that over-centralization of the Internet to be concerned about actually creating, and it's definitely something you should consider before choosing services like Cloudfare and S3. But I don't think it's so bad today that the Internet isn't the Internet anymore.
Judging by how frequently I read this sort of comment on HN, there are definitely many people who are concerned here.
Sure, there has been consolidation in the industry, but it is much easier to exist on the internet now than ever before.
What remains now is an accounting of how some of the most sensitive C code on the Internet was tested prior to the discovery of this bug (by Cloudflare's own accounting, the underlying error seems to have been present long before it became symptomatic), and what they're doing about that now.
How would you classify people using all of the crawled data though? It seems sort of hand-wavy to claim nobody deliberately exploited the bug if they used leaked session tokens in crawled data to get access to user accounts.
For example, my Cloudflare-using site is in the 1B to 10B requests a month bracket, meaning 112–1,118 anticipated leaks per Cloudflare's calculations. It's enough that you'd expect to possibly find some in the search engines, but we haven't.
One potential explanation is that our usage is heavily concentrated within a geographic area (the Nordics), and this just isn't where the big search engines are doing their web crawling from.
We also didn't take into account that the distribution of the ~6,500 sites that triggered the bug weren't even across our infrastructure. We distribute load for any site across some fraction of all the servers in any given PoP. Based on whether you're on a cluster of servers with more or fewer vulnerable sites you'd have been more or less likely to leak data.
For the general case the stats are directionally correct and, if anything, conservative (we believe, if anything, overstate the risk). But a particular customer's situation could deviate from the expected probabilities for a number of reasons.
With CloudFlare all the requests are passed through the CDN. With dynamic, authenticated requests the benefits of this are dubious. The responses can't be cached anyway, an additional HTTP(S)-level hop is required for each request (and many additional IP-level hops). The drawbacks are now obvious: sensitive, unencrypted data of thousands of different sites resides in the same, not isolated, shared process memory. With such architecture we will always be one C pointer manipulation bug away from another leak.
Look at what happens when anyone gets hit with a large scale DDoS attack these days, the first thing people will tell them, myself included is "CloudFlare". How do you handle a massive spike in load on a tiny, usually empty server? CloudFlare is the go-to tool for this purpose, especially for people who can't afford to pay Akamai, BlackLotus, Amazon or others. And they're impressively cheap for small businesses too. You can't hide from a DDoS attack or save your tiny DigitalOcean box if your web server isn't hidden away behind a service like this.
I have nothing against CloudFlare, they've proven themselves extremely trustworthy in many ways and even in handling of past security issues they've been very responsible. But this time that's not the case - the downplaying of the issue is really damaging their reputation in my eyes. I'm a paying customer and I'm quite disappointed to see this.
You can't cache responses to authenticated requests, so in case of DDoS attack against an HTTP end point that requires authentication the best CloudFlare can do to safe the back-end is to drop requests, which makes the attack successful. I really can't see benefits of passing authenticated traffic through CloudFlare.
> the best CloudFlare can do to safe the back-end is to drop requests, which makes the attack successful
It can also drop the attack requests, which is generally not from authenticated users, and pass the real traffic to the backend.
Only in the case of real user, authenticated traffic is CloudFlare not a solution. In cases of high unauthenticated user or high fake user traffic or in cases of attacks not operating on HTTP at all, CloudFlare will solve the problem perfectly. Some of the nastiest attacks, like reflection attacks, don't even touch HTTP at all but will quickly knock most servers offline - even lead many server providers to nullroute you to protect their other customers. It'll also get you out in cases of high real user, unauthenticated traffic - like being posted on HN, reddit, etc.
Even without caching, it's useful to run traffic through frontend servers that are close to users. Terminating SSL as close to visitors as possible speeds things up. Getting visitors off their ISP network and on to quality backhaul helps too.
https://news.ycombinator.com/item?id=13719455
FWIW, as an example, Grindr uses CloudFlare: do a Google search for "authorization: grindr3" and you will find a URL which (no longer cached but you can still get snippets) contained an authenticated grindr request, which would be enough to have had temporary access to that person's account.
https://costumla.com/wild-west-costumes-for-men.html
"... U Edge Certificate Authority1 0 U San Francisco1 0 U California�1p ���U� N��U � � �f"�0! %���U@ @GET /v3/profiles/[REDACTED] HTTP/1.1 CF-RAY: 33282514b8d957d7 FL-Server: 15f76 Host: grindr.mobi X-Real-IP: [REDACTED] Accept-Encoding: gzip Client-Accept-Encoding: gzip X-Forwarded-Proto: https Connect-Via-Https : on Connect-Via-Port: 443 Connect-Via-IP: 104.16.85.62 Connect-Via-Host: grindr.mobi CF-Visitor: {"scheme":"https"} CF-Host-Origin-IP: [REDACTED] Zone-ID: 22252132 Owner-ID: 2607399 CF-Int-Brand-ID: 100 Zone-Name: grindr.mobi Connection: Keep-Alive X-SSL-Protocol: TLSv1.2 X-SSL-Cipher: ECDHE-RSA-AES128-GCM-SHA256 X-SSL-Server-Name: grindr.mobi X-SSL-Session-Reused: . SSL-Server-IP: 104.16.85.62 X-SSL-Connection-ID: 15d1b8de6025864d-DFW X-SPDY-Protocol: 3.1 authorization: Grindr3 ... accept: application/json user-agent: grindr3/3.0.13.16790;16790;Free;Android 6.0 CF-Use-OB: 0 Set-Expires-TTL: 14400 CF-Cache-Max-File-Size: 512m Set-SSL-Name: grindr.mobi CF-Cache-Level: byc CF-Unbuffered-Upload: 0 Set-SSL-Client-Cert: 0 Set-Limit-Conn-Cache-Host: 50000 CF-WAN-RG5: 0 CF-Brand-Name: cloudflare CF-Age-Header-Enabled: 0 CF-Respect-Strong-Etag: 0 Set-Proxy-Read-Timeout: 100 Set-Proxy-Send-Timeout: 30 CF-Connecting-IP: [REDACTED] Set-Proxy-Connect-Timeout: 90 Set-Cache-Bypass: 0 Set-SSL-Verify: 0 CF-Force-Miss-TS: 0 Set-Buffering: 0 CF-Pref-OB: 1 Set-Keepalive: 1 CF-Pref-Geoloc: 1 CF-Use-BYC: 0 CF-IPCountry: [REDACTED] CF-IPType ..."
edit: I have spent the last fifteen minutes pulling the search snippet (edit: and now an hour; but all on my phone, so this is harder than it should be, and I also am distracted by other stuff). The way you do this is by walking through the parts you can see to get nearby context. (I do this a lot to pull content purged from or inaccessible to or simply updated in Google's cache.)
In so doing, while I haven't been able to still have access to the session key (likely too long and unique for a snippet), I have pulled the X-Real-IP address field of this Grindr user and a profile identifier they were checking out (both of which I redacted above, but you could trivially get yourself now using that context).
CloudFlare: if you think there isn't private data that was leaked, OR EVEN PRIVATE DATA STILL ACCESSIBLE, you are a bunch of fucking idiots. 1) Clearing the cache isn't sufficient, as for anything built from short plain text words we can pull the snippet. 2) IP addresses and session keys count as "private data". 3) GET requests actually are often sensitive information :/.
I personally saw an oAuth bearer token for a user of Fitbit. It was there clear as day.
Maybe I saw the only one that leaked, but that seems...unlikely.
The FitBit stuff I recall seeing other day was OAuth1. It would look like oauth_token=XXX and oauth_consumer_key=YYY, along with oauth_signature=ZZZ. If that's what you're seeing, it's not a bearer token and that data isn't secret. Those values are essentially record locators.
OAuth1 has an undisclosed "consumer secret" (embedded in the app) and "token secret" (acquired at login) that are used to sign the request. So these requests don't leak any usable tokens. If you also see oauth_nonce and oauth_timestamp values, then the signature is also protected from a replay attack.
The response body of the request that originally acquired the token/secret would leak these secrets.
I'm sure there are bearer tokens out there for other services, but if fitbit is using OAuth1, you'd have to snag that original token acquisition to call the api as that user.
Edit - I hope this doesn't come off as argumentative. I just want to point out the good design choices made by the guys that weren't actually compromised by this. OAuth2 is really easy to implement, but it kinda feels like a step backwards security-wise.
Fitbit uses oAuth 2
No, no it isn't. There are about three proper ways to implement it. Almost nobody whose site I've come across that uses OAuth2 has it properly implemented.
Token revocation and authentication is hard, folks.
They even have a paragraph talking specifically about cookies in GET requests:
> This is not to downplay the seriousness of the bug. For instance, depending on how a Cloudflare customer’s systems are implemented, cookie data, which would be present in GET requests, could be used to impersonate another user’s session. We’ve seen approximately 150 Cloudflare customers’ data in the more than 80,000 cached pages we’ve purged from search engine caches. When data for a customer is present, we’ve reached out to the customer proactively to share the data that we’ve discovered and help them work to mitigate any impact. Generally, if customer data was exposed, invalidating session cookies and rolling any internal authorization tokens is the best advice to mitigate the largest potential risk based on our investigation so far.
For all we know they cherry picked the responses they tested from a single site that doesn't handle anything sensitive.
I don't think you understand how Cloudbleed works. It doesn't matter what site they picked; every single vulnerable site can leak the exact same info. It's literally impossible to cherry-pick that data.
1. Information where improper disclosure is illegal. For example, health records. Or credit card details, which would presumably be a violation of PCI DSS.
2. Information that can be actively exploited, but can also be fixed so the previous disclosure is harmless. This means passwords, authentication tokens, etc.
3. Information that is merely private in nature.
Cloudflare is focusing on the first two items. The third one is hard to quantify; what one person may consider private, another person might not care about. And there's not much that can be done about this kind of disclosure (beyond scrubbing caches which they're doing anyway). Also, it's difficult to automatically identify this type of content (whereas cookies, passwords, credit card numbers, etc are pretty easy to detect) and Cloudflare probably doesn't want to have their employees spending their time reading through all of the cached data they can find looking to see if there's private info (both because it's a lot of work and a huge waste of time since it won't affect anything, and because it's private info; chances are nobody's going to see it normally, and having employees reading your private messages doesn't help anyone and is a violation of your privacy).
I don't know how they could determine how many health records were leaked without looking at and classifying the data, so they've already gone ahead and done that. They presumably know how many steamy Grindr messages were in there.
Additionally, there are laws about messages as well. Email laws generally don't specify smtp only.
I wouldn't call the disclosure harmless. It's unknown if anyone made use of the leaked information before Cloudflare knew, so accounts should be treated as compromised unless it's shown otherwise.
Also, leaking user credentials to any system that handles payments and health info would also breach PCI/HIPAA . This broadens the scope of systems effectively breaking the law.
Another thing to keep in mind is that many(most?) token based authentication systems don't invalidate tokens. So any tokens captured will be valid until they expire, and they can't be "changed" without invalidating every outstanding token (changing the server key)
> Another thing to keep in mind is that many(most?) token based authentication systems don't invalidate tokens.
In my experience, changing your password generally invalidates all outstanding tokens. And yes, this does mean invalidating all of them instead of just the leaked one, but that's not usually a big deal.
I have some pretty fundamental moral issues with this position. Cloudflare is measuring the risk that you were affected by something they can fix, not measuring the risk that you were affected by something you consider important. They are absolutely downplaying the importance of all of this private information: saying "this isn't to downplay" is like adding "we have no affiliation with" at the bottom of a massive trademark violation: it means nothing. That something is difficult to measure and might be impossible to fix it does not somehow make it unimportant.
Let's translate this: we find that some company has been dumping a bunch of random chemicals in our water supply. They respond with an attempt to make us not be concerned, but they concentrate their analysis on a handful of toxins that can be "corrected" with an antidote or a chelation or some other form of direct mitigation in the water itself, while giving lip service to something which merely increases your cancer risk and carefully not mentioning the stuff that will just make you violently ill for a couple days, as it is difficult to know what will cause that, it is a dosage mediated effect, and they can't fix it.
Meanwhile, people keep reporting that they are measuring the water coming out of their tap and keep finding junk in it, and some of the stuff they are finding would make people sick if they drank it. Are you seriously telling me that you think this kind of statement is legitimate?
I am going to maintain: these Cloudflare headers are themselves private information. Even if you ignore Authorization headers, just knowing what URLs people are browsing is something that we would normally consider to be a serious problem... I mean, all BEAST leaked was the size of the files you were downloading, and people take that seriously: the fact that there are even worse potential problems shouldn't distract us from the base issue.
Whether information is merely private is a judgment call that we may be ill equipped to make since we lack important context.
Being outed as gay may range from merely inconvenient for my coworker where friends and family already know and folks from the company have much evidence to assume so to life-threatening for a person that's living in a state where LGBT* people are actively prosecuted by the state and live under the threat of death penalty for merely living their life. Anything in between is possible.
Thankfully CloudFlare spent a week cleaning up the leak in search engines and caches before publicly announcing the issue, so a lot of the evidence is gone.
Just think of how easy it would have been to exploit this bug: buy a load of random domains, host some plausible looking but malformed content on it then setup a few thousand bots to hit those sites and harvest the responses. Keep that plugging away for a few months and you'd easily collect a lot of supposed-to-be-private info.
That would have been an order of magnitude worse than having to sift through caches for scraps. Not to say that what's happened is 'good', but it could have been a whole lot worse.
At this point, these kind of problems are on us as a community that we keep using unsafe tools. Every time we choose one of these languages we are implicitly trading security for performance (a.k.a. money).
My advice to friends and family currently is to use their "log out all active sessions" buttons, and to reset only crucial passwords. As Techcrunch mentioned in their coverage, many sites will likely pay for identity theft insurance to cover the change of a leak, rather than force password resets and lose their users trust. It's a risk analysis that each site needs to make individually, and it may be the right call to eat the losses of ~10 possibly compromised users instead of forcing millions to reset their passwords.
Wow, so just as bad as we thought.
We did not find any passwords, credit cards, health records, social security numbers, or customer encryption keys in the sample set.
BUT WAIT, THERE'S MORE
The sample included thousands of pages and was statistically significant to a confidence level of 99% with a margin of error of 2.5%.
Oh, so it could actually be as a high as 2.5% leaking encryption credentials. And if none of the data was found to leak anything sensitive where the fuck is the dataset? I've been around way too long to take a "study" like this at face value without third party verification.
I also enjoy the straight up lie at the end:
We are continuing to work with third party caches to expunge leaked data and will not let up until every bit has been removed.
That sounds great right? Well, its too bad that a lot of 'third parties' are a box sitting on the corporate network edge that hasn't been touched in 5 years. Deleting all of this data from third party caches is not physically possible. In fact it might actually make things worse because it's destroying evidence of which credentials were leaked.
IMO a leak this bad should be enough to sink cloudflare. A provider of SSL was randomly spitting out private data onto public websites. OVER A MILLION TIMES. Entire CA's have been shut down for leaking a couple hundred certificates. This has leaked private data over a million times, cloudflare is a joke
A != B
The only question is what is in these leaks exactly.
This leak is being downplayed by webmasters because it's so incredibly bad that there's no way of handling it. The credentials of practically any internet user could have been leaked. The only "safe" way to handle this is to give everyone in the US new credit cards and SSN's and to reset accounts and security questions for every user on a site with cloudflare
This issue is a drop in the bucket when it comes to the amount of sensitive data leaked.
So, yes, responsible websites can mitigate session cookies being leaked.
That said, I am not impressed by Cloudflare's transparency which in this case consists of downplaying things, blaming Google and Taviso and not really taking responsibility.
One of the caches they worked with was Baidu, which has direct ties to Chinese intelligence. Just because it isn't publicly available doesn't mean people aren't still pouring over it looking for useful data.
Anything short of invalidating any and all tokens, cookies and sessions is irresponsible and CloudFlare should communicate is as such.
They don't get into the full detail of how they knew which sites to look for. For example, if I created a site with these settings, used it for a week, and then deleted it...is my site counted? What if I changed the settings back to normal? Does it slide under their radar? Did they really look for every site that had those settings at any time, or just any site that had those settings at the specific time they looked?
To me, it still feels like there's a window where a bad actor could have created a site that triggers this behavior, and then pumped it with a scraper, over and over.
"Of the 1,242,071 requests that triggered the bug, we estimate more than half came from search engine crawlers."
This is very important to sort out, most people don't think about security much less the storing of credential or identifying data by search engines, this is a huge part of the incident response.
I think that what we should take away from this is that even though the bug existed it was responded to in a reasonable manner.
In total, between 22 September 2016 and 18 February 2017 we now estimate based on our logs the bug was triggered 1,242,071 times.
1) we have found no evidence based on our logs that the bug was maliciously exploited before it was patched;
2) the vast majority of Cloudflare customers had no data leaked;
3) after a review of tens of thousands of pages of leaked data from search engine caches, we have found a large number of instances of leaked internal Cloudflare headers and customer cookies, but we have not found any instances of passwords, credit card numbers, or health records;
and 4) our review is ongoing.
Based on looking at 1% of their log entries for 10 days of the total ~150 day period.
So uh, Cloudflare has logs of all page loads since september at the very least? And I guess with response sizes, since that seems like the most reasonable way for them to come up with such a number.
Awesome.
Guess they only kept a counter of requests. A bit of maths and you get an estimate for the amount of leaks.
> For the last twelve days we've been reviewing our logs to see if there's any evidence to indicate that a hacker was exploiting the bug before it was patched. We’ve found nothing so far to indicate that was the case.
right after they explain:
> Cloudbleed is different. It's more akin to learning that a stranger may have listened in on two employees at your company talking over lunch. ... you can't know exactly what the stranger may have heard, including potentially sensitive information about your company.
So they say they've found no evidence after explaining that the situation is analogous to a case where no evidence can be found even if information were stolen.
There seems to be some question about whether any passwords were actually leaked, but in principle some password data could have gotten out.
What would actually have been leaked, a cleartext password or the encrypted/hashed password stored on an affected server (or both)?
I have accounts on a number of sites that were affected by Cloudbleed, as indicated by http://www.doesitusecloudflare.com/, and I've already changed my passwords on those sites. But the passwords I use are randomly generated and should be practically unguessable given a hashed version. Would I have been at risk if I had left my passwords alone?
This is towards the edge of my knowledge circle but...
It doesn't appear that any hashing takes place client side[0] with myBB before the password is transmitted to the server, but I may have missed it when rechecking the code.
Presumably that would mean in my case users who logged in potentially had their cleartext passwords pass through Cloudflare's network when they submitted their login details.
Despite being behind https, when https is terminated at Cloudflare's edge and re-encrypted to my origin server (I use full/strict crypto setting), there was the possibility for decrypted data to be held in memory and then the possibility for it to be injected into the html pages of other sites when the bug was triggered.
It would be a very very small probability of actually happening, but not zero.
No details from the server should have leaked because only the content of transmitted data that passed over Cloudflare could have been leaked and the comparison of the provided password to the hashed/salted one in the database would have occurred on the server and not passed over the network.
So the answer is it would depend on the way the sites you use handled password transmission and comparison. Personally I'd change them, and actually have done so for sites I use.
[0] https://community.mybb.com/archive/index.php?thread-88668.ht...
Depends on the app. The apps I write will do a PBKDF2 w/ 20K rounds on a password before sending to the server, which then hashes that hash again before storing/comparing. I suspect most apps just sent the username/password plaintext over SSL, meaning the plaintext passwords would have been leaked if the app used Cloudflare to terminate SSL.
That said, if someone got the PBKDF2ed password hash from the SSL payload, they could just as easily use that to log into an account, but they'd have to at least do a little bit of work to break the hash if they wanted to actually crack the password.
In other words, you were absolutely right to change all your passwords.
Only GET responses should be cached (not requests), so it seems unlikely a password would leak.
I am happy with the way CloudFlare have responded to this.
Accounts commenting that have made 10 comments in 5 years, people defending CF that, after going through their profiles, are clearly linked to CloudFlare or maybe even employees. Top relies suddenly being near the bottom and replaced by posts supporting CloudFlare with statements that don't match what was actually in the blog post.
Just seems really really fishy to me
For this reason, please don't post unsubstantive allegations to HN. If you have evidence, or significant suspicious, email them to hn@ycombinator.com instead.
First, no, please don't do that. It's one of the worst patterns in active discussions.
Second, if anyone thinks they see gaming or abuse, they should let us know right away at hn@ycombinator.com so we can investigate. Please don't post about it in comments, for a couple reasons: (1) if there really is abuse going on, we need to know, and we don't see most comments; (2) most of us internet readers are orders of magnitude too quick to interpret our own cognitive biases as abuse by others (e.g. X seems obvious to me so anyone arguing ~X must be a shill). This places the threshold for useless, nasty arguments ('you're a shill. no, you're the shill') dangerously low. Combine that with the evil catnip power of all things meta and you get the most malignant strain of offtopicness there is, so we all have to be careful with it. And yes, genuine abuse does also exist. It's complicated that way.
We detached this subthread from https://news.ycombinator.com/item?id=13767122 and marked it off-topic.