How to Bypass Cloudflare: A Comprehensive Guide
zenrows.com
zenrows.com
I discovered our company's help documentation (and integration guides), hosted by readme.com, were completely de-indexed from Google for the past 3 months.
Our Readme docs were formerly our #1 source of organic (free) leads.
After investigating, Cloudflare (as configured by Readme) was blocking Googlebot when using Cloudflare Workers. Cloudflare was returning a 403 for Googlebot, but returning pages as usual for regular users.
The cause: we were using Workers to rewrite some URLs at the edge (replacing Readme's default images with optimized + compressed images, using Cloudflare's own image optimization service).
By using Workers to do this, it resulted in Readme's Cloudflare account receiving requests from our domain with "googlebot" useragent, but from an IP that wasn't verified as a googlebot IP address (I assume the Worker was requesting the Readme site using the Googlebot user agent but with whatever IP address is used when using CF Workers).
I emailed Cloudflare support but it was clear it would take a lot of time to get them to understand the issue (and probably longer to fix it).
So, we had to spend a lot of time figuring out how to allow Googlebot requests past Cloudflare's "fake bot" firewall rule.
In our own Cloudflare account, we have all security settings at the lowest sensitivity possible (or turned off completely). We serve over 500 billion requests a month (10+ TB of bandwidth), and the amount of blocked traffic to seemingly legitimate clients was surprisingly high.
I love Cloudflare (and own quite a bit of their stock) but I'm beginning to rethink my stance on their service. They make it extremely easy to enable powerful features with little visibility or control over the details of how those features work.
Another SEO nightmare is their "Crawler Hints" service. I highly recommend no one uses this if you are ever the target of automated security scanners (e.g. ones used by bug bounty white hat hackers). With "crawler hints" enabled and with a white hat hacker running a scan of your site hitting random URLs... results in bingbot, yandex, and other search engines attempting to index every single one of the URLs hit by the security scanners used by hackers.
Basically, it's a mess, and the only way to really fix it is to bypass cloudflare or spend a lot of time and money with Cloudflare debugging.
Next quarter I'm faced with the decision of either doubling down of Cloudflare and getting an Enterprise plan with them ($20k+) or just ripping them out of our stack and going back to our old AWS Cloudfront set up which has fewer POPs, but was much less of a hassle.
Is Fastly a viable alternative for you?
The vast majority of such requests are dodgy scanning operations likely looking for email addresses or exploitable forms.
1. For many years I have had great results with not sending a UA header. It is also, IMO, an effective means to discover the true number of websites that refuse to fulfill a request in the absence of a UA header, which IME is extremely small. For that small handful of sites, one can send a "fake" UA header of one's choosing. sec.gov is an example of such a site.
2. http://developers.google.com/static/search/apis/ipranges/goo...
Or maybe google crawler also runs on GCP and it's indistinguishable from regular $5 compute users
https://developers.google.com/static/search/apis/ipranges/go...
To add to some of the other experiences here about no-UA: I've tried that before too, and it was notably worse than pretending to be Google; lots of sites just return "Internal Server Error" or similar messages.
For example the Cloudflare Blog's RSS feed is very often blocked from public-cloud IP ranges with specific clients. This is an endpoint that is intended to be public, is cachable and even intended to be accessed by bots! This is a common issue that should be very easy to solve technically but highlights how Cloudflare is not a set-and-forget solution. If they can't configure their own blog (a super simple case) correctly it is clear that using the tool correct requires special care and monitoring of the limited visibility that they provide you.
Two things that have happened to me:
* Cloudflare has decided that I'm a bot and stalled me, given me capchas, or just blocked me outright.
* Cloudflare has shown me marketing claiming that 40% of traffic is bots.
I'm not particularly impressed.
Was this definitely the cause? It's somewhat surprising to hear that requests would be rejected if the user agent doesn't match a set of hard coded IP addresses.
Were you able to resolve this in the end? If not and the cause is what you suspect then perhaps changing the user agent in your worker might be a workaround.
It’s fairly common for DDoS/scraping prevention, Googlebot (and most other crawlers) publish their IP ranges for that reason[0][1][2]. I don’t work at Cloudflare though, so no insider knowledge of what you folks are doing.
[0] https://developers.google.com/search/docs/crawling-indexing/...
[1] https://developers.facebook.com/docs/sharing/webmasters/craw...
[2] https://developer.twitter.com/en/docs/twitter-for-websites/c...
No, I haven't confirmed it. We jumped straight to fixing it without debugging the root cause. It's possible the cause is something totally different (I should have added this caveat in my original post). I was just speculating.
Regardless of using Workers or not, Cloudflare requests URLs with the header `cf-connecting-ip` and `x-forwarded-for` set to the actual client's IP address, and the website behind CF should be using this header to get source IPs.
Most useful services for that are https://shodan.io/ and https://search.censys.io/. I've had decent successes with Censys on finding real IP addresses of websites behind Cloudflare. Of course you might also have success by checking history of DNS records for a particular domain.
Seems to work great when you use something like a GUID, and no need for IP whitelisting.
How is using CF’s origin CA preventing the connection to the real backend in order to bypass Cloudflare? you cam just ignore the SSL error couldn’t you?
To others out there who explore this: As with all scraping, be gentle! If you start pounding on someone's origin server directly, you're much more likely to be noticed than if you're pounding on something behind a CloudFlare cache. Set rate limits, scrape during off-peak hours, etc. Be a good scraping citizen.
Only afterward, they will protect it with Cloudflare/Akamai/Cloudfront, etc, which means you'll have a DNS history trail to lookup.
I've had good luck with services such as SecurityTrails, Virustotal and CompleteDNS for DNS history.
Once you find the original A record IP address, you can put it in your hosts file, pointing to the domain. Then, you will be able to access the site, bypassing the CDN without issue.
Edit: I think the modern term is content native advertising, although I'm perfectly happy to keep using the word infomercial.
There's a term for that, infomercial :-)
The first 31 pages pretty extensively provide information on bypassing Cloudflare, which would be very useful to and save a lot of time for someone who is tasked with implementing Cloudflare bypass software.
This is followed by 1 page that talks about their product for doing this.
In other words it delivers pretty much what the title said it would, and then follows that with a small mention of their product.
That's not clickbait by any even remotely reasonable definition.
- Giving more background than is appropriate to the subject (explaining what cloudflare is in an article about bypassing it)
- Lots of fluff about "what we're going to cover" like it's a poorly-written highschool essay
- Asking and answering questions rather than stating things: "Can Cloudflare be bypassed? Thankfully, the answer is yes!"
I'm not entirely sure what drives these things, but they seem to be very common in this sort of content marketing article. I'm guessing a lot of it is SEO-driven.
This particular article has more actual content than most, but still ultimately devolves into an ad, of course.
> I'm not entirely sure what drives these things, but they seem to be very common in this sort of content marketing article. I'm guessing a lot of it is SEO-driven.
I suspect that this is them trying to get into Google's "frequent question"/"people also ask" [0] box, because that seems like a common search term ("can you bypass cloudflare")
I find the mention of a series A fundraising round at the top interesting too. Do the funders really expect something other than an escalating technical arms race that eventually outpaces them?
I certainly wouldn't mind if all advertising was done this way. Unless the information here has been copy-pasted from somewhere else, there's enough value there to stand on its own.
Of course, those trying to profit from online advertising services seek to collect the same (fingerprinting) data. Do Cloudflare terms of service/privacy policy allow Cloudflare to do anything they want with this data, or are there limits.