A Tour Inside CloudFlare's Latest Generation Servers
blog.cloudflare.com
blog.cloudflare.com
(edit to include links)
http://www.intel.com/content/www/us/en/network-adapters/conv...
===
Finally, while 10Gbps Ethernet can run across standard Cat5/6 cable, we elected to use SFP+ connectors. We chose this to have the flexibility between optical (fiber) and copper connections. Some network card and switch vendors lock down their equipment to only support proprietary SFP+s, which they charge a significant premium for. We spent significant time testing a combination of SFP+ vendors before finding FiberStore, a SFP+ manufacturer from which we could directly source SFP+s at a reasonable price that worked in the network gear we wanted to use.
As 10Gbps, and even 40G/100G, gain traction, you start running into cable length issues:
Attenuation and maximum length of AWG 7 copper wire:
Bit Rate (Gb/s), dB/m, Loss/m, Max length (m)
1, 0.721, 15.3%, 106
10, 2.28, 40.8%, 33.5
40, 4.56, 65.0%, 16
100, 7.21, 81.0%, 10.6
So DDoS could be...
1) an L3 packet you would drop, but still saturating your uplink (e.g. DNS amp)
2) a request for some static asset (img, html, octet...)
3) a request for a dynamic page (hitting your app farm)
CloudFlare provides something like an automatically populated CDN which includes defense against #1 and #2, and they're using distributed data centers, and servers with 10GE and SSDs to run their network. They said 23 data centers (locations), but didn't mention how many of these servers they run.Apparently they are able to run HTTP sessions over IP ANYCAST without any issue (I read claims you could only do UDP), so that's pretty cool. BTW - I wish it was easier to setup ANYCAST on your own... it seems like a major investment at the moment.
It's interesting they don't need more CPU power -- I would have expected more CPU would be deployed, since having minimal CPU could provide an attack vector.
Another thing I'm curious about is how they are distributing load between those cute servers they built? Do they segregate specific customers/domains onto specific IPs and then route those IPs to specific boxes in each data center? Or does everything come in "equal" and then get round-robin/least-load divided up by some massive load balancers? Basically, I wonder how much of the load balancing do they try to do "client-side" or "client-based" versus how much do they do strictly on the back-end, and what devices they are using for it?
As far as CPU, there isn't a whole lot of cycles dedicated to decoding and replying to network packets, and serving static content is incredibly trivial. The interrupts from high loads of network traffic (usually small packets) are arguably the biggest impact on the CPU, which is why you get network cards that offload L3 processing much more efficiently than your CPU can. Network appliance vendors rely on them to help do things like transparently filter traffic on 40GB/s interfaces in real time, which would probably be impossible with a normal CPU.
They mentioned in a previous post that they mainly rely on L3 routing to load balance. And I don't remember if they specified this, but they hinted in the post that all customer traffic can be served by any frontend box; it just pulls the customer data from their main storage and caches it on frontends as needed, as any good proxy does.
In general, though, Cloudflare's posts continue to be totally fascinating.
1: http://www.telegraph.co.uk/technology/news/8753784/The-300m-...
Yes, you can make a lot of money buying and selling stocks, but what is HFT actually contributing to the economy other than paying HR costs for its employees?
Market efficiency relies on extreme competition between investors that are highly intelligent and able to make quick decisions. When you can sum up much of those decisions in algorithms, what better way to deliver competition than HFT?
Opponents argue markets were efficient enough pre-HFT and that any benefits delivered by HFT are more 'academic' than practical. Opponents also argue that the level of uncertainty ("inexperience") with HFT brings more risks than argued benefits, as well as all of the "usual" criticisms of 'trading' vs. 'investing'.
As you may be able to tell, I myself am sitting on the fence on this issue, albeit with both feet pointing towards the opponents' side.
I'll try and illustrate with a contrived example. Say i want to buy 1000 shares at no more than $1 each. By using HFT, a trading firm can put together the order 'package' by combining many smaller trades at varying prices such that the average price comes out at the lowest (or target) price.
The competitive edge for a trading firm comes from being able to consistently fulfill orders, meaning they get more orders / customers.
Does that kind of explain the utility of HFT? Yes it allows a trading firm to make more money, but they way they do that is by providing a better service - not by simply exeucting a huge volume of trades and making fractions of a cent from each one.
Either way, it doesn't seem to be about actually funding companies. Maybe stock trades can help influence companies to change strategies or leaders, but I don't see the point of doing it in a forum from which the companies will never see the investment.
I don't think I understand the stock market as being anything more than a gambling game for people with a ton of money. Not sure what the IPOs of Facebook, Groupon, or Zynga did for anyone other than top execs who were already making a ton of cash per year, or the traders who bought and sold options.
In practice, it's probably fairly neutral, so long as no-one is playing silly b@#$%^rs, front-running other people's trades or similar.
I don't claim to understand it completely, but that's the justification I've heard before.
http://www.cpu-world.com/CPUs/Xeon/Intel-Xeon%20X5698%20-%20...
Wow, 4 generations of servers in 3 years. Talk about iterating quickly.
Was Cloudflare bootstrapped or did they start with a huge investment? 23 datacenters full of equipment sounds like a lot to me.
[1] - http://venturebeat.com/2009/11/25/cloudflare-floats-2-05m-eq... [2] - http://gigaom.com/2011/07/12/cloudflare-funding/
"We are indeed looking at Codel. We were actually working on backporting BQL+Codel to the 2.6.x kernel but the Google guys finally got the network stack under control enough for us to deploy >3.3. The 16MB of buffers hasn't hurt us much yet, and may in the long run save us from switches that have too shallow a buffer for the high contention ratios we run on the switch." -LinuXY
You tested it, but you did not mention why you ultimately decided against them. Was there something specifically less good about the AMD CPU's or is Intel giving you a discount for keeping your servers all intel? (i.e. NIC's and SSD's etc)
and
> We were willing to sacrifice a bit of clockspeed and spend a bit more on chips to save power. We tend to put our equipment in data centers that have high network density. These facilities, however, are usually older and don't always have the highest power capacity. We settled on our G4 servers having two Intel Xeon 2630L CPUs (a low power chip in the Sandybridge family) running at 2.0GHz. This gives us 12 physical cores (and 24 virtual cores with hyperthreading) per server. The power savings per chip (60 watts vs. 95 watts) is sufficient to allow us at least one more server per rack than we'd be able to get if we went with the non-low power version.
So a combination of additional instructions and power savings.
Wow 50%! Is this because raid controller performance hasn't kept up with the evolution from spinning disks to SSD's, or have raid controllers always had that much overhead?
They actually only wanted load balancing and I'm sure that their purpose built solution does a better job of being balanced while avoiding increased risk from striping or performance loss from mirroring or parity (I'm curious what level(s) they were using). Though, cutting out the RAID layer when they didn't need it does save them a trip through the controller, which is more important these days when compared to SSD "seek" times.
If CloudFlare ever offered optimized hosting (with PHP + MySQL), I would sign up in a snap and move all my websites there.
We finally left, after months of this, and have had no problems with downtime since. I really wanted CloudFlare to work - I was really excited about it when I signed us up. But at least for a bigger site with heavier traffic that relies on being up as much as possible, I can't say I'd recommend it until it straightens out its downtime issues (especially when paying for "always online").
Particularly for relatively low volume sites which have a short burst in traffic on occasion, CloudFlare can keep those sites running during the peaks.
I think the most important thing is transparency and correct expectations. If they set clear expectations, and they are transparent about how well they are meeting them, then it just comes down to delivery.
I found their status dashboard here: https://www.cloudflare.com/system-status. Unfortunately it doesn't show much long-term historical performance, it would be nice to see 30 days even 180 days of performance history to really evaluate them.
regal, did you find that when you had downtime on your site that it was reported in their status dashboard, and that was an accurate depiction of the service they provided? I think the worst-case scenario is getting hit with unreported downtime, because that brings up all sorts of questions.
Agreed. So long as a customer knows what he/she's signing up for, and gets that, everything's fine. I might have misread what the "99.99% uptime guarantee" was supposed to be for and gotten too excited about it / taken it too seriously when I first signed up, or maybe this is for something else that's too complicated for a part-time tech guy like me to understand.
When I'd log in when the site was down, half the time CloudFlare would have the green arrow next to the site with a "Site Online" type indicator; other times it'd have the brown dot-dot-dot "Site Offline" indicator. I'd confirm numbers on this but apparently the service doesn't save this or makes it no longer available to you on account termination. Pingdom Tools would report the site as down, and when visiting the site, it wouldn't load, or would take 10+ seconds to load. There would also frequently be a "This website is offline; no cached version is available" page from CloudFlare when trying to load the site, even on the homepage, despite the guarantee to supposedly be saving and serving cached copies of the site in the event of downtime (and despite that being what I thought we were paying for, mainly) - sometimes those cached copies would show up too; though more often, there'd just be this page:
http://image2.romantika.name/2013/01/cloudflare-website-offl...
There was one time where I had just upgraded a server at a host and load averages just spiked. Host techs were clueless as well as the server admin I hired to diagnose the issue. They surmised that it was a DDoS attack and to turn off CloudFlare because it might be a cause (what?). Well I'm not a server admin so what do I know and I asked CloudFlare about it. They said there were no problems on their end. After a month of stressful, intermittent downtimes, I decided to just switch to Softlayer, and lo and behold, the issues went away. Turned out that one of the SSDs in the RAID array was dying but the techs at the other host just never bothered to look at it.
If you ever do figure out the issues, try giving them another shot. Their features are excellent and I hate to say it but my applications are now so dependent on them to the point where important parts will break if I even try to move to a competitor (and there are none).
(I fully understand designing for the event - but the emphasis on it in the post makes it seem that you're under constant threat. I am assuming it is your customers that are actually being DDoS'd and Cloudflare just needs to be built up to stand against DDoS in this case??)
Seems like the system is robust, but I was looking for information on their policies regarding access-log retention and couldn't find much information online. Seems like they got a subpoena in the Barrett Brown case, and not sure how that all worked out.
It does seem like doing it on the target machine will reduce latency a bit, though, since the hardware TCP offloaders usually repeat the TCP handshake (this time to the actual server) after confirming that it's valid.
Am I overlooking something?
Uh, I've never had one fail gradually...
> "Intel reports that the 520-series drives have a mean time between failure (MTBF) of 1,200,000 hours (about 137 years)."
Yes, but they have a maximum write cycle, you can blow through the average consumer drive in a month and a half of concerted writing.
16MB vs 512KB is only larger? Not gigantic or even extremely large in comparison to the extremely tiny cache? Oh dear.
That's some fancy math and it probably is completely off, but another explanation is NIC performance doesn't follow exponential scaling for unexplained reasons.
I'd also be interested to know if polling mode was tested with any of the cards, and why it didn't work out
http://blog.cloudflare.com/how-the-cloudflare-team-got-into-...
We're continuing to experiment with polling-based network queues.
Cloudfront, Cloudflare, Rackspace via Akamai etc
Definitely several things we can learn from here. We have to do something similar, though we're still at a much smaller scale. That said, the wall'o'scaling is looming large and we're finding that even initial steps of building our own hardware is paying dividends.
I'm quite interested in the Disk I/O lesson's you've learned, specially when dealing with large amounts of RAM. We have to store large indexed data stores (NoSQL, usually Redis) for persisten, extremely high-speed access. A lot to learn here from what you did with SSD's to back that up, especially the lack of RAID.
And it seems some Gen4 will get Ivy Bridge 8 Core Xeon E5, 256GB Ram and Intel SC3700 SSD?
And why only 10Gbps Per Server, surely you could fit one more Solarflare in?
Other then that i really hope Cloudflare could expand beyond the current PoP.
https://support.solarflare.com/index.php?option=com_cognidox... (requires login)
<edit> We are the OEM (e.g. design and build) for large scale storage arrays for Amazon & Netflix, too, but not compute servers.
A CDN is a way to off-load the bulk of the requests to your webserver by moving the content as close as possible to your end-users, thus reducing the number of hops required to get to the content, which in return increases end-user satisfaction with your product due to a decrease in page load time.
The theory is that if a user gets a snappy service they are more willing to spend their money, and so e-commerce sites and sites that tend to monetize their users in some way find benefits in using services like these.
I hope that explains it adequately. To label cloudflare a mere CDN is a dis-service to them but for explanation purposes it might as well be, I'm sure someone from CloudFlare is able to give a much better explanation of just why their offering is not just an ordinary CDN but goes much further than that.