Facebook is the first to jump into ARM servers
semiaccurate.com
semiaccurate.com
Next, certain types of workloads don't make sense. Anything CPU-bound or poorly scalable (eg. traditional database workloads). Again FB should have plenty of work that scales out relatively effortlessly (note the relatively) and have moderate memory and bandwidth requirements per process. Though in the aggregate you'd expect very large usage of both!
For suitable workloads, working back of the envelope, ARM will hopefully lead to some highly competitive, if not new record, scores on metrics like requests/joule or requests/$. Enery consumption and server cost being important at that scale..
Now for the in-memory caching or database workloads, which want either more memory or faster CPUs, Flash can be used to address capacity at the cost of a couple of orders of magnitude extra latency - albeit a couple less than hitting disk. Back of the envelope, anything that looks too much like a traditional database workload I'd leave on grunty x86 machines. Ditto for any CPU-bound.
So, lets speculate on how to build a machine based on ARM for the types of workloads we care about. Lets assume we don't design a new core but work with a vendor on a System-on-Chip using ARM hard macros. This is very back of the envelope and we'd need to break out spreadsheets to get this nailed down right.. but lets have some fun..
Our hypothetical SoC would be
. Cache-coherent quad-core Cortex A-9 @ 2GHz . PCIe interface on chip . 1Gb ethernet interface on chip . SATA ports on chip . Memory controller
connected to 4GB ECC memory per server. I'll get back to storage.
Now, this should be small and quite low power server. Within a 1U sled we should be able to pack say 6 or perhaps 8 of these, along with dual power supplies (for the entire sled of machines). If possible I'd have distributed power redundancy by including a battery in the sled rather than hooking up to external UPS.
I'd use an internal 1Gb switch which itself is connected to the top of rack switch. We get local cheap communication between the servers, plus we make cabling significantly easier and keep the cost of the top-of-rack switch down. A more whacky alternative would be to use short-range radio. Fewer wires, potentially more bandwidth, but something I'd like to hammer on in the lab before going anywhere near the datacenter.
Now, for the pesky 4GB per server memory limit and storage. We have a few interesting options. We can add flash per server, or, given the $/GB perhaps one machine in the sled gets it and acts as a local memcached server with the ability to fall back to accessing remote ones. We could also have a local file server with, for example, 1TB of storage via 2 flash-augmented disks (eg. seagate momentus xt disks). With good staging of data, we could even make this sled a good building block for throughput-oriented data-intensive work (eg. mapreduce type work). We have lower IO bandwidth but we've also kept the CPU performance down to levels where we have a fighting chance of feeding them.
Obviously, you need your software to run there. Chalk one up for relatively easily ported open source code without being dependent on a slow-moving vendor.
The above is rampant speculation and there are many interesting design points - it's great to see someone trying new things and taking advantage of changing hardware ratios to profit.
That could be about right. I know the number was 20,000 at the end of 2008 from sources within Facebook.
One application for this is in mission-critical transaction processing where computing power is not the active limitation. Instead, the idea is that each core executes the same set of instructions with the same set of input data, and in the end if one disagrees with the other, the entire CPU rolls back the current transaction, takes itself out of the mesh, and alerts an operator... who, if it's an IBM mainframe, walks over with a new processor, yanks the old one, and replaces it. Hot. Downtime: none. Transactions lost: none.
Sure, it would need more I/O capability + ECC, but still -- it's potentially a low-cost, low-power, highly reliable competitor to some of IBM's POWER and PowerPC processors.
Without distribution a key/value store is a hash table. To have any value added it's going to need to be a distributed key/value store. Distributed key/value stores, however, run just as well on commodity hardware. Problem is very few companies work at a scale at which per machine efficiency, power or other cost savings are going to matter; that's a very small market and one that's difficult to sell to.
On the other hand, a customized Linux distribution with a nice UI for deploying key/value stores would be a good idea.
I think even on a small scale, or perhaps that should be especially on a small scale, people could profitably look at the mysql/nosql/memcached (the latter optionally with persistence via flash) from some of the appliance vendors.
I have no financial interest or direct experience, but the data sheet on this:
http://www.schoonerinfotech.com/datasheets/Schooner_DS_Memca...
looks worthy of a closer read.
Personally, I don't use the appliances because for our extremely latency-sensitive apps we do it all from scratch for various reasons. For large scale storage we build off commodity machines. One of our prototypes runs very happily off machines from scalableinformatics.com - highly recommended.
Horses for courses and all that..
This story is completely false. Facebook continuously evaluates and helps develop new technologies we believe will improve the performance, efficiency or reliability of our infrastructure. However, we have no plans to deploy ARM servers in our Prineville, Oregon data center.
Cheap Linux/x86 servers own most of the web space. But the rest of the world runs much more heterogeneously. (And that list above ignores real legacy stuff.. You can buy all of that currently (with the exception of the older windows stuff))
Although originally developed first for 32-bit x86-based PCs (386 or higher), today Linux also runs on (at least) the Alpha AXP, Sun SPARC, Motorola 68000, PowerPC, ARM, Hitachi SuperH, IBM S/390, MIPS, HP PA-RISC, Intel IA-64, AMD x86-64, AXIS CRIS, Renesas M32R, Atmel AVR32, Renesas H8/300, NEC V850, Tensilica Xtensa, and Analog Devices Blackfin architectures; for many of these architectures in both 32- and 64-bit variants.
We could also say: pretty much everything under the sun.
oh, and supercomputers: http://www.top500.org/stats/list/35/osfam
I see an analogy between x86 and ARM and Porter's generic strategies. ARM is like the cost leadership, with x86 being the differentiation strategy - where the cost is the cost of energy. x86 is able to do a lot more things than ARM, but with higher energy costs. And then there is perhaps CUDA etc. for the segmentation strategy.
They roll a heterogenous server setup intentionally to minimize impact from hardware faults, bugs, etc. You build your software at a high level where the hardware under is abstracted away.
Namely, it can run Z books along with POWER7 blades and x86 blades in the same (admittedly very large) box. Basically, run your compute workloads on Power, database on x86, and control on Z. Seems like a pretty good idea to me!
[1] http://www-03.ibm.com/systems/z/news/announcement/20100722_a...
So for a midsized datacenter with say 200 racks, each ARM based rack could have 320 cores in it, giving 64K cores.
Since a FB page may take 200ms to load (or some other load time internal to FB response time targets), such a setup could handle 64K * 5 (i.e. 5 users per second at 200ms/page) or 320K simultaneous users per second; figuring 30 seconds of viewing time per user, this gives you almost 10 million users' worth of capacity.
(all numbers back of the envelope, hypothetical)
In general, I don't understand your comparison of RISC vs x86. The original premise of RISC, namely having very simple chips that can be clocked real fast because of their simplicity has long since been abandoned. Modern RISC chips have all the crazy complexity of x86 chips: you'll find big pipelines and dynamic register renaming and multiple execution units galore. And modern x86 chips internally look a lot like modern RISC chips as well: after they convert native instructions into micro-ops, there doesn't seem to be much difference.
A part of my point is that it hasn't been holding x86 back for a long time. A part of the article's point is that perhaps it will soon. Reread my comment with a skeptical tone.
To my mind Intel's quixotic purchase of McAfee only shows that they've cornered themselves into a market that's peaked.
You're right about the "haphazardly designed ISA" but they've still made it run very fast by throwing a lot of engineering effort at it.
As for McAfee, I don't know what they bleep they're trying to do there, but I suspect they can afford it (especially since AMD didn't keep their eye on the ball for so long). Their previous failed communications and media ventures didn't seem to materially hinder their bread and butter CPU/chipset business.
If you write clean ANSI C code, you're basically compatible by default. If they failed to run on i386, that means they did something special that assumed a different architecture. In that case you're likely to hit the same problem whether they use arm, alpha, infineons, intels, or whatever else.
Source: http://www.eetimes.com/electronics-news/4206387/ARM7-40bit-v...