SeaMicro builds box with 512 Atom CPUs
gigaom.com
gigaom.com
So the intention must be maximising I/O. What sort of workloads are so shared-nothing that they can parallelise to this many non-shared-memory CPUs efficiently? Content Delivery Networks? Seems incredibly niche; niche enough that the CDNs probably have already built their own.
And what exactly is the I/O bottleneck on a Xeon system that a bunch of Atom systems can do better? FSB/Memory throughput maybe? The Nehalems already have a 192-bit, 1333MHz DDR3 memory interface per CPU and gigantic caches, along with I/O that doesn't share data paths with memory accesses.
It's exactly the opposite. Not counting multiplies, pshufb, and a few other instructions, SSE operations are mostly reasonably performant.
The real problems are that the Atom has:
1) no out of order execution, which results in an enormous performance penalty on all code.
2) very high penalties for many "transfers" between operations: for example, the result of an arithmetic op cannot be used for addressing for quite a few clocks. This penalty is humongous when dealing with pointers-to-pointers in managed code, lookup tables, and so forth.
3) a completely un-pipelined multiplication unit that takes 7 cycles, compared to the pipelined 3-cycle Core 2 unit.
4) an enormous number of instructions that were microcoded internally (and thus extraordinarily slow) in order to reduce TDP and chip size.
5) only two integer ALUs, which can only dual-issue in some specific cases.
Overall, a 1.6Ghz Atom is significantly slower than a 1.6Ghz Pentium 4. A Core 2 is at least twice as fast per clock per core as a Pentium 4. And a Core i7 is another 40-50% faster per clock per core...
If you want something that gets out significantly more performance per watt than a Core 2 or Core i7, look at the ARM Cortex A9.
As for SSE, I was under the impression (I don't know where I read that) the Atom achieved its massive power reduction through a very aggressive simplification of the floating point hardware. Maybe the HT capabilities somewhat mitigate the lack of out-of-order execution. In my experience, floating point sucks a bit, but most other uses seem fine.
I am writing this on an Atom netbook while a Django app is importing a largish volume of data for a couple tests (it takes two hours to do it with a SQLite database) and Transmission downloads a torrent (need an alternate Ubuntu install disk). The browser is still responsive as long as I don't try to view videos. My worst complain with this machine is the hard disk being really slow, but considering price, size, weight and battery life, I guess I can't be too picky.
Yup these atoms are essentially I/O processors, just running enough buffer cache management, filesystem, and driver code required to keep other components (network and disk) at high utilization.
The benefit for atoms here is the short latency. They have short pipelines (better for the little actual computation required for I/O driving), and there are presumably more cores available here than the equivalent Xeon system (reducing queuing delays).
So it's basically a physically smaller supercomputer running low-power CPUs. It probably wins on real estate and power metrics, and probably loses on cost vs. racks of consumer stuff. Is there a market for that? Note that the investment came not from a VC fund, but from the DoE...
Their patent applications have some technical details: http://www.pat2pdf.org/pat2pdf/foo.pl?number=20080320181 http://www.pat2pdf.org/pat2pdf/foo.pl?number=20090216920 It looks kinda like a cheap x86 Blue Gene...
I think their argument is that a typical web-style datacenter is throwing up hundreds and hundreds of commodity servers to do some sort of distributed processing tasks (think of something like a hadoop cluster or a ton of app servers behind a load balancer).
With this setup their argument may be you're no longer in need of a giant network/switching infrastructure to do these types of compute tasks. Plus all the power/cost savings of only needing a couple boxes just to get those thousands of CPUs. I wonder what kind of RAM these things are going to have, as obviously that's pretty crucial.
But yeah, losing one of these things would cause more than a little headache of an outage. Seems like an interesting approach that's worth more thought though.