The Future Google Rackspace Power9 System
nextplatform.com
nextplatform.com
Here's the truth:
Google uses lots of compute power (insightful!)
Google isn't shifting to Power.
Google does have an active R&D program looking at Power.
TheNextPlatform misses the whole point here: That Zaius board has 32 DDR4 slots (commercially available servers from eg Dell max out at 24) and it has 2 NVLINK slots! (!!)
Those NVLINK slots are what Intel should be worried about, because that's where Google is prepared to pay money. They are building computers that lock themselves into NVidia and doing it gladly.
Intel better find a way to compete with NVidia on deep learning.
https://blogs.nvidia.com/blog/2014/11/14/what-is-nvlink/
http://www.tomshardware.com/news/nvidia-nvlink-boosts-perfor...
I'd be curious to hear what Intel is developing to compete with this.
NVLink is a communications protocol developed by Nvidia. NVLink specifies a point-to-point connection between a CPU and a GPU and also between a GPU and another GPU. NVLink products introduced to date focus on the high-performance application space.
and [1]:
NVLink – a power-efficient high-speed bus between the CPU and GPU, and between multiple GPUs. Allows much higher transfer speeds than those achievable by using PCI Express; estimated to provide between 80 and 200 GB/s.
It isn't a barrier for adoption because swapping racks out of a datacenter is easy, and they fit on standard datacenter floor tiles.
What is a barrier is that damned 48V.
Disclaimer: I run a hosting company.
Interesting what is the issue with 48V, I saw equipment for it seemed to be overpriced. Remember pricing out some stuff and as soon as 48V option for power came into play then price rose quite a bit.
Or is it that voltage is not high enough to be efficient for a large data center?
The computers have extremely simplistic power supplies that basically can't fail, and just DC->DC transform from 48v to 12, 5, and 3.3; and the large scale power supplies that convert three phase 240v (or whatever you're supplying it with) from the datacenter to 48V are much higher efficiency than the ones that would have been in the server (which you would have fed them, usually, something like single phase 208v).
Redundancy is supplied by just hooking multiple transformers in the middle rack to the + and - terminals on each PSU, instead of a convoluted multi-module redundant PSU (which always uses a single backplane, and backplanes in redundant PSUs fail surprisingly frequently).
The total round trip efficiency of this system is about as high as you can realistically get. 80Plus Titanium is 90-95% efficient (depending on load), but has efficiency losses in rack level distribution which 48V tries to correct.
However, 48V DC can be very dangerous to work with, and a lot of tech workers refuse to work with it. Now, if you believe it is dangerous or not (I've seen arguments stating that it is no more dangerous than single phase 208v) is immaterial, this is the opinion of a lot of workers.
The cost of 48V gear is expensive if you're not in a datacenter already setup to handle it. Facebook obviously doesn't have this problem because they build entire datacenters from scratch.
I personally don't believe in it because it doesn't buy me anything that single phase 208v doesn't give me, I do not pay enough in electricity to have the overhead of dealing with it.
While you could get fried from a single rack server's failing power supply, it is likely that it will die of some other cause before zapping you if you take simple precautions.
DC on the other hand, puts "the forces of nature" up close and personal to the back of the rack.
A wrong move will vaporize the thickest of metal screwdrivers and can easily do irreparable damage to a human.
Something as simple as not removing your wedding ring can result in an inadvertent bridge between + and - ; after which, bad things happen.
48VDC isn't sufficient to vaporize a metal screwdriver at the amperages used in most 48V datacenters, however, I know people who work with higher voltages than that, and they own expensive ceramic tools just for safety reasons. If I personally ran 48VDC, and my workers requested ceramic tools, I would not hesitate to expense those.
And yes, if you're working around DC >12v, you should be seriously using every method available to make sure you don't become a conductor, including only using one hand at a time to touch power relays (to avoid having your heart stopped).
Theoretically you should have zero leakage, and you should also trip on bigger currents (but a rack wouldn't use more than 10A maybe?)
Datacenters are not "built" with racks. They are not permanently affixed to the floor. Most datacenters do not keep empty racks on the datacenter floor, and keep the floor open.
If you are renting by the rack, no, most datacenters won't swap racks for you. You'd need to be renting entire cages for them to consider it.
Open Rack isn't really useful for small scale providers, it's more use to hyperscale companies like Facebook, Google, Amazon, and Microsoft. I don't think anyone we'd consider "medium scale" has adopted it (if I'm wrong, I'd love to see a story hit the front page about it).
Only one of those companies have adopted it, the others have considered it and, although they manufacture custom hardware and thus could take advantage of it easily, they have not done it.
To simply put, what are the incentive to switch over to Power9 platform?
But if Google switches to an entire different eco-system dragging it back into x86 won't be easy because all of their platforms are built for a completely different architecture.
So it's not certain Intel will win.
Purley Platform, Skylake Xeon offers:
Up to 8 Socket and 28 Core per Socket 6 Channel Memory Controller, 12 DIMM per Socket Support of Intel Xpoint NVDIMM 48 PCI-E Gen3, OmniPath 100G Connection
Offers up to 1.5TB Memory on a 2S Server, or 6TB Xpoint. If you push to the limit of 128GB TSV DRAM and 512GB Xpoint, that is potential of 3TB on 2S Server and 12TB of Xpoint.
Not to mention Intel's Network Controller. The Whole ecosystem from Intel Cloud is actually quite amazing. Both from Hardware innovation and Software Compiler they are working on. It is the same lock in as the PC Windows industry, and unless you get a dramatic new way of doing things. You cant simply switch the Mobile Industry to x86 or vice versa all by yourself. Even if you are as big as Google. Then you get 10nm Intel Server in 2018/2019.
Again, I dont see the incentive making the switch.
Again when you take into account of power, performance, and other parts of the Server components, the TCO of CPU is relatively small. Even if you gain 10% of TCO improvement, you have to factor in the future roadmap of the CPU, as well as the Software development, compiling and testing cost involved.
If you want to run an entire CPU on an FPGA, ASICs will clearly be better. FPGAs are mainly useful for things that need to be (re)programmable.
For things like programmable network routing/packet filtering, FPGAs could be a very effective solution.
Also what is the issue with the memory controller?
Power8 or Power9 has better memory controller then Intel Xeon. More memory channel, higher bandwidth, and higher memory capacity.
1. CAPI doesn't seem to get mentioned to much around here, but imagine an FPGA directly accessing some shared system memory. It's neat.
You can do some pretty cool things from a HW designer's perspective inside the accelerator, and in the main application. Since the accelerator is cache-coherent, and able to map the same virtual addresses as a given process (and attach to multiple processes' address spaces) the device can do "simple" things like follow pointers, which used to require building a command / data packet, DMA'ing it to the device, and then waiting for a response packet. This, effectively, frees up the main CPU to do other things, rather than wrangle data. It also means that bottlenecks move.
So, it "requires" CPU support, just not in the way usually meant.
Disclaimer: I work on this with some very smart people @ IBM. Opinions are my own.
When a PCIe device is in CAPI mode, the PCIe protocol is used as a transport layer, but the CAPI protocol rides on top, and hardware in the CPU's PHB (the CAPP unit) and hardware in the accelerator (the PSL in this case) cooperate to present the common address space to the process and to the accelerator itself. [1] If a CAPI-capable card's plugged in to a non-CAPI-capable slot, it remains a PCI card. If a non-CAPI card's plugged in to a CAPI-capable system, it remains a PCI card. If both sides match on protocol versions and the kernel contains the cxl driver, the kernel will switch the slot into CAPI mode, and the CAPP unit and PSL effectively take over the PCI link on either side.
[1] http://events.linuxfoundation.org/sites/events/files/slides/... - see page 13+ for some GPU / NVLink materials, and page 24+ for CAPI materials (oh - and I worked on the product who's data is quoted on page 29 [2]!
[2] https://www.ibm.com/developerworks/community/blogs/fe313521-...
CAPI is much more flexible and interesting. It allows a CAPI-capable connected device access to a process's virtual memory. Essentially, you can extend the CPU's capabilities with CAPI. Usually this would be an FPGA (and the utility of FPGAs for machine learning is very much a research topic), but I could easily see a DSP being useful for voice recognition. GPUs can take advantage of it too, but ML work is usually just offloaded entirely to the GPU.
CAPI is very very cool and designed by some very smart people. I'm excited what people will do with CAPI and FPGAs.
http://www-304.ibm.com/webapp/set2/sas/f/capi/CAPI_POWER8.pd... is a good intro.
[Disclosure: I work on CAPI at IBM]
The x86 server processors have too much legacy they can't get rid of, and that limits how far they can push it.
http://www.penguincomputing.com/products/rackmount-servers/o...
Or is it the "if you have to ask, it's too much for you" as seem to be the case with the IBM power systems?
[0]: https://www.phoronix.com/scan.php?page=article&item=talos-wo...
[Disclosure: IBMer]
http://www.anandtech.com/show/10230/ibm-nvidia-and-wistron-d...
For a platform to succees, it should provide low barrier of entry - may be a low capacity P9 system at a low cost ~ 1K ? Will be a better strategy for OF.
I know Power is targeting cloud computing applications, but IBM should consider low cost entry level gear for getting some market share at the lower level which can transition to higher margin markets .