Intel Lines Up ThunderX ARM Against Xeons
nextplatform.com
nextplatform.com
https://www.cavium.com/OCTEON-III_CN7XXX.html
Looked awesome. The ThunderX is the Octeon ported to ARM ISA far as I can tell. The Octeon II's with 16 cores plus Gbps I/O and crypto accelerators in PCI card form can be had on eBay for $200-300. Now, I have performance sheets from Intel confirming it performed as well as I hoped.
Great news for Cavium. ;)
After considering it all I went with the E3-1230v5. Near 6700k multicore performance, near 6700k single thread performance, max 80 watt TDP, and idles much lower. I also get a M.2 slot which allows pretty amazing performance (well above SATA) and a tiny footprint (gum stick or so).
I don't buy systems often though, the machine it's replacing is a quad core 3.0 GHz machines (Q6600) with 8GB ram from 2007. So I don't mind spending a bit more if it's going to last. Hopefully with 32GB ram this new machine will last quite a while.
The hybrid seemed like one of the best designs. Probably why AMD started copying it with their semi-custom business followed by Intel. But those are still full-on x86 processors with all kinds of circuits you probably don't need plus maybe some backdoors for AMT, etc. Cavium's design is much more interesting and maybe safer. I've mainly encouraged people to clone it with RISC-V cores rather than straight-up buy them. ;)
The newer ones are pretty impressive in performance and accelerators, though:
It's getting close, that's for sure - and there may well be some niche where it already makes sense, but I wouldn't know which one.
Probably the best value proposition is that the gap is likely to shrink, and investing now may help you prepare for a future switch when it makes more sense.
Proprietary blobs such as Intel's ME that cannot be removed from Intel Xeon systems pose a security risk. If the ARM systems are found to lack them and they achieve the status of being good enough, security minded individuals and organizations will buy them because of that.
As long as that happens, Cavium will be able to use the revenues to develop the next generation of chips to better compete on traditional metrics like Intel was able to do against the RISC vendors of the past. As a security minded individual myself, I am well aware of that and I look forward to supporting such developments by purchasing such hardware for applications where it can be used.
There are others thinking this too. Some of them are watching the Talos workstation motherboard come to market with that in mind. The people behind it are even talking about IBM's executives being willing to make the POWER9 processor's microcode open source if the POWER8 model succeeds.
Do you expect the above to be anything more than a rounding error of the total server market? I doubt Cavium is going to finance its development from that.
Is ARM's TrustZone not the same as Intel's ME? At least I remember that AMD uses a ARM Core with TrustZone in it's x86 CPU's to basically achieve the same (or something similar) as Intel's Management Engine.
[1] quoted from: http://www.forbes.com/sites/moorinsights/2016/05/30/cavium-s...
AFAICT the benchmarks seem to focus on traditional use cases like CPU-bound/memory-bound computation. But the ever-growing use case is the physical hardware that underlies the commercial "cloud computing" focus. IMO much of that market is running I/O-bound code and if I were sourcing hardware for one of those providers I would want to offer a tier which could serve enormous numbers of clients on a single node.
Power consumption is another huge element of data center costs and scaling up ARM-based servers is likely to reap a big savings.
Also; most of these CPU's aren't designed in a vacuum: faster CPUs tend to have better I/O subsystems. E.g. memory latency (which is I/O from a CPU's perspective) is much, much worse on ARM, probably because it hasn't historically mattered as much.
So I'd expect arm to make sense on special purpose I/O related stuff, where you can be sure the limitations aren't an issue. Do some fast e.g. SSDs include ARM chips?
Perhaps
For some definitions of "on the market". The ARM server chips can be hard to get still.
https://www.apm.com/news/appliedmicro-announces-x-gene-3-the...
http://www.nextplatform.com/2016/04/28/arm-server-chips-xeon...
https://semiaccurate.com/2016/04/25/appliedmicros-x-gene-3-a...
"The point Intel is making in the chart above is that it has mastered NUMA scaling, which is no surprise since the company has been at it for decades."
SGI was one the pioneers of ccNUMA (Origin 2000?) in the late 90s. That would constitute decades.
Around the same time, Intel delivered the ASCI Red supercomputer using P6 processors with a snooping coherency scheme.
I don't think Intel had significant NUMA experience (at least in any commercial products) prior to AMD K8 (2003), and their first NUMA product would have been IA64 (E8870 chipset?) or the Nehalem microarchitecture chips. That is not quite a decade, let alone decades.
Anyway, the article's author likely made the mistake of thinking that multisocket implies NUMA. It is an easy mistake to make for those who do not know the details of how systems are implemented.
This was the system in which the directory logic was formally verified: http://www.sgidepot.co.uk/origin/compcon97_dv.pdf
I agree the author probably made a simple mistake, but it is misleading to people at the "enthusiast" level who then parrot things like this as flamewar fodder. A lot of respect is owed to computer engineers at corporations that are defunct or a fraction of their former greatness (e.g., SGI, DEC, HP).
As a further example, Bob Colwell's book, "The Pentium Chronicles", is a great read and teaches some powerful lessons about how to successfully do engineering in the kind of large organizations that can scale production of a design. But if you take it at face value, the book makes no meaningful attempt to inform the reader that other companies had done Out-of-Order well before P6. It leaves a naive reader with the idea Intel pioneered OOO as well, which is no way true.
I guess I'm just (maybe overly) sensitive to sloppiness and omission leaving an inaccurate impression of the history, especially when it's so easy these days to fact-check.
Certainly for 2 and 4 sockets Intel uncore (i.e. the connecting fabric) is state of the art and has been for a while (but not 2 decades).
Also agree that the crux of the point is multisocket scaling. Though with enough sockets (if memory serves, in mid-2000s, the rough consensus was 4), I don't think computer architects think there is any way to scale well without a NUMA approach of some kind, so it's not just an academic point. With today's memory bandwidth and "memory transaction volume", perhaps even 2 sockets wouldn't scale well in a uniform memory access configuration.
Maybe CPU margins are going to drop too? ARM is likely the successor. ARM already won mobile market, maybe it start eating some of the cloud market? Do we really care on which CPU architecture runs S3 servers and etc.? That gives cloud providers huge leverage in negotiations. E.g. AWS can demand more discounts for Intel CPUs and if they won't comply slowing, but surely move some portion of it's infrastructure to ARM.
Not at all, specially when the language runtimes are rich enough to abstract OS and hardware architectures.
Even GPUs can be treated as an "acceleration device" by rich GPGPU runtimes.
Intel ate the server market in the late '90/eary '00 thanks to the fat margins it had on the desktop.
That's not currently the case in the ARM world were only Apple and maybe Qualcomm have the needed margins, and Apple currently doesn't seem interested in the server market.
Somebody above the "we" in that statement has to care.