Intel Xeon E5 v4 Review: Testing Broadwell-EP With Demanding Server Workloads
anandtech.com
anandtech.com
This is useful for me when comparing ranges of cores, power, price, and maybe it will be for you too.
I'll filter to my range of options, then make decisions on $/W, $/Ghz, etc..
https://docs.google.com/spreadsheets/d/1PcjgdtSV-2JLJXDpktjg...
http://danluu.com/clwb-pcommit/
Maybe we need even more cores soon...
So, now it has jumped from (disk -> network -> memory -> ...) to (network -> disk -> memory -> ...), which is a big change.
I've definitely noticed that when optimizing a query and deciding between a high number of seeks vs. a table scan, older versions of MSSQL will tend to be pessimistic of drive latencies and just go with the full scan (potentially incorrectly / prematurely). In an uncached scenario on an SSD, this is probably sub-optimal. My guess would be that instead of looking at actual seek latency, the optimizer was using reasonable guesses for spinning disks. I'm guessing newer versions are more SSD aware though.
It's more and more like network/disk -> L3 cache -> L2 cache. DRAM is pretty slow.
Because PCIe controller is anyways on the same chip as L3 cache, there's no reason to send the data on a long trip to DRAM and back. Until, of course, when the cache line gets evicted for reason or another.
Intel Xeon E5-2630 v4 10/20 2.2 GHz 25 MB 85W $667 US
Intel Xeon E5-2630L v4 10/20 1.8 GHz 25 MB 55W $612 US
Intel Xeon E5-2623 V4 4/8 2.6 GHz 5 MB 85W $444 US
Intel Xeon E5-2620 v4 8/16 2.1 GHz 20 MB 85W $417 US
Intel Xeon E5-2609 V4 8/8 1.7 GHz 20 MB 85W $306 US
Intel Xeon E5-2603 v4 6/6 1.7 GHz 10 MB 85W $213 US
A better calculation would be per GHz-core. Because a core at 2.2 is not the same thing as a core at 1.7.
(Of course, licensees of per-core-licensed software are screwed again.)
http://ark.intel.com/products/82932/Intel-Core-i7-5820K-Proc...
There is no way Intel would be charging over $4000.00 for a chip if they had any competition in that space.
I'm not trying to shill and hope I don't come across that way. I just want to share a semi-insider perspective to help others understand where MS is coming from. Do with that understanding what you will. I don't have a horse on this race.
However when it comes to CPU's the price was really good for a while now, hopefully the prices for SAS drives will soon be as good as CPU's aswell. I mean computing power is propably cheaper than storage at the moment.
If you don't need to run fast queries against your analytics or keep logs for something like 20 years you really won't get too much data. Especially when the only thing you index are the content of business documents.
Sort of like: buy one get one free (of slower clock speed).
I am impressed. (by the effective marketing deflecting the importance of clock speed).
2.2 GHz is guaranteed – without AVX, with you only get 1.8. 2.8 GHz is possible, assuming there are thermal reserves, power consumption is not hitting a limit, etc.
It is shameful that in 2016 we still don't have, say, parallel rendering in browsers. All hope is for Servo.
Any real world examples? Especially considering that at this point, sacrificing cores to boost the remaining ones seems to be a really bad deal with current silicon. Core power requirements appear to decrease faster than their actual computational speed does if you go low-power. Even if you lose 40% of performance due to overhead, if the same-TDP CPU package is twice as fast with more cores, you still win. (And who's to say that your implementation can't be improved in the future?)
As for "most useful things people actually want to do", it seems to me that a lot of relatively computationally expensive software still isn't using lots of cores where they're available in practice today.
One significant example is computer games. Since the advent of GPUs with effectively hundreds or thousands of parallel computations available, rendering hasn't been the bottleneck it once was. Today the bottleneck might instead be the game control logic that runs on a CPU, and is often still either single-core or divided among at most a small, fixed number of cores doing different tasks.
Another common real world example is graphics and image processing software. You'd think there might be a lot of natural data parallelism to exploit, but software in this area has made relatively little use of algorithms that scale to arbitrary numbers of cores so far.
A third example would be real-time processing, say operations on high speed network traffic. In this case you can sometimes dispatch different packets to different cores to process them in parallel, but the amount of processing you can do on any given packet might well be limited by the speed of a single core, because the overheads for cache misses or inter-core communications are prohibitive. If your processing needs to consider more than one packet at once, so you can't just spray packets at different cores as they arrive, then this can become a very significant real world bottleneck.
This isn't to say that none of these problems will ever be solved as we develop more understanding and better tools, but even in 2016 the state of the art is far from using as many cores as we have available efficiently for a lot of real world use cases. Manual parallelisation often has architecture-level implications and few development teams have the experience and foresight to get it right consistently with today's programming tools. Automatic optimisation to exploit data parallelism is an interesting research field but still in its infancy, and many mainstream programming languages have far from ideal semantics for such optimisations because of aliasing issues and the like. Either or probably both of these areas will have to advance considerably before we can assume that scaling out into more cores is generally going to give better performance than scaling up with faster CPUs and related hardware architecture.
Is there an alternative RAM database that you like better that is multi-threaded?
Single-treading is one of those. There are times when having more than one thread to help process things would come in very handy, though I recognize that the cost of adding this can be very high.
It's something that will have to be addressed eventually for a single Redis process to take advantage of newer hardware with very low ceilings on CPU power, but huge numbers of cores.
Also clock-speed doesn't mean a processor is slower, there are processors with a slower clock speed and still have a higher IPS.
(are E3 cores different from E5 cores?)