Intel’s High-End Cascade Lake CPUs to Support 3.84 TB of Memory per Socket
anandtech.com
anandtech.com
Since then, whenever I see a headline with Intel in it, I heavily discount it until I can verify the facts. They’ve damaged my trust, and I suspect many others.
> chips will support up to 3.84 TB of memory per socket ... due to combining 512 GB Optane DIMMs and 128GB DDR4 DIMMs
> ...
> in a 6 x Optane and 6 x DDR4 configuration, they will provide 3072 GB of 3D XPoint memory and 768 GB of DDR4 RAM for a total of 3.84 TB of memory
Optane DIMMs don't really have the performance characteristics of traditional DRAM. It sounds like the real capacity per socket is 12 x 128GB = 1.53 GB, which is the same capacity as the previous generation.
That being said, I'm optimistic about Optane DIMMs - it seems like an interesting performance point in between DRAM and SSDs.
What makes this not real AVX? Because there is not one dedicated 256-bit unit?
My understanding is that without any penalty was not true, but it is possible the new gen cpus have changed that, I have not firsthand benchmarked it.
Except literally running at half capacity. That doesn't qualify as a "penalty" to you?
The point was that AMD advertised support for AVX in a context where you'd expect it to be performance-comparable, and what they shipped was "support" for AVX in the sense that the code wouldn't crash, but wouldn't provide any performance benefit over SSE either.
Intel reduces clockspeeds when running AVX2 and drastically decreases clockspeeds when running AVX512. AMD's slightly smaller unit doesn't need to slow down, so the actual performance difference is smaller than it would appear.
Changing frequency takes time. Also, if you are putting through an AVX instruction and a few integer instructions at the same time, the integer instructions will downclock the whole time the AVX pipeline is in use (plus the time before/and after while the clockspeed is being adjusted) decreasing performance for more than just the AVX.
There are a couple articles written on the topic. Here's one.
https://blog.cloudflare.com/on-the-dangers-of-intels-frequen...
https://software.intel.com/en-us/forums/intel-isa-extensions...
I’ve been writing SIMD code for some time now, and I disagree.
You can read the documentation I’ve generated https://github.com/Const-me/IntelIntrinsics AVX intrinsics begin with _mm256_.
You’ll find that for some AVX operations, such as _mm256_fmadd_ps, _mm256_adds_epi8, _mm256_blendv_ps, both latency and throughput of Ryzen is equivalent to Skylake. For some others, e.g. _mm256_add_ps, _mm256_mul_ps, Ryzen is close to some older Intel, Haswell or Broadwell. Only for very rarely used stuff, e.g. _mm256_broadcast_ps, _mm256_madd_epi16, Ryzen is significantly slower than Intel.
AMD decided that consistent clockspeeds were more important than wider units. They decided that most users won't use enough AVX instructions to overwhelm the AVX unit (especially with consistent clocks), so there was no justification in ballooning the die size and increasing power consumption without a decent payoff.
https://en.wikipedia.org/wiki/Advanced_Vector_Extensions#New...
Optane is almost entirely hype, with no substance.
As it has been from the very first announcement. Intel is being very misleading marketing Xpoint as "memory". It's not even saying it's "close to memory" now. They're literally calling Xpoint memory (as in the same meaning we use for RAM).
Considering Xpoint has much lower performance than RAM, I'm starting to believe that maybe the FTC should intervene over "false advertising". It's just too bad the U.S. doesn't have as strong laws as Europe does for false advertising.
Data center admins are going to be able to get the capitol for real DDR4 over Optane. They'd need to be ordering a lot of servers before the price point becomes significant enough that they'd order an Optane configuration and decide to benchmark them.
I'd be curious about the real benchmarking at that point. It there a performance boost to simply have more memory, even if some of it is Optane, or is it better to have a lower capacity without the Optane sticks? I have a feeling this will vary heavily by application too.
I would imagine it depends entirely on whether your working set fits in main memory already. if yes, you're not gonna get much speedup from more memory. but if you have to keep swapping from a traditional SSD, it's gonna be hard for faster main memory to make up for that.
The advantage here being latency (and thus single queue throughput)
No, it's not, that's the whole point of pcie, and dimms share the memory bus anyway, so the whole point is moot.
M.2 slots are a whole another story, they are usually connected to the chipset instead of CPU on desktop motherboards (and finding this information for any particular board is difficult). The chipset is connected with CPU through DMI, which equals to pcie 4x and this is shared with everything on chipset - satas, gigabit ethernet, usb 3...
PCIe x4 is 3.94GB/s. PCIe x16 is 15.76GB/s.
> DDR4 dimm read throughput is about 20GB/s per dimm
Per channel, not per dimm. Two dimms typically go into a singular channel.
> The advantage here being latency (and thus single queue throughput)
Agreed
The storage dimms have some notable benefits over pcie nvme ssds - latency to the processor is extremely low, they give extra total storage I/O since they're leveraging a different controller on the cpu with its own bandwidth, while still leaving pcie bandwidth for traditional storage. They can add capacity to a small system (a 1U server can hold 24 such ssds)
As I see it though, these optane dimms are just fast swap drives which do not offer the same consistent throughput as true memory dimms. They are transparent to the OS which is normally in charge of deciding what stays in memory and what goes to swap. Use of an Intel sdk is required to leverage much of the benefits. It's also Intel-only as far as support goes for now. As usual the marketing doesn't mention the throughput or pricing and instead of comparing to memory they compare it to storage.
http://www.admin-magazine.com/HPC/Articles/NVDIMM-Persistent...
https://thessdguy.com/an-nvdimm-primer-part-1-of-2/
https://www.extremetech.com/extreme/270270-intel-announces-n...
And the reason for that is because their chips/products don't improve as much every year, and Intel is becoming more desperate each year as it faces more competition from Arm, and now AMD, too. So it's trying to "make-up the difference" by being more misleading about it. It's not just the competition either, but the fact that Intel needs to make it sound like their next-gen chips "are so much better" than last year, and make you buy the new ones, too.
I'll give an analogy. Let's say a year ago Intel's CPU got an IPC boost of 10%. But this year it's only getting a 5% boost. Intel will either increase power consumption of the chip to get that 10% or will increase boost speed while keeping the base clock speed the same or lower it for the next-gen chip, just so it can say "the peak performance is once again +10% for this new generation!".
Kaby Lake, Coffee Lake, and whatever Lake is coming next that isn't on 10nm all suffer from this kind of thinking. It's the same kind of thinking that created the whole scandal with the 28-core "5GHz chip". Intel's marketing is all about creating the perception that their chips are significantly better than the last generation. It's why they now have "Xeon Gold", too.
At this point I wouldn't trust any of Intel's announcements until they are verified by trustworthy parties (of which there are only a handful, because most don't go deep enough into their reviews to actually spot what Intel has done to their products to mislead customers).
What's wrong with that? High-end workstations often have the same sockets as servers don't they? Did they say it wasn't a server socket?
"What Platform Did It Run On? Intel did not answer this platform directly, however it was clear that the CPU was aimed at the LGA3647 server-based socket given from our examination of the demo system. It was unclear how Intel was going to promote this as an extreme workstation-type system, however Intel did note that they expect only a select market to be interested in this type of processor: a niche of a niche."
supposedly coming out in Q3, i.e. soon
Edit: emphasis on threadripper, which is a consumer product.
Servers is also where you get back-and-forth discussions with customers large enough to be worth listening to. Getting useful feedback from the users of desktop CPUs is difficult and time-consuming. But someone like Facebook or Google, who buy thousands of CPUs, and know how to use them, can give you quick and reliable feedback. They can test prototypes and suggest changes within weeks.
It doesn't, and AMD doesn't market Threadripper for server use. It only came up in this discussion because of a misunderstanding of Threadripper's capabilities.
https://www.tweaktown.com/news/58052/amd-epyc-32c-64t-3-2ghz...
They have gotten greedy and lazy without any competition.
Years before that they were interfering with compilers and cheating at benchmarks to screw over AMD which was actual, willful, evil deception, but THIS damaged your trust?
As others have pointed out, the addition comes from their support of Optane "dimms" which have much higher densities than DDR4 memory in the same form factor.
This is going to give the folks who write schedulers some more variants to play with. First it was per core memory, then it was per socket memory per core, and now it is memory type per socket per core. So you don't want your process running on Core3 with its pages mapped to Optane memory that is attached to the other socket (worst case).
(edit: obviously i've replied to the wrong comment, this is aimed at Intel's deceptive techniques which is currently the top comment)
And, with high core counts, we can start playing with things like dedicating cores (and their L1 caches) to single tasks, kind of specializing them as we do with mainframes.
that is why I tend to eye the ~6 core machines for desktop use, higher single core mhz on them
I still think we can push the envelope a bit further for most common desktop software.
There is no such thing as too many cores ;-)
The software developers could do a lot for memory throughput hungry applications (datastructure layout, not using bloated strings for everything, not using linked lists, etc.), alas memory based optimizations almost never happen and tend to require a great deal of hardware understanding + non scripting language(s).
L1 and L2 grow linearly with the number of cores. If you have cores to burn, pinning processes makes a lot of sense.
Pinning makes most sense in NUMA indeed. Nowadays my impression is that often times the multicore/socket hardware tend to be chopped down by virtual machines, though.
I'm going to wait 1-2 years to upgrade since I just got the 1st gen.
AMD becoming competitive again is one of the better things to happen in tech in recent years
If you are gaming however, 32 core is mostly useless and you should probably go with single thread performance. Intel i5 are still the CPUs of choice for gamers, though AMD Ryzen 5 are not bad either.
Disclosure: I work at Intel partly with the Optane SSDs (the PCIe version you mention).
But that hasn't stopped anyone because when talking about computers and data storage, unless you got some wacky old russian terenary system running in some bunker, everything should always be expressed using a base 2 system.
How will the numbers look after the next Spectre/TLBleed/etc patch?
I'd be more impressed by a boring processor that works.
In most other areas of engineering where there is still some concept of liability, the process is: validation first, save hot-rodding for v2. CPU industry is largely unencumbered by liability, driven by specious benchmarks, and has no use for that process.