Since then, whenever I see a headline with Intel in it, I heavily discount it until I can verify the facts. They’ve damaged my trust, and I suspect many others.
Since then, whenever I see a headline with Intel in it, I heavily discount it until I can verify the facts. They’ve damaged my trust, and I suspect many others.
> chips will support up to 3.84 TB of memory per socket ... due to combining 512 GB Optane DIMMs and 128GB DDR4 DIMMs
> ...
> in a 6 x Optane and 6 x DDR4 configuration, they will provide 3072 GB of 3D XPoint memory and 768 GB of DDR4 RAM for a total of 3.84 TB of memory
Optane DIMMs don't really have the performance characteristics of traditional DRAM. It sounds like the real capacity per socket is 12 x 128GB = 1.53 GB, which is the same capacity as the previous generation.
That being said, I'm optimistic about Optane DIMMs - it seems like an interesting performance point in between DRAM and SSDs.
AMD decided that consistent clockspeeds were more important than wider units. They decided that most users won't use enough AVX instructions to overwhelm the AVX unit (especially with consistent clocks), so there was no justification in ballooning the die size and increasing power consumption without a decent payoff.
https://en.wikipedia.org/wiki/Advanced_Vector_Extensions#New...
What makes this not real AVX? Because there is not one dedicated 256-bit unit?
Except literally running at half capacity. That doesn't qualify as a "penalty" to you?
The point was that AMD advertised support for AVX in a context where you'd expect it to be performance-comparable, and what they shipped was "support" for AVX in the sense that the code wouldn't crash, but wouldn't provide any performance benefit over SSE either.
Intel reduces clockspeeds when running AVX2 and drastically decreases clockspeeds when running AVX512. AMD's slightly smaller unit doesn't need to slow down, so the actual performance difference is smaller than it would appear.
Changing frequency takes time. Also, if you are putting through an AVX instruction and a few integer instructions at the same time, the integer instructions will downclock the whole time the AVX pipeline is in use (plus the time before/and after while the clockspeed is being adjusted) decreasing performance for more than just the AVX.
There are a couple articles written on the topic. Here's one.
https://blog.cloudflare.com/on-the-dangers-of-intels-frequen...
https://software.intel.com/en-us/forums/intel-isa-extensions...
I’ve been writing SIMD code for some time now, and I disagree.
You can read the documentation I’ve generated https://github.com/Const-me/IntelIntrinsics AVX intrinsics begin with _mm256_.
You’ll find that for some AVX operations, such as _mm256_fmadd_ps, _mm256_adds_epi8, _mm256_blendv_ps, both latency and throughput of Ryzen is equivalent to Skylake. For some others, e.g. _mm256_add_ps, _mm256_mul_ps, Ryzen is close to some older Intel, Haswell or Broadwell. Only for very rarely used stuff, e.g. _mm256_broadcast_ps, _mm256_madd_epi16, Ryzen is significantly slower than Intel.
My understanding is that without any penalty was not true, but it is possible the new gen cpus have changed that, I have not firsthand benchmarked it.
Data center admins are going to be able to get the capitol for real DDR4 over Optane. They'd need to be ordering a lot of servers before the price point becomes significant enough that they'd order an Optane configuration and decide to benchmark them.
I'd be curious about the real benchmarking at that point. It there a performance boost to simply have more memory, even if some of it is Optane, or is it better to have a lower capacity without the Optane sticks? I have a feeling this will vary heavily by application too.
The storage dimms have some notable benefits over pcie nvme ssds - latency to the processor is extremely low, they give extra total storage I/O since they're leveraging a different controller on the cpu with its own bandwidth, while still leaving pcie bandwidth for traditional storage. They can add capacity to a small system (a 1U server can hold 24 such ssds)
As I see it though, these optane dimms are just fast swap drives which do not offer the same consistent throughput as true memory dimms. They are transparent to the OS which is normally in charge of deciding what stays in memory and what goes to swap. Use of an Intel sdk is required to leverage much of the benefits. It's also Intel-only as far as support goes for now. As usual the marketing doesn't mention the throughput or pricing and instead of comparing to memory they compare it to storage.
http://www.admin-magazine.com/HPC/Articles/NVDIMM-Persistent...
https://thessdguy.com/an-nvdimm-primer-part-1-of-2/
https://www.extremetech.com/extreme/270270-intel-announces-n...
The advantage here being latency (and thus single queue throughput)
PCIe x4 is 3.94GB/s. PCIe x16 is 15.76GB/s.
> DDR4 dimm read throughput is about 20GB/s per dimm
Per channel, not per dimm. Two dimms typically go into a singular channel.
> The advantage here being latency (and thus single queue throughput)
Agreed
No, it's not, that's the whole point of pcie, and dimms share the memory bus anyway, so the whole point is moot.
M.2 slots are a whole another story, they are usually connected to the chipset instead of CPU on desktop motherboards (and finding this information for any particular board is difficult). The chipset is connected with CPU through DMI, which equals to pcie 4x and this is shared with everything on chipset - satas, gigabit ethernet, usb 3...
I would imagine it depends entirely on whether your working set fits in main memory already. if yes, you're not gonna get much speedup from more memory. but if you have to keep swapping from a traditional SSD, it's gonna be hard for faster main memory to make up for that.
Optane is almost entirely hype, with no substance.
As it has been from the very first announcement. Intel is being very misleading marketing Xpoint as "memory". It's not even saying it's "close to memory" now. They're literally calling Xpoint memory (as in the same meaning we use for RAM).
Considering Xpoint has much lower performance than RAM, I'm starting to believe that maybe the FTC should intervene over "false advertising". It's just too bad the U.S. doesn't have as strong laws as Europe does for false advertising.
Years before that they were interfering with compilers and cheating at benchmarks to screw over AMD which was actual, willful, evil deception, but THIS damaged your trust?
supposedly coming out in Q3, i.e. soon
https://www.tweaktown.com/news/58052/amd-epyc-32c-64t-3-2ghz...
Edit: emphasis on threadripper, which is a consumer product.
Servers is also where you get back-and-forth discussions with customers large enough to be worth listening to. Getting useful feedback from the users of desktop CPUs is difficult and time-consuming. But someone like Facebook or Google, who buy thousands of CPUs, and know how to use them, can give you quick and reliable feedback. They can test prototypes and suggest changes within weeks.
It doesn't, and AMD doesn't market Threadripper for server use. It only came up in this discussion because of a misunderstanding of Threadripper's capabilities.
They have gotten greedy and lazy without any competition.
And the reason for that is because their chips/products don't improve as much every year, and Intel is becoming more desperate each year as it faces more competition from Arm, and now AMD, too. So it's trying to "make-up the difference" by being more misleading about it. It's not just the competition either, but the fact that Intel needs to make it sound like their next-gen chips "are so much better" than last year, and make you buy the new ones, too.
I'll give an analogy. Let's say a year ago Intel's CPU got an IPC boost of 10%. But this year it's only getting a 5% boost. Intel will either increase power consumption of the chip to get that 10% or will increase boost speed while keeping the base clock speed the same or lower it for the next-gen chip, just so it can say "the peak performance is once again +10% for this new generation!".
Kaby Lake, Coffee Lake, and whatever Lake is coming next that isn't on 10nm all suffer from this kind of thinking. It's the same kind of thinking that created the whole scandal with the 28-core "5GHz chip". Intel's marketing is all about creating the perception that their chips are significantly better than the last generation. It's why they now have "Xeon Gold", too.
At this point I wouldn't trust any of Intel's announcements until they are verified by trustworthy parties (of which there are only a handful, because most don't go deep enough into their reviews to actually spot what Intel has done to their products to mislead customers).
As others have pointed out, the addition comes from their support of Optane "dimms" which have much higher densities than DDR4 memory in the same form factor.
This is going to give the folks who write schedulers some more variants to play with. First it was per core memory, then it was per socket memory per core, and now it is memory type per socket per core. So you don't want your process running on Core3 with its pages mapped to Optane memory that is attached to the other socket (worst case).
What's wrong with that? High-end workstations often have the same sockets as servers don't they? Did they say it wasn't a server socket?
"What Platform Did It Run On? Intel did not answer this platform directly, however it was clear that the CPU was aimed at the LGA3647 server-based socket given from our examination of the demo system. It was unclear how Intel was going to promote this as an extreme workstation-type system, however Intel did note that they expect only a select market to be interested in this type of processor: a niche of a niche."