AMD Unveils “EPYC” CPUs Featuring Up to 32 Cores and 64 Threads for the Datacenter
wccftech.com
wccftech.com
A good summary (with pictures) by a reddit user: https://www.reddit.com/r/Amd/comments/6bjvy6/amd_2017_financ...
I have met software engineers who could not believe me when I pointed out that multi-threading was invented, made sense, and, in fact, was thriving - on single-core computers (with no hyper-threading)!
If I have an IDE, 2-3 VM's, a couple of browsers, continuous integration running in the background, webpack/ts-loader with off thread hinting I can easily have 4-5 processes running that all benefit from having a full core to play with.
It's for that reason when I had to build a new desktop for the new job I went with the Ryzen 1700, each core isn't that important (as long as it's comparable with the core in my current jobs i5-3570K which it broadly is), it's having eight of them.
For developers I think more cores is still better (I'd add Spotify and Slack to your list of things that are always running) and yet we still prefer shiny laptops to powerful desktops.
Laptops have horrible postural positions unless you use desktop monitors at which point why not just use an actual desktop which will annihilate the laptop on performance anyway.
My laptop gets switched on maybe once a month.
I just imagine it will make it easier for intel's marketing to imply these are toys rather then true enterprise grade parts.
Sora on the other hand means "gravel" in Finnish. Suprisingly there are few Finns named Sora too.
https://en.wikipedia.org/wiki/Explicitly_parallel_instructio...
or maybe have a power consumption tracking dongle that sends the information back to MS for billing while the MS techs just check if those are connected to the correct machine.
But these days Microsoft is also beginning to exclude third-party browsers from its store, so I guess they don't care they're repeating the same violations all over again. Too much money on the line, when the worst case scenario is a slap on the wrist financially and some "monitoring" by federal agencies.
Sounds like an IBM machine
Sure, you can build a RAID that fast already though.
Few HDFS data sets stored in RAM or some more Data Science Spark workloads.
A few hundred megabytes for our "enterprise" financial software...
http://www.oracle.com/us/products/servers-storage/sparc-m7-1...
This is just the mainstream catching up.
I was under the impression SPARC was on life support
http://www.ebay.com/itm/182537878712
Of course that's just RAM. If you want a server with it, it's not much more, just over $1K shipped:
That said I still want my OS to fit on a fdd and have 64k demos generating the universe.
More ordinary servers typically stop at 6TB since one Xeon can support 1.5TB so an ordinary quad socket board typically with 96 DIMM slots can go up to 6TB with 64GB DIMMs. You can configure a machine like this at http://www.thinkmate.com/system/superserver-4048b-tr4ft and see that for the relatively low price of $110K you can get a machine with 6TB of RAM.
It sounds like you know what you're talking about, so I'm sure it was inadvertent that you wrote "physical RAM limit of the Linux kernel". It's primarily the x86-64 architecture and 5-level paging is coming which extends the linear address space to 57 bits (128 PiB) and the physical address space up to 52 bits (4 PiB). Still, one wonders how long it will take for that to be inadequate.
https://software.intel.com/sites/default/files/managed/2b/80...
Merge 5-level page table prep from Kirill Shutemov:
"Here's relatively low-risk part of 5-level paging patchset. Merging it
now will make x86 5-level paging enabling in v4.12 easier.(It would be interesting to know whether this backplane-based system actually manages to achieve 4.8 GHz QPI speed (9.6 GHz symbol rate), or whether the physical aspects limit it to lower speeds. Given that processors only have four QPI links this would further increase communication overhead -- in a 8-way system some sockets are separated by two hops).
MooseFS (a scale-out storage system) keeps metadata in RAM for low latency and is known for pushing the limits of commodity hardware in big installations. I presume the same applies to any kind of low-latency database of similar size.
mere mortals at what point is there so much stuff in memory that you need a couple hours of hold up time just to flush it out to SSD
If your buying these machines with multiple TB of RAM, then buying flash arrays that can drive multiple GB/sec of IO bandwidth shouldn't be a problem either. At say 8GB/sec write bandwidth flushing 1TB is a little over two minutes. Although, why you have that much "dirty" data in RAM might be another question. A database machine using that amount of RAM is going to care more about the IOP rates of the disk, so that its flushing updates to disk at the same rates they are arriving. Meaning that the RAM won't need to be flushed to disk if the machine/power/whatever fails. Disk arrays with >1M IOP/s have been around for over a decade, and given a SAN can be wired together to increase aggregate performance.http://pro.radeon.com/en-us/frontier/
13 TF FP32, 16GB HBM2 RAM (480 GB/s).
Initial leaks point to a PCIe 3.0 64x interconnect
</rumor>
Well we do already know that the interconnect between modules uses the same PHYs as PCIe 3.0.
With this extension Zen does SHA1 @ 2 cpb, SHA-256 @ 3 cpb and SHA-512 @ 2 cpb (off the top of my head). (All of which are faster than the fastest BLAKE2 implementation I know on Haswell).
I've wondered about trade off between SHA256 vs BLAKE2. In the future there'll be no debate since more and more computers will have SHA instructions. But right now I'm wondering about the speedup of BLAKE2 vs SHA256 with hardware. On the other hand, many computers, especially servers don't have SHA2 instructions for the foreseeable future which will make BLAKE2 a very good option.
However, none with SHAEXT; they just weren't there yet. But the Zen numbers should give you a good idea.
Note that these benchmarks are made using a plain C implementation of BLAKE2 (the reference one), which is not vectorized by any compiler. The fastest (AVX2) BLAKE2 implementation is about 40 % faster than the scalar C implementation (on Haswell).
As far as I'm aware no mainstream crypto library ships optimized BLAKE2 versions. I believe some Go packages do/did make up their own version (not the one from Samuel Neves), but at least one of them mixed SSE and VEX/AVX insns with the predictably bad results (60 MB/s or so) - perhaps this is fixed by now.
So in summary, BLAKE2b is imho the best candidate on perf, and if you use a good implementation it should be within ~30% of SHA2 (512) with SHAEXT — with the numbers we have so far. I understand that Zen's aggressive (=good) power mgt makes it somewhat difficult to benchmark hot loops consistently, so we'll have to wait and see for practical results, I guess.
http://www.amd.com/en-us/innovations/software-technologies/s...
Think of it as AMD's answer to Intel SGX, albeit with a quite different design.
As such this would be ideal for things like VDI and web / cloud hosting where the quantity of VMs is very high, but the load from each is typically not.
Which is why every virtualization platform out there lets you oversubscribe CPUs. That's a solved problem, what's the benefit of having 100 VMs run on 32 slow cores vs 16 fast ones?
As mentioned above, context switching is expensive and extra L1 cache is valuable. Time-sharing can also have a huge effect on latency (because requests must wait until their server is scheduled), even when the throughput is still good.
Even if time-sharing performs well most of the time, when it goes wrong, the performance problems can be opaque and hard to debug. In general solution that "really" does something will save engineer-days as compared to a thing that does it at the same price/performance trade-off, but virtually.
Intel does similar, the E5-2650 v4 is 2.2 GHz to 2.9 GHz.
As an example, the 22-core Intel Xeon E5-2696v4 has a base clock of 2.2GHz. With one or two cores active, it can turbo up to 3.7GHz. With three cores active, the maximum is 3.5GHz, and it decreases by 100MHz per active core until ten cores are active. With 10 or more active cores, the limit is 2.8GHz, provided that the chip is still within its power and thermal limits.
So, if you don't need the high peak clock speeds, 32 half-speed cores would be preferable for datacenters to save money on electricity and cooling design.
Regarding HT, 2 threads is really a sweet spot for a 4-wide CPU. More than that and the competition for cache resources, execution units and register file become significant.
POWER8 is special because a factor of x2 is because each power 'core; is pretty much two distinct smaller cores that can gang together to speed up one thread (it also helps on per core software licensing), while the other x2 factor is for very specialized loads (this is also true, or used to be, for SPARC).
IIRC XeonPhi which is also a specialized cpu has 4xHT.
--------------
Xeon E5 2699 v5
32C/64T @ 2.30 GHz
L1 Instruction Cache: 32 KB x 32
L1 Data Cache: 32 KB x 32
L2 Cache: 256 KB x 32
L3 Cache: 46080 KB
--------------
AMD EPYC
32C/64T @ 1.4 GHz
L1 Instruction Cache: 32 KB x 32
L1 Data Cache: 64 KB x 32
L2 Cache: 512 KB x 32
L3 Cache: TBA
--------------
It's going to be interesting to see what the performance per dollar amounts to on both sides.
I'm guessing the Xeon will land somewhere around $4K. Although, I have no idea about the EPYC.
Regarding the L3 cache: if nothing changed from Ryzen it will be 64 MByte.
Wouldn't AMD need to license Thunderbolt from Intel? There are repeatedly rumors about a license agreement between Intel and AMD regarding AMDs GPU IP, if that is true maybe they get access to Thunderbolt.
[1] https://thunderbolttechnology.net/contact/thunderbolt-develo...
[2] http://www.intel.com/content/www/us/en/processors/xeon/xeon-...
AM4 is PGA and 1331 pins. :)