64 Core Threadripper 3990X CPU Review
anandtech.com
anandtech.com
https://www.phoronix.com/scan.php?page=article&item=3990x-th...
"When taking the geometric mean of the benchmarks for this article today, The Threadripper 3990X came out overall 26% faster than the dual Xeon Platinum 8280, which is a very nice accomplishment since such a configuration currently retails for $20,000 USD worth of processors alone."
https://manpages.ubuntu.com/manpages/bionic/man8/numactl.8.h...
See https://docs.microsoft.com/en-us/windows/win32/procthread/pr...
OTOH, GPUs are still harder to program than CPUs and I can only imagine what the SFX crew would hear when they have to explain the movie can't be released in summer because the GPUs are too hard to write software for.
Not taking away from those developers, just sharing my experience. When you do get something that works really well (the newest version of Julia makes it easy to write multithreaded code), it's hilarious running htop and seeing 64 threads of green running.
I don't think people fully grasp that lots of applications memory requirements scale with the number of threads. Just running gcc against C make files can burn 1/2 a GB a process, multiply that by the 8 cores most people have and 16G of ram leaves plenty of memory for Firefox/etc.
OTOH, 128GB of ram in a 256 thread machine, is borderline out of ram. Build a project where its more like 2G of ram per process (looking at a lot of recent big public google projects) and its more like I need 1/2 a TB and a bunch of swap to keep the project from OOMing.
So at this point I might use that as a general rule, 2G per thread. So your looking at another ~$1.5k+ or so in RAM for this machine for most purposes.
Yes, running "make -j" on these machines feel amazing at first, but you will soon either 1) run out of memory, or 2) notice that you don't need so many cores anyway due to build dependencies.
We did do the math and decided that for us 4GB per thread is likely safe up front; 2GB per thread might have been pushing it. Another reason to avoid the 3990x (for us) was the tricky scaling of the previous generation's 2990WX; We don't have faith that all of our code would run well on the more memory-channel constrained machine.
I mean, if you're just running a single fairly small build, sure, it's likely overkill...
It sure is fun to see the entire top half of htop being taken up by individual CPU bars, but this is another program whose assumptions have been overtaken by the current silicon: At some point it is not that helpful anymore and makes looking at the processes themselves harder.
Sadly the development of htop seems to have stalled, patches/PRs for that and many other issues remain unmerged and unanswered.
You should have seen it with the Xeon Phi - the Phi could do 4 threads per core and the big ones had 72 of them, for a mind-blowing 288 threads. It was, however, much less forgiving than either of these beasts - L2 cache was Atom-sized and cache misses were not cheap.
If you press 'F2' and look at the available meters you should see half and quarter sizes.
Note that this CPU has 128 threads, so half of that picture. For my usual terminal size I might get 0-2 processes listed below that...
I would love to have this kind of problem :)
_edit_: or just all the deciles and 1%ile + 99%ile
I just wish they could better market themselves to CIO / CTO / CEO rather than prosumer, enthusiast market.
Or they could try and convince Apple to use it on Mac, I am very certain the rest will follow.
[0] https://www.phoronix.com/scan.php?page=article&item=3990x-th...
[1] https://www.anandtech.com/show/15445/amds-fy2019-financial-r...
[2] https://www.anandtech.com/show/15433/intel-q4-fy-2019-result...
Mobile growth was huge, and this was even with the older 12nm chips, and at best, mid-tier laptop models. This year I suspect will be even better as the 7nm 4000 series beats the Intel competition across the board, and as OEMs seem to be finally putting the AMD chips in higher-end models.
Intel continues to have ongoing chip shortages that's impacting OEM bottom lines, so we'll see if that changes the landscape in 2020 as well. https://www.digitimes.com/news/a20200117PD203.html
This surprised me as well, but your explanation makes sense.
I'm curious as to how the new Ryzen 4000s will affect laptop market share. Because offering 8C/16T at 15W is just terrific.
Those top 5 vendors has long term relationship with Intel. So their AMD offering seems to be very lacking or purely as a play to have more bargaining power with Intel. Not to mention Intel's consumer marketing is far better than AMD.
Unless consumer ( not us, but average consumer ) reacts favourably, I dont think it will make much of a difference in terms of volume and unit shipment. It might do well in Gaming Laptop, but it seems most gamers prefer Nvidia.
However, what I really care about right now is higher speed ECC UDIMMs. It's really hard to get ECC UDIMMs at speeds higher than the JEDEC standard.
You could in practice run quite a few VMs before you'd run into trouble, especially if most of them are idle at any given point.
I don't even think there's a benchmark out there for such a test.
Or more realistically if they're usual VMs and idle the vast majority of the time, that should be faster too.
One caveat (at least with how vSphere 5.x worked) is the hypervisor has to claim all CPUs at the same time in order to do work, even if the other guest CPUs are idle. For example, if I have a 4 core VM on a 6 core host, it has to wait for 4 of the 6 to be free before the VM gets to do anything. So sometimes VMs with less CPUs can outperform one with more for the same workload. Getting proper measurements on your loads (peak/avg CPU, memory, disk IOPS etc) is critical to a good migration.
https://www.youtube.com/watch?v=jvzeZCZluJ0
What he did was he took a single 32 core AMD CPU to replace all the computers in his house, including gaming PCs. At around 00:54, he mentions that the cores on the CPU are not "weak cores".
So the answer to the question depends heavily on how cache-resident the problem you are throwing at them is. They'd do great mining bitcoin and be a total disaster as memcached hosts. More typical workloads will be somewhere in the middle.
https://videocardz.com/84697/amd-ryzen-threadripper-3990x-re...
I've not gotten my hands on a TR3. But rome gives you the option of 1, 2 or 3 Nodes Per Socket. Various benchmarks (eg, stream) published by vendors such as Dell show better performance in 4NPS mode. My own testing on Netflix workloads shows a similar speedup.
I run my dekstop (2990WX) in numa mode, exposing the real topology to the scheduler, and I find that it seems to help compilation times.
That's why i'm wondering if 2-NPS or 4-NPS is available with TR3, and what impact it would have.
Its all about cores from now on.
Well, it has been for the last 10 years, but things are going to get silly from now on.
What happened with graphics cores will happen with CPUs now.
https://gist.github.com/cavinsmith/ed92fee35d44ef91e09eaa877...
https://bisqwit.iki.fi/story/howto/openmp/
The two guides I found useful to get started.
Learning high performance computing will let you take advantage of other things too. Like, for example, current desktop workstation processors with "just" four cores.
What computer language will be using to take advantage of that?
If you want low level you can use c, c++ which are both dangerous, rust, zig (I just did some multithreaded stuff in zig and it's fantastically easy).
could you share a little more about your experience here?
Like does zig have its own parallelism/threading in its standard library or do you have to link to a c library like pthreads?
1) A bit about what I'm doing with zig. I am working on an FFI interface between elixir and zig, with the intent to let you write zig code inline in elixir and have it work correctly (it does). https://github.com/ityonemo/zigler/. Arguably with zigler it's currently easier to FFI a C library than it is with C (I'm planning on making it even more easy; see example in readme).
2) The specific not-in-master-branch feature I'm working on now is running your zig code in a branched-off thread. Fun fact about the erlang VM: if you run native code it can cause the scheduler to get out of whack if the code runs too long. You can run it in a "dirty FFI" but your system restricts how many of these you can run at any given time. A better choice is to spawn a new OS thread, but that requires a lot of boilerplate to do and it's probably easy to get wrong. Making it also be a comprehensive part of erlang's monitoring and resource safety system is also challenging, and so there's a lot to do to keep it in line with Zigler's philosophy of making correctness simple.
3) Zig does have its own, opinionated way of doing concurrency. I honestly find it to be a bit confusing, but it's new (as of 6 months) and is not well documented. I believe the design constraints of this are guided by "not having red/blue functions", "being able to write concurrent library code that is safe to run on nonthreaded/nonthreadable systems"
4) The native zig way of doing concurrency is incompatible with exporting to a C ABI (without a shim layer) so I prefer not to use it anyways.
5) Zig ships with std.thread. I believe it's in the stdlib and not the language because some systems will not support threading. But since I'm writing something that is intended to bind into the erlang VM (BEAM), it's probably on a system that supports threading. Also I believe that std.thread will seamlessly pick either pthreads or not-pthreads based on the build target, which makes cross-compiling easy.
6) So yes, figuring this all out is not easy (zig is young, docs are not mature), but once you figure out what you're supposed to do, the actual code itself is a breeze, this is the code that I use to pack information to connect the beam to linux thread and launch it: https://github.com/ityonemo/zigler/blob/async/lib/zigler/lon.... I really hope the docs come with guides that will make this easy in the near future.
I'm relearning c++ right now just because I am building a poker solver for a toy project.
You are right btw, if 32 core cpus become common because of a race to the bottom in prices I imagine there will be a massive increase in demand for programmers with the experience to program massively parallel systems.
edit: and rust :)
That being said, I like it, but I tend to use Python more. I wish I had more of a chance to use Rust in my daily life, but I don't use it at work =/
If you're using Python to invoke highly-optimised native-code, then your performance will be excellent (as shown by the various Python numerical libraries), but performance-sensitive code shouldn't run in the Python interpreter.
As others have said, Python also lacks true multithreading (its threads are capable of concurrency but not parallelism, on account of the GIL), but you do have the option of just running a bunch of Python processes in parallel. I imagine that's a workable solution at least some of the time, but I've never explored this, so I don't know how good the library support is.
Edit: Someone else mentioned 'mpi4py' which seems to be a Python library for multi-process work.
Python is not at all suitable today for parallelism. Which is one reason why languages like Go and Elixir are gaining so much traction.
As a long time Python dev, including work on parallel applications, I have to agree. It's always annoying in Python, and entirely Un-Pythonic.
https://www.python.org/dev/peps/pep-0554/
This just may be the way forward in the Python ecosystem
It's a band-aid. If you want to run Python code in parallel, without large overhead, then CPython is simply not your environment to do so, and Python is not a good choice overall in that kind of endeavour.
Not out of the ordinary to see a SQL Server query using 20+ cores for a query with a parallelized plan
That said, one of my units has a clock speed that doesn't match any of the retail models (I guess they didn't end up selling that model?), and another doesn't seem to work with threading (or whatever AMD calls it) enabled. But that's a small price to pay for the money saved.