Jam 80 Cores, 768GB of RAM into E-ATX Case with This Tiny Board
tomshardware.com
tomshardware.com
I mean I understand why, because the raspberry pi 4 melts all competition into the ground, but why nobody is asking is my concern.
If you need atomic parallelism you should be fine with a 25W Atom 8-core machine that you can passively cool.
Also would be interesting to see how this performs in a atomic parallel scenario? My guess is my HTTP server would not perform so well because the selector thread would not be able to service 79 other cores but I might be wrong about that.
I'm pretty sure the RAM will throttle the 80 cores if they work on a joint problem though!
Someone likely would just docker/vm to partition that 80 cores into micro-service handle httpd proxy and app, db backend.
Those app/VM can easily be converted to app running pi modules.
20 pi modules likely have much better DDR, SSD, Network bandwidth. Probably scale from 2 pi to 200 pi as easily as typical vm setup - and it comes with GPUs for free for those need them.
20 pi modules of 8GBs is only 160GBs.
Ignoring that: Bigger nodes are better in all practical scenarios. With 768 GBs of RAM, this singular big server can likely keep in-memory a large collection of information (ex: all of English-Wikipedia likely fits inside of that RAM).
20x Rasp. Pi cannot access all of English-Wikipedia in RAM. This means that you can't index, you can't search, you can't analyze the pages. Even if you could: you'd need to have a collaborative external memory model, which is not easy to program.
80-cores with access to all 768GBs can do many, many more things than 20x 4-core Rasp. Pi working on only 32GB at a time.
Also, to have a number of benchmarks involve memory, caching, etc.
> GPUs beat everything
that’s workload dependant, and also subject to support, no? Doesn’t GPGPU basically mean CUDA currently, and therefore beholden to Nvidia support for your hardware/software platform?
It seems like RDNA (5xxx and 6xxx cards) are less supported, but reports are that OpenCL kinda-sorta works (RDNA cards have a very different architecture than Vega / CDNA)
Depends. If the programmers take the path of least resistance, sure it gonna be CUDA. Personally, I made quite a few things on top of DirectCompute i.e. vendor agnostic.
GPUs win in TFlops (like 20 or 40 TFlops these days), and TFlops/watt (A modest 20 TFlop GPU these days would be under 400W).
GPUs also have 1TBps memory bandwidth thanks to HBM2 (on the high end), or at least 500GBps (thanks to GDDR6 on the low end).
Any serious compute problem with "obvious" parallelism will run on a GPU these days. CPUs are for sequential problems... (and running many sequential problems in parallel: which GPUs are kinda bad at actually. No GPU would ever be able to serve web-pages like a CPU: branch divergence is just too high)
The headline sounds like a brag, but isn't exactly impressive.
There have been multiple EATX Dual EPYC motherboards on sale for some time now. The limited board area means only one DIMM per channel, but with eight channels and 128GB DIMMs that still means you could have 128 cores and 2TB of RAM in a single EATX system with "ordinary" AMD hardware.
An example that might solidify the idea: pack a Wikipedia snapshot into it for search, and serve ~1m queries per second on it (12k qps per core).
Naively, assuming identical instruction sets (I know they're not), 16 threads at 4 GHz is less than half as good as 80 at 2 GHz. But that can't be the whole story.
ARM only has 128-bit SIMD through NEON. Its reasonably well designed, but nothing beats the brute force of just doing 256-bits at a time (or 512-bits in the case of Intel's AVX512)
No. This core is terrible at encoding.
EDIT: And encoding is limited to ~16 cores in practice. It seems like after that, the communication between threads get too much to be useful. Unless you plan to be doing 5-simultaneous encodings at a time, then you're gonna have to find something else to do with all those cores.
For a decently sized video (say a TV episode) there's usually like 100 split points to divy out to encoders.
2 Intel(R) Xeon(R) CPU E5-2690 @ 2.90GHz
768 GB of RAM (384 GB per Processor)
18 TB of SSD Storage (Intel 3 Series or Samsung 840 Pro Enterprise Series)
If the vendors were ready for something like this on the software side, this would be great for edge compute when low latency response is required - remote utility substation handling and reacting to a large array of sensors feeding at 60 data points per second. In some use cases going to the control centre and back would be too slow to benefit. Basic grid control is well handled, but I could see optimizations benefiting from this. Vendors and utilities are way behind on this though.
The dataset is 2-3B records with 5-12 64-bit values each, stored in a few dozen files using the Apache Arrow format. If we take the midpoint of this range that's 170 GB just with raw data. With the overhead of data structures, I was running the process with ~400 GiB of RAM and could have done more on a beefier machine.
It took about 20-30 minutes to run the full algorithm on these tens of billions of data points and this approach was perfect for this use case. No overhead of Spark and all of its dependencies, just one program, a bunch of input files, and it's done when I get back from lunch.
I'd assume that any workload that benefits from 80 cores would benefit from 160-threads (on those 80 cores). Apple's decision to avoid SMT on M1 kinda-sorta makes sense, from the perspective that phones probably don't have throughput-sensitive workloads like servers.
But if databases / other systems with lots of I/O or RAM-heavy wait times start coming up, surely SMT would easily improve performance without much costs in area or power?
---------
It seems like the lower-power E1 core (Efficiency core) has SMT. So the ARM / Neoverse team has the experience to bring SMT should they desire it. This suggests that there's some design reason they left SMT off the table.
The N1/N2 cores are more "general purpose", so I'd assume that they'd see more workloads than E1. If E1 benefits from SMT, why not N1/N2?
I'm not current on this one, but I recall when it first came out that disabling hyperthreading was the only solution.
Has it been solved yet, or are some chipmakers avoiding enabling hyperthreading now as a result?
That's the confusing thing: they have the tech on one core, but not the other.
Seems like its ideal workload is lots of compute on a cache-friendly quantity of data.
If you hit a memory-bandwidth bound on 80 threads, there's no point going up to 160 threads.
In most situations, I expect code to be memory-latency bound on a single thread. (Ex: node = node->next style traversals are quite common, and you cannot progress until the memory has responded). This is exceptionally common in interpreted code (Java, Javascript, PHP, Python), especially OOP-code.
So your 80 cores are sitting there waiting for RAM-latency to respond. Wouldn't it be nice if they could execute 80-other-threads in parallel while waiting? This converts a RAM-latency problem into a RAM-bandwidth problem.
------------
ONLY problems that are memory-bandwidth bound on 80 threads will benefit from this architecture.
OTOH, if your reorder buffer can’t keep your backend full, adding threads may be cheaper in terms of silicon area.
The Neoverse E1 is a 2-wide decode core with SMT (wtf??).
The Neoverse N1 is a 4-wide decode core (1-thread per 1-core).
-------
You're right in that the N1 is narrower than Skylake / Zen. But N1 isn't too shabby: it has 8 execution pipelines: 1 branch, 3x 64-bit integer, 2x 128-bit vector, and 2 load/store.
Furthermore: the core that ARM decided to shove their SMT-effort into is the E1, which is probably 1/2 the size of a N1 (well, at least 1/2 sized decode).
It's within 1.5x of Cortex-A55 perf on single threaded workloads, and with efficiency as the mantra, SMT was worth it there. (but we'll see what happens in future designs... there's both the in-order Cortex-A5xx and the OoO Cortex-A6x which is Neoverse E for the power efficient role)
I keep seeing talk about Ampere. Can I buy from them or do I need to be a business or some sort and speak with a rep?
nvm: they're not being sold quite yet.
> The Ampere Altra Q80-28 SoC with 80 cores runs at 2.80 GHz and consumes around 175 Watts.
175w is nothing to scoff at, but also isn't completely ridiculous. GPUs will end up consuming more than that in a typical ATX gaming machine.
Even ITX cases without riser cables end up positioning the GPU right at the bottom of the case, so it gets fresh air (https://cdna.pcpartpicker.com/static/forever/images/userbuil...)
Though my inner troubleshooter is thinking "oh man more connections to re-seat"
That said, you really only run into that kind of stuff when you're entering hobbyist mode. If what someone wants is a compact gaming PC, there are cases like the CoolerMaster NR200 (https://i.redd.it/fq9y7mevznb71.jpg) which are affordable and just as easy to build in as any mid-tower. The only difference is you'll be using a mini ITX motherboard, and an SFX power supply.
If I'm doing something computationally expensive, I can feel the room heatup for sure: just 2 degrees F in practice (~1C or so). Enough to notice + enough to see it in my thermometer I keep in my room... but not enough to be worried about anything.
And honestly, I don't spin up all 16c/32 threads that often.
If you're worried about your components overheating in a small case, there are absolutely ways to cram a ton of performance into a small package. You can fit top of the line gaming hardware in as little as 10-15 liters of case volume.
If you're worried about your PC becoming a space heater, you'll need to go with less power-hungry components, but you can still absolutely build a tiny and capable gaming PC.
N1 cores are weaker so you should really compare 80 threads vs 80 threads which would be a single-socket Epyc or Xeon which fit into the same or smaller board.
> similar RAM
IIRC, EPYC is 2TBs of RAM support from LRDIMMs. Maybe 4TB now, but you need many, many DIMMs for that: like 16 DIMMs or something on a dual-socket EPYC.
64 core/128 thread micro atx https://www.newegg.com/asrock-rack-romed6u-2l2t-amd-epyc-700...
Bonus board just because it's badass: https://www.newegg.com/asus-pro-ws-wrx80e-sage-se-wifi/p/N82...