Can My Water Cooled Raspberry Pi Cluster Beat My MacBook?
the-diy-life.com
the-diy-life.com
Ryzen 2700 launched at $300 (and tapered down to ~$200) and would pretty much run circles around such a Pi cluster.
Just saying. These Pi clusters can be a cool and fun thing to build but if you're looking for compute power, you'd be better served by a mid-range desktop.
So eight Pies alone do not make a complete cluster, just as one Ryzen alone does not make a complete computer. A few times eight times cheap can turn out to be quite expensive, depending on how much cheap actually is.
I haven't paid much attention to component prices recently but my rule of thumb for budget builds is to start with $80 for each component that isn't a GPU or CPU. Go with stock cooler, grab 80 for PSU, 80 for mobo, 80 for RAM, 80 for storage, see where you end up. In the same ballpark, but the real computer will pack a whole lot more punch (even if you had less RAM total).
[0] https://www.jeffgeerling.com/blog/2020/raspberry-pi-cluster-...
> It's slightly more cost-effective and usually more power-efficient to build or buy a small NUC ("Next Unit of Computing") machine that has more raw CPU performance, more RAM, a fast SSD, and more expansion capabilities.
But is building some VMs to simulate a cluster on an NUC fun?
I would say, "No." Well, not as much as building a cluster of Raspberry Pis!
---
You don’t build a RPi cluster for speed or cost or efficiency. You do it because it’s fun. And that’s okay.
For my region, the actual distributors start at 44 EUR for 2GB and 64 EUR for 4GB.
As always, a benchmark is the only thing that will prove it.
If you care about benchmarks, a Pi4 will score around 200 points (1 core) or 550 points (4 core) at 1.5GHz on GB5. At 2GHz you can reach around 700 points multi-core.
Ryzen 2700 will easily go above 6000 points in multi-core bench without overclocking.
So even in a theoretical, embarrasingly parallel workload with minimal sharing between cluster nodes, the Zen will be faster than eight pies. It's not even a contest if you also need some I/O, shared memory, or heavy SIMD.
The video in the OP showed 8 pis, and assuming 4 cores each that's 32 cores total.
A raytracing benchmark would be interesting because you could divide the work up-front and not have to worry about communication between nodes in the cluster, and each node doesn't need that much memory.
Maybe the ryzen 2700 beats the 8 PI cluster, but clearly there's an X number of PIs that will beat the ryzen. Maybe it's X = 10 PIs, or 20, or 200? Idk, but it could also just be 8. There's no way to know for sure without a benchmark.
Also: by the time you've wired up 8 raspis you end up using quite a bit of power just to connect them all together with a switch.
Raspberry Pi 4s need a maximum of 15 watts each. So 120 watts just for the computers. Even if you discount the power consumption of the switch, my 240 watt Ryzen computer is still going to beat that joule-for-joule.
Edit: one more thing, that 240 watt system also powers a 75 watt GPU, so it's definitely more wattage than really required for the CPU alone.
https://www.raspberrypi.org/documentation/hardware/raspberry...
https://www.pidramble.com/wiki/benchmarks/power-consumption
https://raspi.tv/2019/how-much-power-does-the-pi4b-use-power...
As an example: even if I halted my desktop at 0 MHz and it still magically took the same time to calculate the first million digits of pi as a raspberry pi, it still would be using far more power.
If that is the case, then M1 should be the slowest CPU Apple has used in the past 10 years.
Apple M1: https://browser.geekbench.com/v5/cpu/search?utf8=%E2%9C%93&q...
Raspberry Pi 4: https://browser.geekbench.com/v5/cpu/search?utf8=%E2%9C%93&q...
There's also no control for the thermal throttling of the M1, which is probably why the 100,000 example is performing worse.
there's no control because that's part of what he's measuring.
Find all primes up to: 10000 using 256 processes.
Time elasped: 0.51 seconds
Number of primes found 1229
Find all primes up to: 100000 using 256 processes.
Time elasped: 36.71 seconds
Number of primes found 9592
Find all primes up to: 200000 using 256 processes.
Time elasped: 149.55 seconds
Number of primes found 17984
EDIT: ran it again: Find all primes up to: 200000 using 256 processes.
Time elasped: 145.76 seconds
Number of primes found 1798410k using 48 processes: 0.82 seconds (using the original single threaded script this was actually faster at 0.65)
100k using 48 processes: 28.41s
200k using 48 processes: 99.68s
--- EDIT - looking at resource monitor python.exe is only using 7-8% of total available CPU resources
--- EDIT 2 - switching to ThreadPool from multiprocessing.dummy brought the 10k result down from 0.8 seconds to 0.3 seconds, but didn't impact the 100k or 200k results
Find all primes up to: 10000 using 256 processes.
Time elasped: 0.36 seconds
Number of primes found 1229
Find all primes up to: 100000 using 256 processes.
Time elasped: 15.7 seconds
Number of primes found 9592
Find all primes up to: 200000 using 256 processes.
Time elasped: 58.07 seconds
Number of primes found 17984
(this was debian 10)https://github.com/joshjerred/mpi4py-with-multiprocessing-Ch...
To me this would be more fair comparison between single cpu and pi cluster.
Find all primes up to: 10000
Nodes: 1
Time elasped: 0.12 seconds
[1229]
Primes discovered: 1229
Find all primes up to: 100000
Nodes: 1
Time elasped: 4.52 seconds
[9592]
Primes discovered: 9592
Find all primes up to: 200000
Nodes: 1
Time elasped: 17.43 seconds
[17984]
Primes discovered: 17984
I even tried: Find all primes up to: 1000000
Nodes: 1
Time elasped: 383.2 seconds
[78498]
Primes discovered: 78498I found this: https://setiathome.berkeley.edu/cpu_list.php
ARMv7 looks like is raspberry pi, looking at this I can't see any computational value for such cluster setup.
Using numeric arrays chunked into blocks of number ranges would be more efficient (and therefore "crunchier")
[this is from my recollection]
* completely in-register only, no caches. they warn an organic workload will never ever do this and not to do it if you're not confident in your cooling * somewhat more normal math-heavy workloads that don't do any IO but also do access the caches like a normal person * etc
The python benchmark might be fair, though? I wouldn't be surprised if the ARM chip on the pi is like, fast, but then when it came to doing something more holistic - like a bunch of python vm operations - something about the mac is more robust than the pi in a significant way.
Nevertheless this is not as interesting as testing the M1 chip on the latest MacBook offering. I feel a bit misled but perhaps it was just my fondness for the M1 causing this bias.
The reason that 1 is not a prime is to preserve unique factorization. A basic fact from number theory is that every integer uniquely factors as a product of primes, say 21 = 7 times 3. If 1 were a prime, we'd also have 21 = 7 times 3 times 1. That's bad.
In more general number rings, other units (e.g. i) are also not considered primes for the same reason. The exclusion is not arbitrary.
#0 and 1 are not primesThe original code includes 1 as a prime.
He could probably get even better than the Pi cluster by using a (single) GPU.
I wonder what performance would look like on 5 years old 16 core CPU from ebay for like $20, compared to py cluster.
0.8 seconds for 10,000
28.16 seconds for 100,000
10,000: 0.52 seconds.
100,000: 41.48 seconds.
200,000: 157.68 seconds.
If you meant to ask the other commenters here, then something like "Does anyone have an M1 Mac? If so, would you run his benchmarks and post the results here?" would have been clearer.