How much memory bandwidth do large Amazon instances offer?
lemire.me
lemire.me
1 51.1
2 99.4
3 108.2
4 121.3
5 116.4
6 122.9
7 122.0
8 123.2
9 125.6
10 124.9
Xeon E-2288G (aka i9-9900K) 128GB DDR4-2666: $ ./a.out
1 19.7
2 27.8
3 30.0
4 31.7
5 33.6
6 34.4
7 33.9
8 33.2
9 32.2
10 32.2
11 32.3
12 32.1
13 31.9
14 31.5
15 31.2
16 30.9If you have non uniform memory, you may be shooting yourself in the foot. Author used a dual socket system. If you are not careful, and your code is not NUMA aware, you may end up with less memory bandwidth for your application, than a single-socket would have. Can be avoided by prefixing `numactl --cpubind=0` to your command
1 59.9
2 109.1
3 108.4
4 109.1
5 108.0
6 109.3
7 109.4
8 110.5
9 119.3
10 123.9Another factor is mac's have many more memory channels on the m1/2/3 pro and m1/2/3 max. So they can have many more memory transactions, where on most PCs and x86-64 laptops you get 2 cache misses per memory latency (on the order of 70ns), where macs will get many more. So under various workloads all the cores spend much less time waiting on memory latency.
So is this benchmark really useful? Seems like it's engineered enough that it probably isn't relevant for most use cases (like running java, go or python services), and not engineered enough so it's not relevant for performant systems like databases. (That said NUMA simply isn't that hard, even with Java which is also NUMA aware).
1 12.9
2 12.4
3 12.1
4 11.6
E5-1620 v2 @ 3.70GHz, 32Gb 1 16.9
2 29.5
3 39.8
4 37.8
5 35.3
6 35.3
7 34.9
8 34.7 1 3.9
2 3.9
3 3.7
4 3.6
RasperryPi CM4 8GB with data volume set to 4GB (aarch64) 1 4.7
2 4.1
3 3.8
4 3.7The r6i.metal instance says up to 50 Gbps on Network bandwidth: https://aws.amazon.com/ec2/instance-types/r6i/
You could go extreme and start pinning your threads to specific CPUs to tune the code manually instead of relying on the kernel if you know memory bandwidth is extremely important and you won't have much CPU contention to worry about in terms of getting work scheduled in a timely manner.
All that being said, you typically also need to design your application from the ground up to be NUMA aware to take full advantage so that you can set up your allocations to happen on the right zone & whatnot.
PS E:\> .\bandwidth.exe
1 54.0
2 56.7
3 56.1
4 55.3
5 54.6
6 54.3
7 53.9
8 53.1
9 53.4
10 53.1
11 52.3
12 52.8
55 is kinda low for DDR5 I guess..On S23 Ultra (Snapdragon 8 Gen 2+, 12GB LPDDR5X @4200MHz)
u0_a339@localhost ~> ./bandwidth
1 34.4 GB/s
2 38.7 GB/s
3 42.0 GB/s
4 42.3 GB/s
5 42.4 GB/s
6 40.2 GB/s
7 41.0 GB/s
8 40.3 GB/s
On a NanoPi-R6C (RockChip RK3588s @2.4Ghz + 8GB LPDDR4X) pi@nanopi ~> ./bandwidth
1 22.1 GB/s
2 26.2 GB/s
3 27.5 GB/s
4 27.7 GB/s
5 27.1 GB/s
6 27.1 GB/s
7 27.0 GB/s
8 26.9 GB/sRunning on an 2xAMD EPYC 7763 64-Core Processor, with 256GB of ram and 256 threads (EuroHPC LUMI standard CPU node): ~38.
Running on my Macbook Pro M1 16GB on battery power, but with 8GB of data to not swap as I have other apps running: ~59.
On Godbolt, with a size of 100MB (due to limits), this improves the speed using LLVM 17 (https://godbolt.org/z/6dW1h8aev) from:
1 15.1
2 15.9
To (https://godbolt.org/z/sh3489Mxv) 1 12.8
2 13.7 1 54.7
2 50.6
3 49.6
4 48.2
5 47.9
6 47.4
7 47.1
8 46.6
9 46.5
10 46.2
11 46.1
12 45.9
13 45.8
14 45.7
15 45.7
16 45.7
17 45.7
18 45.8
19 45.9
20 45.8
21 45.8
22 45.6
23 45.6
24 45.5
25 45.5
26 45.5
27 45.5
28 45.4
29 45.4
30 45.4
31 45.4
32 45.4 1 52.6
2 78.1
3 71.0
4 74.3
5 71.3
6 72.5
7 70.0
8 69.6
9 68.4
10 68.7
11 68.5
12 68.3
13 68.3
14 68.0
15 67.8$ ./a.out
1 59.2
2 76.9
3 66.8
4 68.7
5 64.0
6 67.1
7 63.9
8 66.0
9 64.0
10 65.6
11 64.0
12 65.5
13 65.6
14 66.0
15 66.0
16 65.8
17 65.1
18 65.2
19 65.1
20 65.4
21 64.6
22 65.2
23 65.5
24 65.2 Model name: Intel(R) Xeon(R) CPU @ 2.20GHz
CPU family: 6
Model: 79
Thread(s) per core: 2
Core(s) per socket: 8
1 11.4
2 21.3
3 30.8
4 39.5
5 48.2
6 56.8
7 64.3
8 71.0
9 61.2
10 64.1
11 67.7
12 72.4
13 74.3
14 78.3
15 81.8
16 84.916-core ARM CAX41
root@ubuntu-32gb-fsn1-1:~# ./a.out
1 17.8
2 35.6
3 49.9
4 64.9
5 79.0
6 92.1
7 101.7
8 110.2
9 117.6
10 124.0
11 130.4
12 136.2
13 140.9
14 143.8
15 149.2
16 152.9
16 Core AMD shared CPX51
root@ubuntu-32gb-fsn1-1:~# ./a.out
1 18.9
2 37.5
3 54.3
4 65.4
5 77.2
6 92.8
7 100.8
8 92.1
9 95.1
10 105.7
11 93.4
12 100.9
13 89.9
14 97.0
15 99.1
16 107.8
16 core AMD dedicated CCX43
root@ubuntu-64gb-fsn1-1:~# ./a.out
1 36.3
2 71.6
3 56.8
4 50.8
5 63.4
6 57.2
7 53.4
8 50.9
9 55.2
10 61.3
11 64.4
12 65.3
13 66.5
14 68.8
15 69.2
16 64.4
Increasing the data volume for CCX43 to 48GB increases the bandwidth using 16 cores to 75.1. Memory bandwidth seems to scale pretty well with the number of cores on ARM. Interesting that the bandwidth for the shared AMD system is that much higher than the dedicated system
1 37.4
2 73.3
3 107.3
4 141.4
5 171.6
6 199.5
7 226.0
8 251.1
9 235.4
10 243.0
11 264.5
12 281.9
13 303.7
14 323.0
15 339.6
16 354.4
17 299.0
18 286.3
19 300.9
20 310.6
21 325.7
22 339.2
23 352.6
24 364.3
25 305.8
26 309.0
27 319.6
28 326.5
29 335.4
30 345.5
31 356.7
32 364.9
And then it settles around there.Another 1.5 times greater speed comes from DDR5-4800 vs. DDR4-3200.
The rest may be from virtualization and other overheads.
1 27.6
2 50.5
3 62.5
4 68.3
5 75.6
6 82.6
7 87.4
8 90.7
9 93.0
10 94.3
...repeating
I expected a bit more, my guess would have been around 120 GB/s. I've been playing with LLMs and this hardware is about as fast as memory gets on consumer Intel without overclocking.Screenshot showing usage: https://i.imgur.com/okdqgG9.png
1 20.9
2 33.4
3 35.7
4 36.2
5 35.2
6 35.4
7 34.7
8 34.9
9 34.0
10 34.2
11 34.4
12 34.3
13 33.8
14 33.6
15 33.4
16 32.9
I expected 51.2 GB/s (2 x 3200 x 1e6 Transfers/s of 8 bytes each)Running the test on a desktop Zen 3 (5900X) with slower ECC DDR4-2666, i.e. with a maximum throughput of 42.7 GB/s, provides for 2 or more threads a throughput around 39 GB/s, with a maximum of 39.5 GB/s at 4 threads.