Simple test, `memcopy.c`:
void memcopy(double* restrict a, double* restrict b, long N){
for (long n; n < N; n++){
a[n] = b[n];
}
}
You can check assembly with:
gcc -march=skylake-avx512 -O2 -ftree-vectorize -mprefer-vector-width=128
-shared -fPIC -S memcopy.c -o memcopy128.s
gcc -march=skylake-avx512 -O2 -ftree-vectorize -mprefer-vector-width=512
-shared -fPIC -S memcopy.c -o memcopy512.s
to confirm they use `xmm` and `zmm` registers, respectively. (The 512 version also uses xmm to finish off a remainder. Ideally, it would use masked load/stores, but I haven't seen any auto-vectorizers actually do that.
I'm using gcc 8.2.1. I believe the option `-mprefer-vector-width` was added recently.
Now I compiled both into shared libraries (drop the `-S` and choose an appropriate file name), and benchmarked from within Julia.
julia> using BenchmarkTools, Random
julia> memcopy128!(a, b) = ccall((:memcopy,
"/home/chriselrod/Documents/progwork/C/libmemcopy128.so"),
Cvoid, (Ptr{Cdouble},Ptr{Cdouble},Clong), pointer(a), pointer(b), length(a));
julia> memcopy512!(a, b) = ccall((:memcopy,
"/home/chriselrod/Documents/progwork/C/libmemcopy512.so"),
Cvoid, (Ptr{Cdouble},Ptr{Cdouble},Clong), pointer(a), pointer(b), length(a));
julia> b = randn(32); a = similar(b);
julia> b = randn(32); a = similar(b);
julia> all(a == b)
false
julia> memcopy128!(a, b); all(a == b)
true
julia> randn!(a); all(a == b)
false
julia> memcopy512!(a, b); all(a == b)
true
julia> @btime memcopy128!($a, $b)
7.517 ns (0 allocations: 0 bytes)
julia> @btime memcopy512!($a, $b)
5.277 ns (0 allocations: 0 bytes)
julia> b = randn(64); a = similar(b);
julia> @btime memcopy128!($a, $b)
10.980 ns (0 allocations: 0 bytes)
julia> @btime memcopy512!($a, $b)
7.300 ns (0 allocations: 0 bytes)
julia> b = randn(100); a = similar(b);
julia> @btime memcopy128!($a, $b)
15.060 ns (0 allocations: 0 bytes)
julia> @btime memcopy512!($a, $b)
14.031 ns (0 allocations: 0 bytes)
julia> b = randn(200); a = similar(b);
julia> @btime memcopy128!($a, $b)
33.867 ns (0 allocations: 0 bytes)
julia> @btime memcopy512!($a, $b)
19.546 ns (0 allocations: 0 bytes)
julia> b = randn(400); a = similar(b);
julia> @btime memcopy128!($a, $b)
53.923 ns (0 allocations: 0 bytes)
julia> @btime memcopy512!($a, $b)
28.376 ns (0 allocations: 0 bytes)
julia> b = randn(800); a = similar(b);
julia> @btime memcopy128!($a, $b)
98.277 ns (0 allocations: 0 bytes)
julia> @btime memcopy512!($a, $b)
57.960 ns (0 allocations: 0 bytes)
julia> b = randn(2_000); a = similar(b);
julia> @btime memcopy128!($a, $b)
232.138 ns (0 allocations: 0 bytes)
julia> @btime memcopy512!($a, $b)
133.229 ns (0 allocations: 0 bytes)
julia> b = randn(20_000); a = similar(b);
julia> @btime memcopy128!($a, $b)
3.688 μs (0 allocations: 0 bytes)
julia> @btime memcopy512!($a, $b)
2.759 μs (0 allocations: 0 bytes)
julia> b = randn(200_000); a = similar(b);
julia> @btime memcopy128!($a, $b)
112.598 μs (0 allocations: 0 bytes)
julia> @btime memcopy512!($a, $b)
117.437 μs (0 allocations: 0 bytes)
The advantage wasn't very impressive here, but it persisted for vectors of length 20,000. At 200,000 we saw the memory bottleneck hit in force, when it busts the per core L2 cache. The 7900x's per core L2 cache is 1048576, which translates to 131,072 doubles.