ARM64 Popcount in Golang and Assembler
barakmich.dev
barakmich.dev
TEXT ·hasPOPCNT(SB),NOSPLIT,$0
XORQ AX, AX ; AX <- 0
INCL AX ; AX <- 1
CPUID ; check cpu caps
SHRQ $23, CX ; shift bit 23 to bit 0 (bit 23 => POPCNT support)
ANDQ $1, CX ; examine bit 0
MOVB CX, ret+0(FP) ; return 1 iff POPCNT instruction supported
RET
It seems Intel CPU's have such a bloated and varied instruction set you have to query to see what flavor of CPU you have to create portable code. Writing compilers for this must be a real pain.Full disclosure : I work for Intel
This is equally applicable to other x86 vendors.
The alternatives are either (a) never introducing new instructions, or (b) making portable binaries impossible. Neither alternative seems better? It doesn't look quite as tortured in C:
bool has_popcnt() {
unsigned regs[4];
cpuid(1, regs);
// You could name 23 CPUID_POPCOUNT_SHIFT or something.
return ((regs[2] >> 23) & 1);
}
In normal C compilers (GCC, Clang) you can use ifuncs to do this resolution once at module load time, and then the correct function for the current CPU is dispatched cheaply afterwards.One way to do it in C++, implement dispatching at higher level, where the function being dispatched takes much more time than a call overhead. Sometimes C++ templates or even C preprocessor can help reduce copy-pasting.
Another way to do it is JIT from byte code. .NET tries to do that, I don’t think it’s too efficient currently, but I believe couple major versions in the future it might actually become good. It doesn’t have performance budget for expensive optimization modern C++ compilers do in optimized release builds. OTOH there’re other optimizations unavailable in C++: JIT may use whatever max AVX version is present, may compile immutable or rarely changing input data into code.
I think it's equally applicable to any processor vendor, x86 has just been around the longest.
Meanwhile any ARM64 CPU will have CNT.
https://developer.arm.com/docs/ddi0595/h/aarch64-system-regi...
Edit: Also there's a ton of them, this lists a couple: https://www.kernel.org/doc/html/latest/arm64/cpu-feature-reg...
- You'll get into trouble when migrating a process between cores in heterogeneous environments (having cores that support different ISA extensions is relatively uncommon today, but Intel has already announced some, and people have gotten into trouble in the past with e.g. reading cacheline sizes when big and little cores used different values).
- This can be a pretty slow operation on some platforms, and the system can make it more efficient by caching it (and can provide additional capabilities information that may not have existed on older CPUs so that there's a uniform interface).
i.e: If a process asks "can I use SVE2" and gets a yes, the OS doesn't move it to any core for which that isn't true.
Unrelated: what does Intel Sports do?
Isn't the "Intel vs GNU syntax debate" only on x86? For ARM, AFAIK there's only one assembly syntax.
https://twitter.com/wongmjane/status/1275177255681982464?lan...
My interest in popcount is for radix trie node compression https://dotat.at/prog/qp/ so these bulk popcount hacks are interesting but not directly useful for my purposes :-)
https://golang.org/pkg/math/bits/#OnesCount
go compiler output...
0x0020 00032 ($GOROOT/src/math/bits/bits.go:115) FMOVD R0, F0
0x0024 00036 ($GOROOT/src/math/bits/bits.go:115) VCNT V0.B8, V0.B8
0x0028 00040 ($GOROOT/src/math/bits/bits.go:115) VUADDLV V0.B8, V0
__builtin_popcount()
... which gcc and clang will replace with best available.'nuff said.
Edit: or OnesCount() in Go.