We ran the benchmark here and confirmed the relative times, but our profile shows runtime.addspecial as expected. The expectation is that finalizers are used rarely, for objects where the lifetime is unknown but typically long. In contrast, a finalizer benchmark is hammering on finalizers in a way that real apps basically never do.
There are also some O(N) things in finalizer registration, which turn into O(N^2) things for N finalizers for nearby allocations (only up to N=1024), again because we expect them to be very rare. Those could be fixed if this use case of tons of finalizers is realistic in practice.
If you have locally scoped C data you would typically defer C.free(p), and if you do that the cost is basically identical to calling free(p) directly.
BenchmarkAddition-16 22961580 51.77 ns/op
BenchmarkAllocate-16 7611168 144.3 ns/op
BenchmarkAllocateDeferFree-16 7065505 144.3 ns/op
BenchmarkAllocateFinalizer-16 1028251 1243 ns/op
In response to the "cgo is slow" questions, this shows that cgo calls are about 50ns on my four-year-old x86 MacBook Pro. Is that fast or slow? It depends on what the cgo call is doing. If it's executing a single add instruction, an extra 50ns is slow. If it's doing something more substantial, an extra 50ns may be nothing at all.