The cost of FFI in Go is a cgo callgate regardless of whether you use Cgo, I can point to the actual implementation if need be. The cost from going from Go execution to C execution is ~tens of nanoseconds, so it's not particularly easy to measure. I would guess its on the order of hundreds of CPU cycles. I've stepped through it in IDA from time to time, it's not a ton of instructions. Of course, "not a ton" is a shit load more than "literally zero", but it's worth talking about I think.
There is a potential for greater cost, because with the way the Go runtime scheduler works, if the call doesn't return quickly enough, the thread has to be marked "lost" and another is spawned to take on Goroutine execution. This happens rather quickly of course and it isn't free since it results in an OS thread spawning. But this is actually very important, more on this in a moment...
Meanwhile C# async uses the old async/await mechanism. This is nice of course and it works very well, but it has the problem that execution does not ever yield until you await, and if you DO call a C function and it blocks, unlike Go, another thread does not spawn, that thread is just blocked on C execution until it's over. That was my experience playing with async in .NET 7, but I don't think it can change because either you have zero-cost C FFI or you have usermode scheduling, you can't really get both because the latter requires breaks in the ABI.
I would be happy to talk more because I am honestly pretty disappointed that there's not really a better way to do what Go tries to do. I'd love to have the advantages of Go's usermode scheduling with preemption and integrated GC sequences with somehow-zero-overhead C calls, but it simply can't be done, it's literally not possible. You can take other tradeoffs, but they lose some of the most important advantages Goroutines have over most other greenthread implementations. Google and Microsoft have both produced papers researching this. Microsoft's paper on fibers basically comes to the conclusion that you literally shouldn't bother with usermode scheduling because it's not worth the trouble:
https://www.open-std.org/JTC1/SC22/WG21/docs/papers/2018/p13...
However, their conclusion about the cgo callgate taking around ~130ns does not match what I've seen. But just to be sure, I searched for a random benchmark and found this one:
https://shane.ai/posts/cgo-performance-in-go1.21/
BenchmarkCgoCall 28711474 38.93 ns/op
BenchmarkCgoCall-2 60680826 20.30 ns/op
BenchmarkCgoCall-4 100000000 10.46 ns/op
BenchmarkCgoCall-8 198091461 6.134 ns/op
BenchmarkCgoCall-16 248427465 4.949 ns/op
BenchmarkCgoCall-32 256506208 4.328 ns/op
Of course this may be fairly optimal conditions, maybe it depends on the conditions of the goroutine stack leading up to it, but I think it's fair to say that "less than 50ns per op" is not an unreasonable amount of time for the cgocall to take as long as we're considering that "0ns" is strictly not an option for what Go wants to achieve. With Go you don't have to care if something blocks or not; everything blocks and nothing is ever blocked. That's not something that can be accomplished without some runtime cost. The runtime cost that it actually takes is very nearly zero, but the runtime cost of integrating that with something that doesn't eat that cost is unfortunately higher, and that's where the CGo problem lies.(I admit that a substantial portion of this problem is actually around the stack pivoting, but if you squint hard enough you can see that this is also inextricably woven into how Goroutines manage to accomplish what they do.)