(Everybody seems to assume "high performance" automatically means in whatever domain they work in. Impossible to interpret these sorts of questions. Try "Will BQN have high performance for ...?" instead.)
I was wondering mostly about the FFI overhead. So the question should really be, is it suitable for real-time low-latency contexts? But I have your answer :)
I still would love to find a use for it one day. If there's ever a chance you could go the futhark route and allow for compilation of CPU/GPU routines I think it would be an ideal kind of syntax.
But as currently there isn't any loop fusion in CBQN (though I'd definitely like to add such at some point), despite the native ops being all nice SIMD where possible, it can still lose to autovectorized C or similar, or even scalar code, due to memory overhead.
FFI is libffi currently (with non-JITted preparation), so on the order of 100ns per call.