Moving more of our stack into Rust wasn't viable, so we shelved the project and focused on other areas.
Turns out this is a pretty common experience.
Now we're trying again, but this time with our own binding protocol. It’s 2.5× faster at the boundary, but more importantly, it shifts the break-even point dramatically. Offloading to WASM becomes viable at much smaller data sizes (e.g. ~100 elements), enabling cases that weren’t previously worth it — including ours.
The only trade off is significantly worse DX - the bindings are manual and more verbose, but I'm sure someone could solve this with codegen/macros on either side. For now we're focused on performance.
It's pretty much neck and neck between our Rust and Zig implementations so take your pick.