Sorry to hijack, but since you are involved, can you explain why tail call optimization would incur a run time perf penalty, as the docs mention? I would expect tail call optimization to be a job for the compiler, not for the runtime.
But the good news is that the common case incurs no overhead.
I am trying to think of a situation where a functional language compiler does not have enough information at compile time, especially when effects are witnessed by types.