I think this is a bit of a misconception. A critical element of Forth is that there are two stacks: a data stack for operands, and a return stack for keeping track of what word to return to after each call completes (and other fun stuff). These stacks grow and shrink independently. Unused operands on the data stack quietly flow along to the next word that takes a peek with no overhead. There's no difference between how the stacks are used within an "expression" and between "function calls"- these concepts don't exist at all like they do in curly-brace languages!
In a typical non-Forth "stack VM", there is a single stack which behaves like the C stack: allocate an activation record with locals on a function call, pop the whole thing off at the end. Intermediate calculations go on top of all this, avoiding the need to explicitly register allocate in bytecode, but operands don't naturally flow between function definitions; they're copied around explicitly as the stack grows and shrinks.