This matches how i think about realtime voice: keep call control deterministic,
and give model steps clear inputs and outputs.
How would you handle retries when a step fails ?
I build low-latency voice gateways .
Does your local API expose interim transcripts
and timestamps while audio is arriving ,
or only the final text after a dictation ends ?