The short answer is that everything is streaming — as tokens come back from ChatGPT we send them as soon as possible to the synthesizer. The long answer is found in our code[0] :).
[0] https://github.com/vocodedev/vocode-python/blob/main/vocode/...
[0] https://github.com/vocodedev/vocode-python/blob/main/vocode/...