So just feed it batches smaller than 1000 characters? It's not like TTS requires maintaining large contexts at a time.
So just feed it batches smaller than 1000 characters? It's not like TTS requires maintaining large contexts at a time.
The simplest examples are punctuation marks which change your speech before you reach the mark, but the problem extends past sentence boundaries
For example:
"He didn't steal the green car. He borrowed it."
vs
"He didn't steal the green car. He stole the red one."
A natural speaker would slightly emphasize steal and borrowed in the 1st example, but emphasize green and red in the 2nd.
Or like when you're building a set:
"Peter called Mary."
vs
"John called Mary. Peter called Mary. Who didn't call Mary?"
-
These all sound like small nits but for naively stitched together TTS, at best they nudge the narration towards the uncanny valley (which may be acceptable for some usecases)... but at worst they make the model sound broken.
I agree, but it seems unusual for this to matter past paragraph boundaries, and it sounds like there should be enough room for a full paragraph of context.
And the current SOTA for TTS includes breathing too, so you can't just put a fixed empty pause between your paragraphs.
People are chunking by paragraphs anyways (or even sentences) and it works, but the top commercial models support maintaining a context or passing in the most recently generated text for that reason.