Great product, giving it a try. Here you saying that 20 seconds is enough, and on a "clone" page there is an instruction about 30 minutes for better result. Is there any kind of instruction about how to create a good sample of the voice? For example, should I speak English, or any language will do? Do you have some stats on corellation between sample length and generation quality? Thank you!