Probably not something we'll get to hear as part of the PR pitch.
Or is the consent statement the thing that will be cloned and is there no separate training audio? Then it might actually work and you'll just have to get close enough that the human you're trying to fool can't distinguish anymore (defeating the need for this tech in the first place, at least in targeted rather than automated cases).
I like your idea of just training on the consent text! That wasn't the case when I tried it as you needed around 3h (optimally) of training data.
If you mean the white noise, I meant that as a brute force attack because, to do it more targeted (to know what it'll accept as seeming like your target voice), you'd likely need their exact model rather than doing your own.