It's not really style transfer, but for a new speaker model, you just need to train each speaker with a dataset of 25 hours audio with time matched accurate transcriptions.
In the case of David Attenborough, I'm pretty sure that amount of data is available with subtitled BBC content.