What would be your use case?
I see lyrebird api being very helpful in helping my users practice listening skills and add a level of creative fun! If we had 10-20 different voices, the flashcards will be read a little differently each time. Right now (since our flashcards is dynamic), our audio feels very monotone. We would love to help you beta test your API and work something out.
Our current production process requires a group consisting of editors, readers and technicians to get together every Friday morning from 7am to record an hour or more of news which is then mastered onto CD, duplicated hundreds of times and mailed out by 11am.
We usually have four readers each week (from a pool of 30 or so) who take it in turns to read the items. Some readers are better than others and sometimes readers don't turn up. Sometimes there are interruptions or disturbances to recording such as another reader in the studio coughing, rustling of papers, etc.
If we had the ability to digitise the voices of our readers it would enable our new (in development) totally digital production and distribution system (podcasts, streaming, etc.) to be produced at any convenient time and to allow our listeners to choose their preferred reader's voice(s).
The studio software side is using FL/OSS software, with Ardour as the digital audio workstation attached to a Delta 1010 digital input system and an Evolution UC33e control surface.
Being able to program the pre-production phase to generate the audio recordings using favourite readers voices which are then fed to the (automated) studio mastering process would give us some amazing functionality and flexibility to produce programmes on-demand with no studio presence required.
The development experience and final package will be documented and published for other talking services to adopt and adapt.