According to Butler v. Target Corp., it was held that although lyrics to a song are copyrightable, the underlying voice is not. As such, there is no copyright protection available to the infinite number of words or phrases a person might utter in their distinctive voice.
Additionally, the synthesized audio can be considered derivative, as it transforms the the "audio" into something entirely different than original, and so falls under 17 U.S.C.A § 103.
So, I'm not sure what you mean when you say there are plans for tighter controls. Care to back that up?
Disclaimer: I am not a lawyer and this is my personal opinion.
What's the difference between using an AI voice and Bill Hader or Jimmy Fallon doing a celebrity impersonation on his show and monetizing that?
The AI image generation space has hundreds of players. Audio has dozens.
One likely outcome is that big tech will come to each of the "successful" companies with close peer competitors and offer to buy them. If they say no, they buy their competitor. Or build it internally.
You'll have to run really fast and hard to survive. I think it's totally doable, though, and this is a very interesting attack gradient.
Best of luck! It's exciting times.