I've built a speech synthesis system with Marytts before, it works, but unfortunately, the quality is not very good, HMM and unit selection speech synthesis are now very old approaches, they are far from the current state of the art, you should try open-source implementations of Tacotron2 or wavnet, you will surely achieve better quality.