This is an interesting take, and I'd guess that the training data for this probably did use podcasts as a source.
Getting very realistic / real world conversational training data for an ai would be hard. Only a subset of us appear on podcasts, radio or tv and probably all speak in a slightly artificial manner when we do.