I read somewhere that Tom's Diner is used for lossy audio codec evaluations because it is a clear, strong a capella voice. It's been said that humans are sensitive to even minor distortions in vocal quality. You can sing a capella in any language, and it is not limited to western music. Sorry, I don't have references to offer.