The server could pick say 5 random offsets in the song, and ask the client to upload the decoded data from one second of music starting at each of those offsets. The server can then compare that to what it has and see if it is close enough.
Or simpler still - the iTunes application will generate fingerprint of the whole audio in each file and send it along with whatever metadata is available. Then serverside the fingerprint and metadata are analyzed and if a certain criteria for a match are not met, iTunes is prompted to upload the file.