The tool uses ffmpeg to load the video, and according to the "anatomy" section it's based on mel-frequency cepstral coefficients, so it's only using the audio to do the alignment.
Feeding it an mp3 might "just work".
Feeding it an mp3 might "just work".