I don't know what type of speech each dataset represents, but the talon results are extremely impressive... I assume it wasn't trained on at least some subset (depending on the train/test split) of this data?
Due to whisper's weakly supervised training on a large amount of automatically scraped data and reliance on a bigger language model, it's far more likely whisper had seen some of the test data before.