- detection: just separating signal from background with confidence
- classification: confidently separating different signals in noise (e.g. PCA)
- identification: requires ground truth (perfect example), but still difficult to do with confidence
If the scientists have ground truth, e.g. "this call means [xyz]", then it’s a signal processing problem to do identification.
If there is no ground truth, then AI would have to have a lot of background data on whale activity at different locations including visual and audio content to try to infer (and that's about it) what the vocalization might mean. And this is where science starts to leave the building.
It's kinda' like inventing the Rosetta Stone using AI: probably not going to be very accurate or repeatable.