Apart from the total lack of external validation (which I suspect will never come), we already know that ECGs have features associated with emotion: heart rate, frequency of premature ventricular contractions, frequency of premature atrial contractions. Are these deep learning detectors better than running a simple regression after using something like a wavelet transform to detect these specific features? This matters because if you have a method that is more computationally efficient, more easily interpreted, and more robust to extrapolation to out-of-domain data, why in the world would you go with DL?
Of course the article is completely silent on these questions. If I was a reviewer I would refuse to accept until they added extra experiments to address this.
This is a symptom of a more general problem in interdisciplinary medical research: Papers are written and reviewed either by engineers who have limited understanding of the clinical background or by clinicians who can code a CNN in pytorch but have no knowledge of classic ML methods or ML theory. And so we see seriously flawed papers get accepted. (I know this is a preprint, but I've seen similar stuff in well regarded journals.)
It's incredibly frustrating as someone who tries to do truly rigorous research in this space, only to be desk rejected because ECG magic sounds a lot cooler than incremental gains.