1. How did they split the data for train/validation/test sets? If all of the sets had the data from the same time period, or the same companies end up in multiple sets, it is a major flaw. For example, if outperformance was persistent for the companies over the time period considered, the model may simply learn to identify specific companies by their filings.
2. What is the variance of the out-of-sample performance? Given that their dataset is very small, and the model performed badly at predicting high returns and reasonably well at predicting good returns, what are the chances of getting those results by luck alone?
3. How has the model performed since then, on the most recent filings?
4. Why use a convnet? Would gradient boosted trees not perform just as well/better but be more interpretable? Methods like the ones in the eli5 package can help get an idea of why the model makes a particular prediction, which could help sanity check the model.