I don't understand why there is any doubt about this. From OP, section "Methods", subset "Data set creation and splits":
"The training, validation and test data sets are generated without overlap from periods in sequence. Successive periods of 400, 12, 40, 40 and 12 h are used to sample, respectively, training, validation, and test data, with the two 12 h periods inserted as hiatus."
I know nothing about machine learning, but I was in grad school with others studying machine learning and I took a Coursera course on the subject, and carving out a training set is utterly standard practice. A paper that didn't do this would get filtered by a grad student reviewer. So I'm not sure what possible sub-category of machine learning your statements could apply to. Can you share any examples? For example, do any of the famous papers in the field fail to hold out test data? Something like AlphaGo or AlphaZero is indeed testing on all new data -- new games it plays with others. Do any papers in well-known machine learning fora fail to hold out test data? Does anything that has bubbled to the front page of HN? It would be interesting to look at if so.
Do you mean p-hacking? That's a much more subtle thing.