I was always curious if you trained a model on a literal visual representation (pixels/image) of the charts or candlesticks, would the model be able to “see” something that we can’t.
Of course not. It's the same data but presented in a harder to process format.
It may be the similar but as they say, a picture speaks a thousand words, the visual features that a CNN might pick up could be something completely different than the features someone could think of. It is all about data representation. Hypothetically, the data representation shouldn't matter, but I think it is like viewing the optimization surface from a different angle, it is possible to get something different out of it.
This is something I got curious about too. It's very likely the answer is no, but I would like to test it at some point.