Eg, can you take a webcam image analysis software and teach it that apples are red without ever showing it something red through the camera?
I'd say the answer is obviously dependent on the design.
Yes if you have enough control over the software. You can introduce the shape of an apple into its dataset, and tell it that an apple shape is such and that it rates highly on the R channel. Then the program will upon seeing an apple for the first time say "Yup, this is indeed a red apple" and be entirely "unsurprised".
No if the system is isolated enough. Eg, if the recognition system is locked in a black box and your only controls are a query of "what is this?" and a command of "remember that the thing you're looking at now is X", then you don't have the access to introduce the information above by any other means than presenting an object to the camera.
Both systems are still 100% physical though.
The question is what humans are most similar to and I think the second option is most likely. Think of that for instance I can't mentally measure my blood pressure, or heart beat, or make myself vomit. There are some things my body does that are outside of my direct control. Given that, there's no reason to assume then that by just reading or listening to information that I would necessarily have any pathway from there to the depths of my visual system.