Can you think of a test or empirical observation that would convince you that a model does "actual reasoning"?
Do bear in mind the tests that very smart people have proposed in the past (from playing chess to holding a conversation to understanding pictures). And consider the implication for your position if you find it hard to devise a well defined test that would convince you to abandon this position.