Yes, there are functional gaps between MLLMs and humans. Their long-term memory is an external mechanism that can use RAG-like approaches, context compression or something like that. The models have problems managing those.
The models can't do continual learning. Although there are promising directions (expert cloning in MoE models, and others).
The only mode of learning available to a model while working on a task is in-context learning. This limits the models to concepts that they developed during autoregressive pretraining and the later stages of training. That is a model can't create new concepts as a result of working on a task (the model's maintainers could choose the task to be represented in the training data later though).
But it's all about functionality.
I guess you have the Leibniz's mill intuition. We can look at how those things work, and there are no experiences or intelligence in sight.